[00:37:36] that's awesome if true [07:24:55] greetings [09:37:57] morning. I'm looking at the toolsdb replication alert [09:41:49] the replica is running "OPTIMIZE TABLE entry" on mixnmatch, it will likely complete on its own [10:42:03] turns out the wmflabs.org->wmcloud.org redirect broke when moving to haproxy, fix is https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329527 [10:43:35] (seems like I only tested that redirect working before moving the 404 page from nginx to haproxy :/) [10:44:48] taavi: nice find, +1d [10:46:01] * dhinus wonders if we can fully deprecate wmflabs.org at some point [10:46:29] (well deprecate is the wrong term, I mean drop all the redirects) [10:47:20] not in the short term at least [10:47:29] (especially if there are still services using wmflabs.o canonically) [10:47:53] would it make sense to have a big tracking task like the one we had for migrating away from icinga? [10:48:29] not saying this should be a priority in the short term, but maybe in the medium term it could? [11:01:57] i've just been letting folks migrate naturally, at some point we might want to do a big push to get the remaining ones migrated but i don't see a big rush for it [11:13:11] with that wmflabs.org issue fixed, rolling out the changes to proxy-6 as well [11:46:05] andrewbogott: godog: T435571 is apparently asking for a local storage VM. I'm tempted to say no given we can't really offer any uptime guarantees on those [11:46:05] T435571: Provide more disk I/O headroom for the Castor cache volume (integration-castor06) - https://phabricator.wikimedia.org/T435571 [12:49:47] taavi: indeed, I think a cinder volume with proper limits/quota (as opposed to the instance flavor that is) seems the way to go [12:54:28] I got https://gerrit.wikimedia.org/r/q/topic:%22bug/T435581%22 out to clean dumps-nfs symlinks and friends [12:54:44] ngl feels very good to clean that stuff up [12:54:47] looking [12:56:48] cheers [13:17:39] the redis alert is for a job that should not be there anymore, looking into why it still is there [13:22:29] speaking of alerts, the toolsdb one resolved on its own as expected, replication is back in sync [14:45:39] andrewbogott: ready for the k8s upgrade when you are [14:46:09] ok! Starting with irc... [14:46:45] * taavi wonders if anyone actually reads the status from the topic [14:47:15] I would guess there are two or three people who read it :/ [14:48:01] OMG I updated to the OIT recommended version of macOS and now it alerts and interrupts me when I paste into a terminal [14:48:07] Anyway... now about to run [14:48:08] sudo cookbook wmcs.toolforge.k8s.prepare_upgrade --cluster-name tools --dst-version 1.33.13 --task-id T433132 [14:48:10] T433132: Upgrade tools to Kubernetes version 1.33 - https://phabricator.wikimedia.org/T433132 [14:49:40] ...and it failed [14:50:13] "{"errors":[{"code":"PRECONDITION","message":"the execution 38543153 has tasks that aren't in final status, stop the tasks first"}]}" [14:50:53] uhh what fails on that? [14:52:40] I also don't love the 'missing' in this line... [14:52:41] | api-gateway | chart | api-gateway | missing | | [14:52:51] taavi: you want a paste of the whole output? [14:53:05] i want anything that includes any more context [14:54:25] suddenly irccloud won't let me paste snippets. So... https://phabricator.wikimedia.org/P96261 [14:54:50] ok, that's from harbor [14:55:07] dcaro: ^ [14:55:41] that cookbook sure seems to do a lot of things in dcaro's homedir :/ [14:56:44] want to just try again? [14:56:53] sure [14:57:27] I'm on PTO, anything broken? [14:57:49] ah sorry, then go away [14:57:51] taavi: did you change something or is this literally trying the same thing? [14:57:56] The tests run as my user yes (the admin part), there's a task to change it [14:57:58] the latter [14:58:10] ok. It's still thinking but seems better this time. [14:58:29] * andrewbogott wants every cookbook's help message to tell me if it's idempotent or not. [14:58:46] if it got past that test, then the thing to do is to file a task to sort out why that test is flaky [14:58:55] Yesterday I deployed stuff, and the tests were passing, so if they are broken it's new-ish [14:59:18] dcaro: seems like a one-off but I'll open a task [15:00:23] 👍 [15:00:28] * dcaro off [15:00:31] d.caro shoo shoo, go to your vacation! [15:01:24] created T436126 [15:01:25] T436126: Surprise harbor failure in wmcs.toolforge.k8s.prepare_upgrade - https://phabricator.wikimedia.org/T436126 [15:02:54] still running but all green this time [15:18:29] finished, now about to run [15:18:33] sudo cookbook wmcs.toolforge.k8s.worker.upgrade --task-id T433132 --src-version 1.32.13 --dst-version 1.33.13 --cluster-name tools --hostname tools-k8s-control-7 [15:18:34] T433132: Upgrade tools to Kubernetes version 1.33 - https://phabricator.wikimedia.org/T433132 [15:18:43] ok [15:21:28] hmmmm [15:21:32] https://www.irccloud.com/pastebin/yJf57kCN/ [15:22:10] ok, let's have a look [15:23:45] hm, so `kubectl get events -n kube-system` says this https://phabricator.wikimedia.org/P96262 [15:23:54] and of course the actual Job and Pod objects are already gone [15:26:46] andrewbogott: found this https://github.com/kubernetes/kubeadm/issues/3282 [15:27:08] I'm in the spicerack logs, is this error from 'sudo -i kubeadm upgrade plan 1.33.13' or from whatever happens after that? [15:27:35] ok, that bug report suggests 'just try it again' again [15:27:36] it will prompt you whether to approve a plan after running that, did it do it for you? [15:27:43] no, didn't get that far [15:27:50] ok, that's safe to run again for sure [15:27:56] great, here goes [15:28:10] plus it says 'preflight' which also suggests it didn't do anything real yet [15:29:51] failed again? [15:30:15] nope, working this time. I pasted the plan into the task, about to 'go' [15:30:32] cool. plan lgtm [15:30:43] fwiw this is what i saw for the job: https://phabricator.wikimedia.org/P96263 [15:30:59] and timed out again [15:31:19] I guess I can try yet again, the upgrade didn't get anywhere before failing [15:31:28] yeah, +1 [15:31:45] seemingly the timeout is much increased in newer kubectl versions so as long as we can get one to work this upgrade then it should be fine [15:32:31] draining is a lot faster the third time around [15:33:35] seems like it is doing its thing now? [15:34:07] yeah, it's actually upgrading things [15:35:56] hm the kubelet on -7 seems unhappy [15:36:49] yeah 'restarting static pods' is taking forever [15:37:02] more specifically, the api server on that box is unhappy which is making everything else unhappy [15:37:11] although... some things report that they're running [15:37:14] I do note the pod is stuck in a Terminating state [15:37:33] huh, we've lost the api server yaml file again??? [15:37:38] `.cookbook-stopped-kube-apiserver.yaml` [15:38:25] ok, yep, cookbook failed with 'we somehow failed to stop static pod kube-apiserver' [15:38:36] I had to manually kick it a bit [15:38:48] the cookbook failed entirely? [15:39:09] yes. [15:39:12] https://www.irccloud.com/pastebin/0EvWgBNi/ [15:39:27] ok :/ [15:39:41] so it's probably the same issue as in toolsbeta again [15:40:10] the good news is that this is the very last step of the cookbook, so you can run `cookbook wmcs.toolforge.k8s.restart_static_pods` agains that node to try it again and then we can try upgrading the next node [15:40:15] seems like. I guess we had two issues there and not one. [15:40:21] since the node itself is healthy now [15:40:25] did you copy over the server yaml already? [15:40:57] yep, it's back now [15:41:46] sudo cookbook wmcs.toolforge.k8s.restart_static_pods --cluster-name tools --task-id T433132 --hostname tools-k8s-control-7 [15:41:47] T433132: Upgrade tools to Kubernetes version 1.33 - https://phabricator.wikimedia.org/T433132 [15:41:57] So that yaml file is managed by kubeadm, not the cookbook right? [15:42:18] yes (except the cookbook moves the file back and again to do the restart) [15:42:38] oh, so does that mean it /was/ the cookbook failing to replace the file? [15:43:24] so what the cookbook does is moves the file away, waits for the pod to stop, and moves it back, to restart the pod [15:43:32] "waits for the pod to stop" failed [15:43:48] ah, ok [15:44:16] we have the same problem [15:44:28] yeah [15:44:29] did something in kubelet change? [15:45:06] so what's failing is kubelet is trying to talk to the api server to tell it about the api server pod stopping. which obviously does not work if you just stopped the api server [15:45:20] aha [15:45:32] aha? [15:45:41] so in -8, which has not been upgraded yet, the kubelet kubeconfig file references the haproxy load balancer address [15:45:58] but on -7, that's been changed, presumably by kubeadm, to only refer to -7 itself [15:46:12] I was just going to ask... haproxy now has... [15:46:30] well that explains [15:46:56] it does but that's a bug in kubeadm isn't it? [15:47:07] yes and no [15:47:16] us moving the file around is already a hacky enough thing to do [15:47:58] (for reference, we restart the pods to avoid a race condition with controller-manager and scheduler trying to talk to apiserver before apiserver is up. restarting them all, in alphabetical order, ensures tha apiserver is up before the other two) [15:48:25] why do we move the file around rather than letting kubeadm do its thing during the upgrade? [15:49:17] what do you mean? [15:49:42] nevermind, I'll ask that later [15:49:52] for now -- how did you unstick things? [15:50:22] `mv /etc/kubernetes/manifests/{.cookbook-stopped-kube-apiserver.yaml,kube-apiserver.yaml} && systemctl restart kubelet` [15:50:39] ok [15:51:40] so https://gerrit.wikimedia.org/r/c/cloud/wmcs-cookbooks/+/1329597/ is what we need to continue, i think [15:52:10] seems fair [15:52:41] ok, let's merge that, pull it in to cloudcumin, and then move on to control-8 [15:53:02] The upgrade guide wants me to first depool 8 and 9 and make sure everything is happy with the upgraded controller [15:53:06] Are you patient with that? [15:54:35] i don't think that's ever been really necessary at this stage, but if you want to do it i don't mind [15:55:03] I'll feel better if I do it! And it's good practice. [15:56:15] ok, everything is forced to -7 now [15:58:15] ok, patch merged [15:58:28] running tests, which will take a while. Good time to make a coffee if you need one. [16:12:43] tests tests tests tests tests [16:18:42] ok we're back! I'm going to restore the haproxy to use all three controllers and then upgrade -8 [16:21:58] taavi: the api server restarted properly this time [16:24:20] static pod restarts completed [16:24:49] now -9 [16:31:38] And now...worker nodes [16:37:15] how many batches are you running? 3? [16:38:36] So far i'm just doing the non-nfs workers, all in one batch. The nfs workers I'll split into 3 [16:38:44] ah cool [16:39:18] while the workers are running, you should go and upgrade the gateway nodes one by one. I didn't get T426948 done yet, so the main thing with them is to be veeery slow to move on to the next one after one of them is finished [16:39:18] T426948: wmcs.toolforge.k8s.reboot needs to be slower with Gateway nodes - https://phabricator.wikimedia.org/T426948 [16:41:21] ok [16:42:03] there are only the two classes of worker nodes, right? nfs and non-nfs? [16:42:35] you could count the gateways as a third [16:43:05] oh, sure [16:43:05] sudo cookbook wmcs.toolforge.k8s.worker.upgrade --task-id T433132 --dst-version DST_VERSION --cluster-name tools --hostname tools-k8s-gateway-1 [16:43:06] T433132: Upgrade tools to Kubernetes version 1.33 - https://phabricator.wikimedia.org/T433132 [16:43:32] oh oops needs dst_version [16:46:04] how long should I wait between gateway nodes? [16:46:43] enough for the pods in the `istio-gateway` namespace to come back up and then be noticed by haproxy as being back up [16:55:57] ok, gateways and non-nfs nodes updated, nfs nodes underway but only about a third of them are done [17:20:16] taavi: I'm running the test suite again but we're all done. Thank you! [17:20:22] \o/ [17:20:36] now we get to do all of that again for the next upgrade [17:21:02] yeah. I'm tempted to just start .34 immediately but I guess I have other things I ought to work on first. [17:24:38] congrats! [17:29:28] https://prometheus-alerts.wmcloud.org has really exploded over the last couple of weeks; it's probably fallout from bullseye deprecation but I'm going to spend a bit of time looking for trends there. [17:30:22] taavi: I assume the 'frequent restarts' alerts are from all the node draining, but what about these? https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolWebserviceErrors [17:30:44] is that just normal 'sometimes user's tools are broken' stuff? [17:31:05] hm, not the first one :( [17:35:55] error rate is decreasing but I don't understand why they were down for so long during the upgrade...