[06:25:01] greetings [06:53:29] I'll reboot cloudcumin hosts shortly for T434750 [07:18:56] morning! [07:19:05] paws seems to be having issues, looking [07:21:04] two of the nodes are down, rebooting [07:21:15] cheers [07:24:59] hmpf... the keys for the worker nodes are not up to date in the cloudcontrol nodes, so can't ssh [07:25:41] wait, maybe wrong user [07:25:47] yep they work :) [07:50:16] things are back online for paws \o/, looking into the puppet alert for tools [07:50:40] ack [07:59:08] godog: there's many canary vms not working (or reporting down/alerting), is that part of your reboots? (should not be related to cloudcumin I think) [07:59:57] dcaro: part of the cloudvirt rebalance yes - T431682 I'll silence once I'm done with the last host decom [07:59:57] T431682: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682 [08:00:05] ack [08:02:50] !log tools truncate haproxy.log to 1G (out of space) on haproxy-8 [08:02:50] dcaro: Not expecting to hear !log here [08:02:58] wrong channel xd [08:17:55] morning [08:35:09] NovaComputeUnavailable page is me btw [08:35:36] ack [08:36:34] godog: thanks, acked the page in victorops [08:37:43] ack [08:41:22] unless there are objections I'd like to roll out some patches moving the TLS termination and rate limiting on the cloud vps proxy from nginx to haproxy [08:43:38] +1 [08:57:07] and done, and sent emergency rollback instructions to cloud-admin@ just in case [08:58:40] \o/ \o/ \o/ [09:55:12] I realized NovaComputeUnavailable should not have paged earlier, https://gerrit.wikimedia.org/r/c/operations/alerts/+/1326225 is the fix [10:12:50] * godog errand and lunch [10:19:01] shipping https://gerrit.wikimedia.org/r/c/operations/puppet/+/1326240 to fix x-forwarded-proto with the new haproxy setup, apparently that's causing some projects to send infinite redirect loops to clients [10:24:23] dcaro: hello, now seems like a good time for me to deploy https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1371 to toolsbeta then tools, wdyt? [10:27:22] thilp: you might need https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1373 [10:28:53] I'm about to go for lunch too, so might not be around until later. Toolsbeta sounds ok, I'd wait for tools until after lunch (we have the pairing slot too if you want) [10:41:51] * dcaro lunch [11:34:18] * dcaro back [11:37:47] I’ll start with toolsbeta now then, thanks! [11:42:09] 👍 [12:21:18] can I get a review of https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 ? I will want to do a new release of the components-cli later today, and that would be the perfect test for the auto-creating releases ci [12:26:40] thilp: oh, just overwrote your commit to api-gateway https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1365 [12:27:00] let me know if you are ok with the new code, I'll copy at least the comment [13:48:28] I found an archived mailing list thread about the issue I'm investigating and then the archive site went down just before I got to the good part :( [14:37:34] Ok, it's back! I'm not the only one getting sidelined by new ceph performance alerts: https://www.spinics.net/lists/ceph-users/msg86138.html [15:36:28] dcaro: did you ever spend time with bluestore caching? It looks to me like we're configured to auto-resize the cache and allow caches up to 6G per osd, but in reality it never uses more than 3 per OSD and I wonder why. [15:37:09] I remember doing some tuning at some point to make it match the osd memory sort of [15:40:26] I can't convince myself that any of the levers are doing anything. We're running ours OSD nodes at 50% free RAM so it feels like there should be some performance gains somewhere in there... [15:41:31] the docs also don't really say if auto or manual tuning takes precedence [15:44:04] Raymond_Ndibe: wm-lol now uses the publish option when deploying using push-to-deploy \o/ [15:47:28] andrewbogott: is our setup different enough from prod that a consult w/ Ben may not be fruitful? Wondering if they are noticing similar issues [15:49:07] their cluster isn't nearly busy enough to trigger any of the things we're seeing. Ben might have thoughts about specific settings though. [15:52:51] Raymond_Ndibe: and now a webservice got push-to-deployed xd https://gitlab.wikimedia.org/toolforge-repos/wm-lol/-/jobs/930830 [16:26:19] * dcaro off [16:26:22] cya tomorrow [17:24:59] * dhinus off [18:07:03] aputhin: officially she can self-approve but do you mind adding a +1 to https://gerrit.wikimedia.org/r/c/operations/puppet/+/1326343 and/or https://phabricator.wikimedia.org/T435123 ? [18:07:36] (this is to get Belinda shell access on some wmcs infra servers) [19:21:43] andrewbogott: did you intend to send that mail to cloud-admin@ instead of cloud-announce@? [19:22:38] nope. resending. [20:07:45] Updates to the deprecation playbook here: https://wikitech.wikimedia.org/wiki/Portal:Cloud_Services/Admin/Deprecation_playbook