[06:57:04] slyngs: if you're around, could I trouble you with reviewing https://gerrit.wikimedia.org/r/c/operations/puppet/+/1319853 please? [06:57:08] thank you! [06:58:07] Seems like a nice to start the day on :-) [07:01:10] brouberol: Techincally I think we need have it run by I/F because your expanding sudo privilegeres, so it needs an approval stamp on the ticket [08:13:32] Any sre to give me an ok for this plan? https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321123 [08:25:04] jynus: don't know the service, does it support sd_notify to systemd? [08:25:49] otherwise I'm afraid `Type=notify-reload` doesn't work as intended [08:26:19] or just use something like `ExecReload=/bin/kill -HUP $MAINPID` [08:26:52] slyngs: yep, I added m.oritz as a reviewer but he seems to be away [08:27:04] no rush though [08:28:31] fabfur: yeah, that was the alternative [08:29:44] left a comment on that, mainly the upstream software must support sd_notify or systemd will never know if the software has been reloaded or not [08:30:49] basically I think systemd will hang or fail when trying to reload (even if versutygw will be reloaded under the hood) [08:34:05] makes sense, yeah, I could test it outside of production [12:57:46] Is there a way to avoid alertmanager slack notifications to repeat for firing alerts, sort of like we do for Phab where it creates a ticket and that's it? [13:01:24] claime: i'm not sure if there's a way to completely disable them, but I think just setting the repeat interval to a a very long time would do the same in practice? [13:01:49] aaah I forgot about the repeat interval! yes, good call [14:02:20] Heads-up that I've depooled search EQIAD for a maintenance, ref T324335 . No user impact is expected beyond latency, but it will probably cause some alert spam in #operations over the next hour or so [14:02:20] T324335: Index the OpenSearch server application logs in the central observability cluster, using ECS format - https://phabricator.wikimedia.org/T324335 [14:23:14] hi folks - following yesterday's etcd switchover, we'll soon be depooling client read traffic from eqiad, in advance of reimaging the cluster there. no impact expected, just mentioning it for awareness. I'll follow up here again before reimages start. [16:56:53] ^^ The above maintenance is finished and search is enabled in eqiad again [17:11:15] does anyone happen to know anything about zarcillo? [0] it appears to contain a surprise etcd client, which is the only remaining client hitting the eqiad etcd cluster and is blocking the reimages there. [17:11:16] [0] https://wikitech.wikimedia.org/wiki/MariaDB/Zarcillo [17:33:17] swfrench-wmf: i think Federico as the owner of https://phabricator.wikimedia.org/T384810 [17:33:41] mutante: indeed, thank you! [18:13:56] would anyone be comfortable reviewing https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1321611 to (hopefully) fix zarcillo's network policy? what this does is add the v4 and v6 addresses of the codfw etcd cluster members (previously it was only the eqiad values). [18:15:11] swfrench-wmf: happy to, looking [18:15:35] thank you very much [18:16:09] trying to find a solution here so that your and Chris' help this morning is not for nothing :) [18:16:44] one question [18:16:51] (left a comment in the commit) [18:27:36] apologies, got interrupted - responded, and hope that makes sense :) [18:28:04] yes, thanks, mostly for my own learning if nothing else :) +1 [18:28:09] and thanks c.danis as well! [18:28:11] thank you both [18:28:24] indeed [18:33:53] I wonder why that doesn't use external-services [18:36:01] do we have one for etcd? if not, we should add one :) [18:42:50] ahhh we have one for zookeeper I think [18:42:55] but maybe not etcd [18:51:49] ah, that's entirely believable [18:52:59] alright, now that all the clients are sorted (many thanks f.ederico3 for popping in and deploying the policy change), I think we're ready to start actually reimaging :)