[01:41:20] [non-urgent] hello data persistence friends - I wanted to signal-boost the announcement earlier today [0] just in case anyone happened to miss it, particularly because I recall you folks have longer-lived automation that invokes dbctl periodically. [01:41:20] [0] See "Short main-etcd (conftool, spicerack) read-only on Tuesday 4th August (13:30 UTC)" [01:42:41] in any case, no action _explicitly_ required, unless you have anything you'd like to pause during the short (~ 5m) read-only period. thanks! [01:43:51] (oh, and I'll be online from ~ 13:00 and coordinating in -sre and / or -operations) [06:45:33] Based on yesterday's conversation, I will be merging today https://gerrit.wikimedia.org/r/c/operations/puppet/+/1320627 including a cleanup of outdated comments on core sections, unless you object [06:55:05] swfrench-wmf: thanks I sent a calendar invite to the team yesterday with your email so people are aware [06:55:28] In any case federico3 cezmunsta ^ [06:59:29] I'm aware; thanks for the reminder [07:30:51] morning folks, could I get a +1 to https://gerrit.wikimedia.org/r/c/operations/puppet/+/1320658 please? These are the last 3 old-style storage nodes running bullseye, I've drained them from the rings, so they now need to come out of the rings and get converted to new-style storage so I can reimage them to trixie. [07:31:33] [I did reimage a couple of old-style nodes to trixie but it's not a great process and risks losing data, hence deciding to drain-convert-reimage these last three] [07:33:07] this seems strange to me: ms-be(106[6-79]| [07:34:19] jynus: ms-be1068 is still old-style [07:34:38] [I reimaged it to trixie old-style and it was a pain, which is why I decided to convert 1069-71] [07:34:42] I see [07:35:06] Emperor: that is a lot of servers, no? should it be 10[66-79]? [07:35:07] probably once I'm done with upgrades I'll convert the last holdouts to new-style storage [07:35:56] I think it is because I would use [679] [07:35:59] cezmunsta: no, 1072-1079 are already new-style, so 1069-1079 is correct [07:37:02] jynus: if you think 106[679] in regex.yaml is clearer, I can make that change [07:37:27] * Emperor doing so it probably is neater [07:38:06] like, I don't think it is a big deal, but I wouldn't use a range for 2 servers only, unless it is a pattern [07:38:56] {{done}} [07:40:06] Emperor: sorry, yep - I was reading it here... as a clustershell rather than regex :) [07:40:34] while for example, |ms-be20[7]* I think it is clearer like that [07:40:36] gerrit should have the new version now [07:44:03] So it was more of a "yeah, I know I skipped 8 on purpose", and making sure an extra 9 wasn't typed by accident [07:44:11] or an extra 7 [07:47:51] 👍 [12:59:50] I removed db1150, db1171 from orchestrator and zarcillo [13:04:14] 104/104 swift hosts running trixie :-) [13:04:24] sobanski: ^-- FYI [13:05:00] \o/ [13:07:17] nice [13:09:29] that has been quite a lot of focused work :) [14:20:09] Emperor: Hi, how hard it would be to make the promethues exporter expose some metrics (number of objects, total size) per container too? Context: T433964 I can try to make a patch to do it for commons thumb containers if you're okay with that [14:20:09] T433964: Experiment: set a short TTL on thumbnails for newly-uploaded images - https://phabricator.wikimedia.org/T433964 [14:23:21] good question [14:27:59] Amir1: AFAICT the prometheus setup currently only does stats at the per-account level. Shifting that to per-container would be really quite a lot more metrics to produce (and store) [14:29:34] @cezmunsta are you using test-s4? [14:29:47] No [14:30:15] Emperor: I double check with o11y to see if they are okay but also if possible, it'd be only thumbs of commons which is only 256 containers [14:35:30] Amir1: I'm not totally opposed, but I'd like a clear indication of the extra workload on the stats reporter host &c. [14:42:20] Can I get +1s on two more changes, please? https://gerrit.wikimedia.org/r/c/operations/puppet/+/1320950 to load two new backends to the codfw rings, and https://gerrit.wikimedia.org/r/c/operations/puppet/+/1320955 to load the three nodes we've just converted to new-style storage in eqiad and drain the last three old-style nodes so we can convert them to new-style in due course [14:57:35] federico3: just checking if you saw my response ^ ... also, which of the rolling restart scripts is intended for mX? [14:59:19] uhm no, when was that? [14:59:50] immediately after your question about test-s4 [15:00:36] https://usercontent.irccloud-cdn.com/file/yDEm7O8W/image.png [15:00:57] it got losts by my client or maybe serverside? not good [15:02:23] anyhow for ms there's scripts/rolling_restart_pc_ms.py but "m" there isn't [15:02:28] see https://wikitech.wikimedia.org/wiki/MariaDB/Upgrading_a_section [15:28:57] FIRING: SystemdUnitFailed: swift_ring_manager.service on ms-fe2009:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:33:01] RESOLVED: SystemdUnitFailed: swift_ring_manager.service on ms-fe2009:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:55:34] federico3: are you still using test-s4? [15:55:41] a quick review for this would be appreciated https://gitlab.wikimedia.org/repos/sre/schema-changes/-/merge_requests/68 :D [15:55:44] yes, do you need it now? [15:57:46] federico3: no, I was going to try something with db-test2002 re where it can replicate from [15:58:36] I just reset it, go ahead [15:58:56] Amir1: looking [15:59:09] thanks [15:59:18] federico3: by reset, do you mean moved back or a real RESET ...? [15:59:58] moved back [17:13:25] federico3: OK, leaving it alone again now if you need it [17:13:36] ok