[05:13:43] I [05:13:57] FYI, I'm rebooting bast1004 shortly [05:14:22] ok [05:14:29] thanks for the heads up [05:17:28] and it's back and can be used again [05:17:58] thank you! [10:01:52] gonna merge a change to add a new kafka tls check, might make some alertspam [13:52:27] I see Debian have dropped their ATS package entirely [13:52:40] yeah that has been for quite a while IIRC [13:53:23] * Emperor got the RM notification to the "this still depends on the obsolete PCRE library" bugs [14:00:21] the maintainer is unavailable for personal reasons, if anyone from WMF SRE wants to step up and help out, I'd be happy to sponsor the uploads to Debian [14:01:50] current upstream releases are fixed for pcre2 [14:09:27] The pcre2 maintainer is mostly friendly too ;) [15:01:10] Please be aware that we're upsizing codfw for the DC switch over: https://phabricator.wikimedia.org/T438299 [15:01:27] And that the DC switch over is happening on Wednesday as planned. [16:02:32] to the on-callers: I've just reimaged and repooled registry2004 to trixie, logs look good and I don't expect anything big. In case of trouble, please just depool it. [17:37:09] <_joe_> ack [18:04:40] _joe_ / federico3: I've been (trying to) put sessionstore1005 back together cf T437915 - I've (hopefully) put built an appropriate fs on the replacement drive, restarted cassandra, nodetool-a status -r seems happy. Do you want to try repooling sessionstore in eqiad, or just leave it depooled since we're depooling eqiad tomorrow anyway? u.random is OOO all week. [18:04:41] T437915: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915 [18:05:55] [it's arguable that we want sessionstore in eqiad available in case of problems with the one in codfw; but it's also possible I've cocked something up in trying to put sessionstore1005 back together because this isn't exactly a highly-documented process, and I've not done it before] [18:07:02] Emperor: how is this impacted by the DC switchover? Also would you want to pool it in now or on tomorrow morning? [18:07:36] federico3: so eqiad is getting depooled tomorrow as part of the switchover. [18:08:23] it'll then be repooled on 1st October [18:09:44] I _think_ that means that absent any other issue, eqiad-sessionstore would be depooled until 1st Oct from whichever point in the switchover flips sessionstorage [18:10:31] in the mean time, having eqiad-sessionstore down is increasing latency for all users [18:10:54] (hence the ticket being UBN, I think) [18:13:15] Given all of which I _think_ it would be sensible to repool eqiad-sessionstore and depool it again if it looks like it's misbehaving. But I wanted oncall's opinions since it would not be me being disturbed if I've fscked it up [18:16:07] maybe I should also annoyingly ping rzl and bblack since they're about to take over as oncallers. [18:17:42] yeah, I haven't been worrying about the cross-dc latency too much specifically because we're about to depool it tomorrow no matter what [18:18:03] but if you think it's ready to serve, no reason not to repool it today -- especially because then we'll know whether we *can* [18:18:55] (whenever we have eqiad depooled for the switchover exercise, the point is that it's an exercise and we can always repool if we discover a problem! so, important to know whether that's true, or whether sessionstore-eqiad is still actually in bad shape) [18:20:42] Right, OK. I'll repool it shortly. Any sign of it being funny we can depool it again and u.random can yell at me when he's back from leave :) [18:21:15] Hang on, just seen the slack ping [18:25:30] To join this conversation up, that was a reference to https://en.wikipedia.org/w/index.php?title=Wikipedia:Village_pump_(technical)#Loss_of_session_data [18:26:02] but those complaints have been going on for a couple of days, so can't be related to what I've been doing this afternoon/evening [18:26:20] So I think going ahead an repooling eqiad-sessionstore is still the way to go [18:28:22] <_joe_> please proceed [18:29:21] <_joe_> I was at dinner [18:30:52] doing so now [18:34:55] [going to take an overdue 10m typing break, then will be back to see how it's going, but initial metrics look OK to my not-expert eye] [18:46:45] looking OK ot me [18:48:12] [thankfully my climbing buddy who I have stood up this evening is understanding] [19:02:49] I think it's OK enough for me to step AFK. [19:03:41] thank you for staying late :( sorry about that