[01:48:23] I've added the bullseye deprecation reminder to the next edition of Tech News: https://meta.wikimedia.org/wiki/Tech/News/2026/34 [09:05:22] gah I forgot to run the clinic duty rotation script yesterday ! [09:49:21] ok fix proposal for the rotate script, https://gitlab.wikimedia.org/repos/cloud/wmcs/utils/-/merge_requests/9 [09:59:11] got an easy one for pywikibot upgrade in paws https://github.com/toolforge/paws/pull/529 [10:04:47] godog: which days are there that are not 'any day between Thursday and the following Wednesday'? [10:06:50] taavi: lol I see what you mean, my intention was to make sure we don't skip a week but as stated in the code it doesn't make sense [10:07:33] so yeah maybe just look back for last thurs if today is not a thurs and be done with it [10:08:04] that'll work until there's a week where Thursday is a holiday and we move the meeting [10:10:11] good point, thank you for taking a look [10:10:51] I'm already regretting opening the can of worms, I'll let the MR be [10:11:52] i.e. abandon [10:12:02] tbh I think just going with a 'find the last Thursday' as you suggested is good enough [10:12:35] it already can't handle the moved meeting case (and forgetting to run the script on Thursday is much more common than that) [10:13:21] fair enough, if the next person that forgets to run on thurs wants to revive the change then I'm all for it [10:18:30] there's the notification script, that did help me last time (it still messes up if the meeting happens on a non-thursday, but well, better than nothing) [10:18:34] * dcaro vanishes [10:21:12] lol thank you d.caro [10:24:48] andrewbogott: hmm, did the fastcci project NFS server get retired instead of being upgraded to trixie? [10:26:32] https://github.com/xoreaxeaxeax/skitter-creek-bath-salts boggles the mind [12:03:52] amazing, thanks for sharing g.odog [12:07:10] you are welcome ! can't say I fully grasp it, but the outcome is clear [12:24:48] what does the red outline that has appeared on the bullseye gsheet mean? [13:39:44] taavi: good question re: fastcci. I seems like that server must have existed when T401812 and T429793 were created but never made the list... maybe it was already shut down even then? [13:39:45] T401812: Migrate WMCS-managed NFS servers off of Bullseye - https://phabricator.wikimedia.org/T401812 [13:39:45] T429793: Update all cloud-vps Bullseye NFS servers - https://phabricator.wikimedia.org/T429793 [13:40:58] If I deleted it in error it will be easy to recreate. If any project admins ever appear... [13:42:47] andrewbogott: hmm. based on the dates of https://phabricator.wikimedia.org/rCLIP8b246920212d87a5ce2969ecf8cff4bd3cbf13da it was deleted last week with the other shutdown NFS servers [13:42:58] so that just leaves the mystery of why it was shut down if there is no replacement [13:44:19] (I'm asking because it's still in the puppet list of projects with NFS enabled) [13:44:27] Yep, that's the same timeline I tracked. But the SAL also has you turning off all VMs in february... let's see if they've been on since then [13:45:06] nah, there's been all kinds of admin work done since Feb [13:45:28] no, wait [13:45:40] that's not right, the things in the action log are just live migrations. [13:46:12] in the purge it was claimed by some user, not a project admin it seems https://wikitech.wikimedia.org/wiki/News/2025_Cloud_VPS_Purge#fastcci_In_use [13:46:59] So my best guess is that everything in that project has been shut off since "turning off VMs per no response to https://wikitech.wikimedia.org/wiki/News/2025_Cloud_VPS_Purge" on Feb 4 [13:47:08] that would make sense [13:47:16] and despite being 'claimed' weeks later no one turned it back on. [13:47:26] So... that leaves us with a couple of options :) [13:47:59] Even if things had been turned back on in Feb they'd be off again /now/ for the bullseye deprecation. [13:48:23] The last time an admin contacted me about the project was in 2021 but I guess I will respond to that email. [13:50:13] andrewbogott: the google spreadsheet of projects ti be migrated says that you deleted something on fastcci on august 5. I guess that just refers to the NFS server, and nothing else? [13:50:28] yeah, that's right. [13:50:49] ok, then that calling the entire project decommissioned seems like wrong. /me updates the sheet to be more accurate [13:51:16] (btw, in case it's not obvious: the nfs /data/ is still present in the project, and recreating an nfs server is pretty trivial. So that's why I'm being casual about the potential incorrect deletion of the nfs server) [13:51:28] (which I double-checked before deleting the server iirc) [13:51:36] yeah, that's true [13:53:49] Only one admin for that project, Daniel Schwen. I will see if I can also ping him on-wiki... [13:55:36] This upholds our timeline: https://commons.wikimedia.org/wiki/User_talk:Dschwen#FastCCI_Server_is_down [14:00:50] https://commons.wikimedia.org/wiki/User_talk:Dschwen#fastcci_project_will_be_deleted_soon [14:01:20] We could also make some kind of public call for project adoption if it's important enough for further effort... [14:02:17] taavi: I'm going to just leave it there unless you or komla feel differently. [14:06:44] hashar: We've been seeing this alert for quite a while "Puppet CA certificate gitlab-runners-puppetmaster-01.gitlab-runners.eqiad1.wikimedia.cloud is about to expire in 10d 7h 20m 16s" -- is that something you or someone on your team is equipped to deal with? Runbook is here: https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/PuppetCertificateAboutToExpire [14:07:25] andrewbogott: yeah my colleagues in the US would be able to handle it [14:07:39] I can file a task and relay that to our team channel [14:07:41] 'colleqgues in the us' meaning dduvall ? [14:07:44] ok, thank you! [14:07:51] * hashar files task [14:08:20] oh hmm https://phabricator.wikimedia.org/T433581 [14:08:33] and jelto on July 31 stated he has rotated the cert [14:09:17] andrewbogott: so tentatively that got fixed 15 days ago [14:10:27] that task links to https://alerts.wikimedia.org/?q=%40state%3Dactive&q=alertname%3DPuppetCertificateAboutToExpire&q=project%3Dgitlab-runners [14:10:48] and if I drop the filters to only keep `alertname=PuppetCertificateAboutToExpire` it shows up ANOTHER one: [14:10:50] summary: Puppet CA certificate Puppet CA: integration-puppetmaster-02.integration.eqiad.wmflabs is about to expire in 22d 11h 36m 42s [14:11:13] I do not know why I'm seeing one and not the other [14:11:33] and that other alert I have no idea where it is alerting [14:12:29] the integration one or the runner one? [14:13:07] well both I guess [14:13:36] I mean I have no idea how the Puppet master CA get monitored but that alert for `integration` should definitely notify releng / me [14:14:05] https://alerts.wikimedia.org/?q=team%3Dwmcs is where I'm seeing the -runner one [14:14:10] and on https://alerts.wikimedia.org/?q=%40state%3Dactive&q=alertname%3DPuppetCertificateAboutToExpire I don't see the gitlab-runners project so .. [14:14:47] OH [14:14:47] https://alerts.wikimedia.org/?q=alertname%3DPuppetCertificateAboutToExpire [14:14:48] so that's a prod alert and not a cloud-vps alert. I'm not sure how/why [14:14:53] here they are both [14:15:15] ok [14:15:29] I re-opened that task, should I also poke people on slack? [14:15:49] I think Jelto is in vacations [14:15:57] so possibly #wikimedia-sre-collab [14:16:48] not -releng? [14:16:57] I think they are managed by SRE [14:17:05] if they're managed by SRE, that's me. [14:17:16] But I would like them to be managed by the project owners instead [14:17:18] 🤯 [14:17:24] so that would be SRE Collab [14:17:44] ok, I will ask around [14:17:53] individual project puppetservers outside of 'our' projects are not managed by sre tools infra, at least [14:18:33] yeah. We provide One True puppetserver, and if you want to use your own instead you are On Your Own (tm) [14:18:51] It's totally possible that those projects don't even need separate servers though. No idea. [14:19:04] oh the other alert about "Puppet CA certificate Puppet CA: integration-puppetmaster-02.integration.eqiad.wmflabs". That host does not exist (not the obsolete puppetmaster terminology and .wmflabs.). So that is a ghost alert [14:19:22] it's not a ghost [14:19:45] why so? [14:20:04] the alert is tagged against the new host (integration-puppetserver-01), the CA is named after the first host it was created at, and has been transferred from host to host ever since [14:20:45] ?!!?!! [14:26:34] Most of the puppet data is on a cinder volume that can move to new servers; that's nice for not having to regenerate every single client cert when you upgrade. [14:31:28] Not After : Jul 21 09:06:41 2029 GMT [14:31:28] Subject: CN = integration-puppetserver-01.integration.eqiad1.wikimedia.cloud [14:31:32] so yeah at least it is valid for a while [14:31:36] with the proper CN [14:31:44] so I have no clue what that other alerts is for