[03:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:41:54] moritzm: I'm going to depool esams to reboot asw1-by27-esams, could you drain the ganeti nodes ganeti3007 and 3005 when you have time? thanks [08:04:25] sure, I'll look into it in a bit [08:17:55] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984 (10ayounsi) 03NEW p:05Triage→03High [08:18:16] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12321134 (10ayounsi) [08:18:37] topranks: new junos mess: https://phabricator.wikimedia.org/T437984 [08:18:55] moritzm: you can use that task for the ganeti draining ^ [08:20:11] ugh [08:20:17] fun :) [08:20:49] seems to only affect a small number of devices right now though?? is it only qfx? [08:21:28] topranks: yeah, looks like a special combination of QFX and code version, unfotunately looks like there is a big risk of spreading to all of codfw A/B [08:22:02] right [08:26:05] ganeti3005 is drained, 3007 next [08:28:04] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12321156 (10ayounsi) Deployed the workaround on asw1-b3-magru (not showing signs of issues yet) ` ayounsi@asw1-b3-magru> start shell user root Password: root@asw1-b3-magru:RE:0% vty fpc0 Swit... [08:29:34] XioNoX: “Identify on which switches the workaround needs to be applied” [08:29:43] what’s the criteria there? [08:30:14] my guess is asw1-b3-magru (I pushed the workaround) and codfw A/B [08:30:52] everything else is either showing the issue, or has been running for a long time and would already have shown signs [08:32:43] hmm ok [08:33:09] it’s non disruptive right? [08:33:27] looks like not :) [08:33:42] only the reboot for the 6 will be [08:39:24] yeah. I guess we should do codfw a/b then [08:39:48] we can divide the work let me know [08:40:23] topranks: it's ok I can take care of it, it's quite straighforward https://phabricator.wikimedia.org/T437984#12321156 [08:40:59] ah give me a few at least, a guy deserves some excitement in his life :P [08:41:52] hahha sounds good, pick the ones you want and add them to the task description once done [08:42:36] ok cool, I'll tackle a few of the codfw row a ones and report back [08:47:06] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12321264 (10ayounsi) [08:55:38] XioNoX: you can go ahead, install3004 is still on ganeti3007, but we can safely ignore it, I've downtimed it for an hour [08:55:58] will the other esams switch need the same? [08:56:41] moritzm: thanks, yeah, I first I thought only one needed to be rebooted, but when writing the task I noticed that the other one too... [08:58:06] sure thing, just drop me a note when the first one is done, I need to do https://phabricator.wikimedia.org/T428878 for esams anyway so I'll do that whent he first switch is rebooted and before draining the second [08:58:11] moritzm: also unrelated, but there is a race condition between the daily homer runs on cumin1003 and 1004, as they runs at the same time, they sometime hit the locks set by the other one [08:58:30] nice [08:58:33] depooling esams now [08:59:34] ah, right. I'll make a patch to run it two hours later [08:59:58] moritzm: or just on one of them is fine [09:01:40] that would mean to disable Homer entirely on cumin1003, which is fine with me [09:02:01] ^ topranks: did you also move to cumin1004 entirely for Homer in eqiad? [09:04:37] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12321357 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=a4925aa6-63a3-4037-92db-c71b5a881bd3) set by ayounsi@cumin1003 for 2:00:00 on 12 host(s) and their services with rea... [09:06:12] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12321375 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=c96fe5ee-ddcb-4a50-935b-5314cd5dd078) set by ayounsi@cumin1003 for 2:00:00 on 3 host(s) and their services with reas... [09:15:32] moritzm: yes I've been using homer from cumin1004 since we did the deploy without any issues [09:19:58] or just the timer :D [09:20:03] (disable) [09:21:35] moritzm: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1341710 only disable the daily "scan" timer, it's still possible to use homer normally [09:23:13] what does "pofile::homer::disable:" control? [09:23:14] IMHO rule of thumb for the active homer hosts should be: those that sync the private repo between them [09:23:27] XioNoX: I'd just nuke the homer binary after merging, then noone can use it [09:23:30] as long as they sync it it can run on any of them [09:23:33] volans: yeah that makes sense [09:23:58] my bad, I missread [09:24:05] didn't see line 15 of https://github.com/wikimedia/operations-puppet/blob/production/modules/profile/manifests/homer.pp#L15 [09:24:11] only line 36 [09:24:12] topranks: the homer deploy/keyholder etc. are no longer deployed [09:25:01] moritzm: right thanks, being lazy sry, I can see from Arzhel's link [09:25:32] perfect, I'll merge in a few and nuke the homer script, so that everyone only uses 1004 [09:26:06] +1 [09:40:36] moritzm: you can un-drain the by27 ganeti node [09:42:46] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12321554 (10ayounsi) [09:51:24] XioNoX: are you still planning to reboot the second switch, despite the load issues? [09:51:52] moritzm: not today, it's not great [09:52:04] ok! in the mean time I'll proceed with https://phabricator.wikimedia.org/T428878 for esams [09:52:45] moritzm: I can do asw1-b4-magru though [09:54:28] are you still aiming to do the second switch for this week? then I would fail over the ganeti master in esams away from ganeti3008 [09:55:12] moritzm: probably tomorrow at a time when there is less load on EU DCs [09:55:27] ok [09:55:50] I'll drain ganeti7002/7004 in a bit [10:05:08] Hi! I wanted to ask a question about docker-registry, and how the tags are garbage-collected. I started in #wikimedia-releng and some context is in https://phabricator.wikimedia.org/T437994 [10:05:55] We publish 2 versions of the same software in one docker image https://docker-registry.wikimedia.org/repos/data-engineering/airflow-dags/tags/ [10:06:44] Old version, `airflow-2*` is running in prod, and is tagged with version tag and with `latest` tag on build, but rarely built (shame on us, we better have CD) [10:07:08] New version is built frequently but published only with version tag [10:08:11] I'm afraid that if I continue building new version, we can accidentally evict old verison from the registry. Is it correct assumption? [10:16:36] btullis advised to ping elukey, did so in the task [10:20:18] atsukoito: o/ [10:21:00] so at the moment we don't garbage collect old tags, because before the S3 migration the docker registry's internal gc code didn't work [10:21:33] on swift we have ~6TB of data, and now on s3 we have around 2.5TB (because during the recent switch I didn't copy the whole history) [10:23:01] on Wikitech we have https://wikitech.wikimedia.org/wiki/Docker-registry#Deleting_images, but the command just drops the manifest from the registry, the binaries stay there. The idea is that the GC algorithm of the registry will do a mark and sweep, and collect dangling/not-referenced docker layers and delete them [10:23:30] but this needs to be done when the registry is read only, otherwise we'd risk to have newly pushed layers to be marked for deletion [10:23:47] There is a task to migrate to Harbor or similar but we are very far from it [10:24:52] so TL;DR - at the moment we don't drop/clean-up, we may do it in the near-ish future but for things explicitly deleted via docker-registryctl etc.. [10:24:58] thanks for the context! i'll be rebuilding the image with new tag until I fix the CI configuration, and until the snapshots is sorted T437829 [10:24:59] T437829: scap backport failed due to expired debian snapshot - https://phabricator.wikimedia.org/T437829 [10:25:30] I plan to fix my CI pipeline to use separate images this week [10:26:08] what is your end goal though? Push new versions of the image overriding latest? [10:27:19] My end-goal is to leverage `latest` for both airflow 2 and airflow 3 because that's how our test logic is done [10:28:04] But I didn't know initially that kokkuri is able to push to a different images, and overabused tags for this [10:29:21] So I'll be creating something like `d-r.w.o/repos/data-engineering/airflow-dags` for old image and `d-r.w.o/repos/data-engineering/airflow-dags/airflow3` for new image [10:29:56] okok! [11:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:34:56] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12322288 (10cmooney) FWIW I tried this on //lsw1-a2-codfw// but my session just hangs when I try to enter the FPC vty session: ` cmooney@lsw1-a2-codfw> start shell user root Password: root@lsw... [11:38:37] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12322320 (10cmooney) My bad on the above - I'd missed lsw1-a2-codfw was one of the affected devices, so the above shouldn't have been a surprise. [12:09:28] XioNoX: ganeti7002/ganeti7004 are drained, you can reboot he magru switch [12:13:16] thx [12:38:53] moritzm: maintenance over, you can put them back in service, thanks! [12:51:58] do we also need to reboot the other magru switch? [13:03:05] moritzm: no I don't think so, Arzhel applied the workaround to stop the logging on it so we are hopefully ok [13:04:03] perfect, I'll go rebalance magru,then [14:06:41] 10netops, 06Infrastructure-Foundations, 06SRE: Add link from cloudsw1-e4-eqiad to cloudsw1-f4-eiqad - https://phabricator.wikimedia.org/T372061#12323044 (10cmooney) 05Open→03Declined Closing this one, we will review how to provision the bandwidth when the C8/D5 switches are upgraded. [14:29:55] Hello. I have a couple of dinky VM requests in T438037 and T438038. I'm obviously happy to make them, but thought it best to run it past you for any suggestions as to location etc. Ta. [14:29:55] T438037: eqiad: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438037 [14:29:56] T438038: codfw: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438038 [15:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:22:34] btullis: sounds good,left a note for both on task [15:29:22] thx [15:36:36] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: EQSIN:Switch refresh diagram and wiring - https://phabricator.wikimedia.org/T423724#12323689 (10RobH) 05Open→03Resolved [15:47:38] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052 (10RobH) 03NEW p:05Triage→03Medium [15:53:39] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12323874 (10RobH) [16:58:41] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Eqiad: replace unmanaged msw in racks A4, B7 & C5 - https://phabricator.wikimedia.org/T438071 (10cmooney) 03NEW p:05Triage→03Low [17:53:26] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12324568 (10RobH) [17:54:39] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12324576 (10RobH) [18:45:55] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12324794 (10ssingh) [19:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:26:59] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12324913 (10RobH) [23:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed