[01:21:10] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:06:10] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:06:26] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:53:53] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12358459 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=21ede500-1983-4730-b871-1c7ca225b230) set by ayounsi@cumin1004 for 2:00:00 on 20 host(s) and their services with rea... [08:03:02] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12358473 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=728b2b9e-70ae-4304-acc5-5669decbb5dd) set by ayounsi@cumin1004 for 2:00:00 on 3 host(s) and their services with reas... [08:31:10] FIRING: [4x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:41:10] FIRING: [4x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:45:05] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12358769 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=3ba5e2c1-099e-4219-85d1-7c50f64c9b83) set by ayounsi@cumin1004 for 2:00:00 on 3 host(s) and their services with reas... [08:46:55] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12358779 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=5d53fa31-8467-4252-9f52-e7035b9810c9) set by ayounsi@cumin1004 for 2:00:00 on 19 host(s) and their services with rea... [09:08:01] 10SRE-tools, 10Icinga, 06Infrastructure-Foundations: get-raid-status-perccli should allow for commands to return non-zero exit code - https://phabricator.wikimedia.org/T320998#12358833 (10SLyngshede-WMF) 05Open→03Declined Script no longer exists. [09:16:10] FIRING: [4x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:18:42] 10CAS-SSO, 10Bitu, 10Continuous-Integration-Infrastructure, 06Infrastructure-Foundations: Update basedn in CAS - https://phabricator.wikimedia.org/T371930#12358883 (10LSobanski) a:05SLyngshede-WMF→03None [09:31:10] FIRING: [4x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:32:17] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12358952 (10ayounsi) [12:25:52] 10netbox, 06Infrastructure-Foundations: Netbox: reports alerting issues - https://phabricator.wikimedia.org/T439130 (10ayounsi) 03NEW [12:45:13] 10netbox, 06Infrastructure-Foundations: Upgrade Netbox to 4.7.x - https://phabricator.wikimedia.org/T371889#12359810 (10ayounsi) [12:58:24] 10netops, 06Infrastructure-Foundations: Restart lsw1-a2-codfw for T437984 - https://phabricator.wikimedia.org/T439134 (10ayounsi) 03NEW [12:59:03] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12359905 (10ayounsi) [12:59:04] 10netops, 06Infrastructure-Foundations: Restart lsw1-a2-codfw for T437984 - https://phabricator.wikimedia.org/T439134#12359906 (10ayounsi) [13:24:26] 10netops, 06Infrastructure-Foundations: Juniper hostkey-algorithm deprecation - https://phabricator.wikimedia.org/T439136 (10ayounsi) 03NEW p:05Triage→03High [13:31:26] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:35:23] 10CFSSL-PKI, 06Infrastructure-Foundations, 13Patch-For-Review: Add a load balanced service in front of PKI - https://phabricator.wikimedia.org/T436809#12360083 (10elukey) There seems to be consensus on dropping mTLS and relying on the HMAC signing provided by cfssl. Current status: * I merged https://gerrit... [13:58:44] 10netops, 06Infrastructure-Foundations, 13Patch-For-Review: Juniper hostkey-algorithm deprecation - https://phabricator.wikimedia.org/T439136#12360238 (10ayounsi) ` bast6003:~$ ssh -o HostKeyAlgorithms=ssh-dss -o PubkeyAcceptedAlgorithms=ssh-dss cr1-drmrs.wikimedia.org Unable to negotiate with 2a02:ec80:600:... [14:05:40] 10netops, 06Infrastructure-Foundations, 13Patch-For-Review: Juniper hostkey-algorithm deprecation - https://phabricator.wikimedia.org/T439136#12360291 (10ayounsi) [16:30:11] 10netops, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-09-18 - 2026-10-09), 07Essential-Work, 07Kubernetes: Calico IPv4 block exhaustion on dse-k8s cluster, blocking new node provisioning - https://phabricator.wikimedia.org/T429773#12361009 (10BTullis) 05Open→03Resolved I'm going to re... [17:31:26] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:14:52] 10netops, 06Infrastructure-Foundations: Junos 25 and gNMI metrics - https://phabricator.wikimedia.org/T439166 (10ayounsi) 03NEW p:05Triage→03High [21:31:26] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:37:32] Heya! it seems our puppet has gotten so bloaty and slow that spicerack's default 300 second timeout is too small. I'm calling _spicerack.puppet().run() from SRELBBatchRunnerBase and I keep getting timeouts that cause my cookbooks to fail [22:06:18] ugh, 5mins is quite a while, which hosts brett? [22:06:31] various cp hosts in ulsfo and magru [22:06:49] cp4041 is an example [22:07:22] I've just been re-running the cookbooks until they pass [22:07:55] oh, actually, it seems it was *only* ulsfo, not magru [22:07:59] how curious [22:09:33] hmm, a manual run took 1m31 on cp4041 [22:10:13] how many hosts are you doing at once? [22:10:41] https://bouncer.i--b.com/uploads/brett/c9db73f2f5c42955.txt [22:10:48] One at a time (of text and upload) [22:11:35] maybe it is intermittent, based on some other conditions [22:12:10] perhaps. Since there's a looong delay of 30 minutes in-between this has been running all day [22:12:18] But yeah, I guess I'll just chalk this up to space radiation atm [22:12:33] lol [22:12:52] happy to dig further tomorrow, if its keeps being a nusiance [22:14:39] second run was 2m7, so definitely some variability [22:16:23] no biggie! Almost done with that dc anywa [22:51:47] ulsfo but not magru might point to some slowness or other issues with the codfw puppetservers [22:51:54] (also could just be a red herring)