[00:25:56] (03PS1) 10Ahmon Dancy: README: Update plugin testing instructions [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327666 (https://phabricator.wikimedia.org/T434726) [00:32:01] (03PS1) 10Ahmon Dancy: plugins-router: set the MIME type of a served plugin [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327668 (https://phabricator.wikimedia.org/T434726) [00:33:08] (03CR) 10Ahmon Dancy: [C:03+2] README: Update plugin testing instructions [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327666 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [00:33:50] (03Merged) 10jenkins-bot: README: Update plugin testing instructions [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327666 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [00:35:14] (03PS1) 10Ahmon Dancy: plugins: Add wm-wheres-my-code-running plugin [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327669 (https://phabricator.wikimedia.org/T434726) [00:40:25] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:52:35] FIRING: DiskSpace: Disk space build2001:9100:/ 2.885% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=build2001 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [00:54:57] (03PS1) 10Krinkle: Page: Clean up tests of PageProps service class [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327671 (https://phabricator.wikimedia.org/T297300) [00:58:37] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [01:08:57] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations: install1005 running out of disk due to squid log volume from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12239992 (10Scott_French) [01:11:25] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1327672 [01:11:25] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1327672 (owner: 10TrainBranchBot) [01:19:07] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1327672 (owner: 10TrainBranchBot) [01:22:40] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:24:53] (03PS2) 10Lerickson: Update the Qlever image to include named subqueries. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327642 (https://phabricator.wikimedia.org/T435122) [01:46:02] (03CR) 10Lerickson: "Note: This actually isn't a good change as stated because I need to cut yet another proxy to include this fix: https://gitlab.wikimedia.or" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327630 (https://phabricator.wikimedia.org/T433935) (owner: 10Lerickson) [02:00:39] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:08:21] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) [02:13:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:00:37] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [03:42:35] RESOLVED: DiskSpace: Disk space build2001:9100:/ 2.934% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=build2001 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [04:15:33] (03PS2) 10Ryan Kemper: wdqs: drop dangling query-legacy-full cert SAN [deployment-charts] - 10https://gerrit.wikimedia.org/r/1278562 (https://phabricator.wikimedia.org/T415073) [04:15:55] (03PS3) 10Ryan Kemper: cumin: add wdqs-public and wdqs-internal aliases [puppet] - 10https://gerrit.wikimedia.org/r/1278603 (https://phabricator.wikimedia.org/T415073) [04:16:23] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1278603 (https://phabricator.wikimedia.org/T415073) (owner: 10Ryan Kemper) [04:19:56] (03PS1) 10Krinkle: API: wfDebugLog for thumberror [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327678 [04:20:05] !log arlolra@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [04:20:44] !log arlolra@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [04:20:45] !log arlolra@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [04:21:28] !log arlolra@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [04:23:00] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324963 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [04:23:00] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324964 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [04:23:10] (03CR) 10Ryan Kemper: [C:03+2] cumin: add wdqs-public and wdqs-internal aliases [puppet] - 10https://gerrit.wikimedia.org/r/1278603 (https://phabricator.wikimedia.org/T415073) (owner: 10Ryan Kemper) [04:24:01] (03Merged) 10jenkins-bot: Add Produnto to extension-list [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324963 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [04:24:05] (03Merged) 10jenkins-bot: Enable Produnto on Beta [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324964 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [04:24:36] !log tstarling@deploy1003 Started scap sync-world: Backport for [[gerrit:1324963|Add Produnto to extension-list (T421436)]], [[gerrit:1324964|Enable Produnto on Beta (T421436)]] [04:24:41] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [04:27:02] (03PS1) 10Krinkle: Block: Disable flaky API test [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327679 (https://phabricator.wikimedia.org/T435272) [04:27:16] (03PS2) 10Krinkle: API: wfDebugLog for thumberror [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327678 [04:40:25] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:41:47] (03CR) 10CI reject: [V:04-1] Block: Disable flaky API test [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327679 (https://phabricator.wikimedia.org/T435272) (owner: 10Krinkle) [04:44:28] !log tstarling@deploy1003 tstarling: Backport for [[gerrit:1324963|Add Produnto to extension-list (T421436)]], [[gerrit:1324964|Enable Produnto on Beta (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [04:44:33] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [04:45:35] !log tstarling@deploy1003 tstarling: Continuing with deployment [04:54:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.26% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [04:58:37] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [04:58:56] the alert is apparently a consequence of the restart of all workers, with APCU being cleared [04:59:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [04:59:24] !log tstarling@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324963|Add Produnto to extension-list (T421436)]], [[gerrit:1324964|Enable Produnto on Beta (T421436)]] (duration: 34m 48s) [04:59:29] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [05:04:50] PROBLEM - Backup freshness on backup1014 is CRITICAL: Stale: 1 (krb1002), Fresh: 140 jobs https://wikitech.wikimedia.org/wiki/Bacula%23Monitoring [05:22:40] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:28:24] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations: install1005 running out of disk due to squid log volume from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12240206 (10RKemper) >>! In T435555#12239691, @Scott_French wrote: > Dropping to Medium, as the trigger workload has... [05:29:43] (03CR) 10Krinkle: "recheck" [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327679 (https://phabricator.wikimedia.org/T435272) (owner: 10Krinkle) [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260821T0600) [06:10:22] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [06:10:22] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [06:10:22] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [06:12:22] (03PS1) 10Marostegui: db1283: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1327846 (https://phabricator.wikimedia.org/T407942) [06:13:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:13:54] (03CR) 10Marostegui: [C:03+2] db1283: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1327846 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [06:15:04] (03PS1) 10Marostegui: instances.yaml: Add db1283 [puppet] - 10https://gerrit.wikimedia.org/r/1327847 (https://phabricator.wikimedia.org/T407942) [06:16:08] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1283 [puppet] - 10https://gerrit.wikimedia.org/r/1327847 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [06:17:44] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1283 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96213 and previous config saved to /var/cache/conftool/dbconfig/20260821-061743-marostegui.json [06:17:49] T407942: Productionize db12[65-90] - https://phabricator.wikimedia.org/T407942 [06:18:03] (03CR) 10Filippo Giunchedi: [C:03+1] openstack: novastats: Update proxyleaks for singular backend object [puppet] - 10https://gerrit.wikimedia.org/r/1308063 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [06:18:05] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1283: Pool back [06:18:14] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [06:18:14] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 1.845 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [06:18:14] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [06:18:29] (03PS1) 10Marostegui: db1283: Remove note [puppet] - 10https://gerrit.wikimedia.org/r/1327848 [06:18:39] (03CR) 10Filippo Giunchedi: [C:03+1] openstack: wmf_sink: Update for singular proxy 'backend' object [puppet] - 10https://gerrit.wikimedia.org/r/1308064 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [06:20:18] (03CR) 10Marostegui: [C:03+2] db1283: Remove note [puppet] - 10https://gerrit.wikimedia.org/r/1327848 (owner: 10Marostegui) [06:20:30] RECOVERY - OSPF status on cr1-magru is OK: OSPFv2: 5/5 UP : OSPFv3: 5/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:21:14] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:23:22] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [06:23:22] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [06:23:22] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [06:25:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:32:09] (03PS1) 10Marostegui: installserver: Do not reimage db1283 [puppet] - 10https://gerrit.wikimedia.org/r/1327849 [06:35:08] !log powercycling netmon2002 [06:35:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:36:31] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:37:15] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:38:10] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:38:11] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [06:38:11] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [06:38:11] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.137 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [06:40:28] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [06:42:40] !log jmm@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org [06:45:46] jmm@cumin2003 upgrade-firmware (PID 3557650) is awaiting input [06:48:13] netmon2002 was a hardware-related Linux deadlock/hang. powercycled it and will update firmware on it [06:50:28] RESOLVED: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [07:00:05] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260821T0700) [07:00:11] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1327643 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:00:37] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [07:03:09] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1283: Pool back [07:06:58] 06SRE, 06Infrastructure-Foundations, 06ServiceOps: Migrate Docker reporting from build2002 to build2004 - https://phabricator.wikimedia.org/T435314#12240276 (10MLechvien-WMF) @MoritzMuehlenhoff I believe those hosts are on I/F side so moving this to Serviceops radar, please tag Serviceops or I if there's a s... [07:13:18] !log jayme@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [07:14:10] !log jayme@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [07:14:18] !log jayme@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [07:14:58] !log jayme@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [07:15:07] !log jayme@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [07:18:24] !log jayme@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [07:18:32] !log jayme@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [07:24:53] !log jayme@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [07:25:21] !log jayme@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [07:25:23] !log jayme@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [07:25:32] !log jayme@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [07:25:34] !log jayme@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [07:25:43] !log jayme@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [07:25:45] !log jayme@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [07:25:54] !log jayme@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [07:25:56] !log jayme@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [07:27:22] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327679 (https://phabricator.wikimedia.org/T435272) (owner: 10Krinkle) [07:27:23] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327678 (owner: 10Krinkle) [07:32:24] (03Merged) 10jenkins-bot: Block: Disable flaky API test [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327679 (https://phabricator.wikimedia.org/T435272) (owner: 10Krinkle) [07:32:33] (03Merged) 10jenkins-bot: API: wfDebugLog for thumberror [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327678 (owner: 10Krinkle) [07:33:02] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1327679|Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678|API: wfDebugLog for thumberror]] [07:33:10] T435272: Flaky api-testing/action/Block.js: "should not allow multiblocks without newblock" - https://phabricator.wikimedia.org/T435272 [07:33:11] T389028: Race condition in DatabaseBlockStore::acquireTarget() - https://phabricator.wikimedia.org/T389028 [07:37:17] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1327679|Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678|API: wfDebugLog for thumberror]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:38:13] (03PS1) 10JMeybohm: staging-codfw: Update to coredns 1.12 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327852 (https://phabricator.wikimedia.org/T428573) [07:41:25] !log krinkle@deploy1003 krinkle: Continuing with deployment [07:44:07] (03CR) 10Jelto: [C:03+2] "lgtm thank you!" [puppet] - 10https://gerrit.wikimedia.org/r/1327541 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [07:48:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.33% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:48:36] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1327679|Block: Disable flaky API test (T435272 T389028)]], [[gerrit:1327678|API: wfDebugLog for thumberror]] (duration: 15m 34s) [07:48:43] T435272: Flaky api-testing/action/Block.js: "should not allow multiblocks without newblock" - https://phabricator.wikimedia.org/T435272 [07:48:43] T389028: Race condition in DatabaseBlockStore::acquireTarget() - https://phabricator.wikimedia.org/T389028 [07:51:49] (03CR) 10Jelto: [C:03+2] planet: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1327529 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [07:55:36] (03PS1) 10Filippo Giunchedi: labstore: clean up traffic_shaping [puppet] - 10https://gerrit.wikimedia.org/r/1327853 (https://phabricator.wikimedia.org/T435581) [07:55:38] (03PS1) 10Filippo Giunchedi: wmcs: remove support for clouddumps client symlinks [puppet] - 10https://gerrit.wikimedia.org/r/1327854 (https://phabricator.wikimedia.org/T435581) [07:55:41] (03PS1) 10Filippo Giunchedi: dumps: remove production support for non-lb mounts [puppet] - 10https://gerrit.wikimedia.org/r/1327855 (https://phabricator.wikimedia.org/T435581) [07:55:46] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Enable CPU performance governor on Relforge, Cloudelastic, and Elasticsearch hosts - https://phabricator.wikimedia.org/T386860#12240330 (10Gehel) [08:00:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1046:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1046 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [08:02:41] (03PS3) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [08:03:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:04:16] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 24 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326848 (https://phabricator.wikimedia.org/T433713) (owner: 10Sadiya.mohammed13) [08:05:02] (03CR) 10CI reject: [V:04-1] helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [08:05:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1046:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1046 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [08:08:57] (03PS4) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [08:10:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1046:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1046 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [08:19:39] PROBLEM - Host ml-serve1015 is DOWN: PING CRITICAL - Packet loss = 100% [08:24:50] FIRING: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [08:30:00] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [08:31:48] (03CR) 10Muehlenhoff: [C:03+2] Failover irc.wikimedia.org to irc1003 [dns] - 10https://gerrit.wikimedia.org/r/1327561 (owner: 10Muehlenhoff) [08:31:56] !log jmm@dns1004 START - running authdns-update [08:33:39] !log klausman@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw [08:33:43] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet [08:34:11] !log jmm@dns1004 END - running authdns-update [08:34:58] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 06Traffic, 13Patch-For-Review: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12240442 (10MoritzMuehlenhoff) There's two still things, that need to happen, let me reopen the task until they... [08:35:36] jmm@cumin2003 upgrade-firmware (PID 3557650) is awaiting input [08:35:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1046:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1046 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [08:36:28] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org [08:38:12] (03PS2) 10Arnaudb: puppetserver: pull puppet via discovery record [puppet] - 10https://gerrit.wikimedia.org/r/1296495 (https://phabricator.wikimedia.org/T420184) [08:40:59] FIRING: KafkaMirrorMakerConsumerMaxLag: Kafka MirrorMaker main-eqiad-to-main-codfw max lag in last 10 minutes - https://wikitech.wikimedia.org/wiki/Kafka/Administration#MirrorMaker - https://grafana.wikimedia.org/d/000000521/kafka-mirrormaker?var-mirror_name=main-eqiad-to-main-codfw - https://alerts.wikimedia.org/?q=alertname%3DKafkaMirrorMakerConsumerMaxLag [08:43:50] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet [08:44:49] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org [08:44:51] !log jmm@cumin2003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org [08:45:05] RECOVERY - Host ml-serve1015 is UP: PING OK - Packet loss = 0%, RTA = 0.29 ms [08:45:28] FIRING: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [08:47:41] !log jmm@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts netmon2002.wikimedia.org [08:48:02] !log jmm@cumin2003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts netmon2002.wikimedia.org [08:49:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1046:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1046 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [08:49:40] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet [08:49:42] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet [08:49:48] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet [08:49:50] RESOLVED: KubernetesCalicoDown: ml-serve1015.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1015.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [08:50:00] RESOLVED: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node ml-serve1015 has a BGP session which is not in the 'established' state. [08:54:55] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet [08:55:28] RESOLVED: KeyholderUnarmed: 1 unarmed Keyholder key(s) on netmon2002:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [08:56:49] (03PS1) 10Atsuko: deployment_server: define airflow config [puppet] - 10https://gerrit.wikimedia.org/r/1327550 (https://phabricator.wikimedia.org/T416709) [08:57:02] (03PS1) 10Atsuko: dse-k8s-eqiad: provision new airflow namespace [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327551 (https://phabricator.wikimedia.org/T416709) [08:57:12] (03PS1) 10Atsuko: dse-k8s-eqiad: add the airflow-experiment-platform ns to the ceph tenant list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327552 (https://phabricator.wikimedia.org/T416709) [08:57:30] (03PS1) 10Atsuko: dse-k8s-eqiad: provision the airflow instance [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) [08:58:37] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:00:18] (03PS3) 10Arnaudb: puppetserver: pull puppet via discovery record [puppet] - 10https://gerrit.wikimedia.org/r/1296495 (https://phabricator.wikimedia.org/T420184) [09:00:26] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Packet loss on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12240494 (10cmooney) [09:00:29] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet [09:00:30] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet [09:00:36] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet [09:00:42] (03PS3) 10Jelto: etherpad: add helm chart for etherpad service [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327571 (https://phabricator.wikimedia.org/T435509) [09:01:06] (03PS5) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [09:01:33] i'll be rebooting redis primaries today, please let me know if there seem to be problems [09:02:50] (03CR) 10Muehlenhoff: [C:03+2] Switch the new URL downloaders to insetup_ferm for the reimage [puppet] - 10https://gerrit.wikimedia.org/r/1327491 (https://phabricator.wikimedia.org/T427282) (owner: 10Muehlenhoff) [09:03:18] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Packet loss on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12240509 (10cmooney) The HE NOC came back to advise they had several sub sea fibre breaks at once which affected capacity to the region. ` On Thu, 20 Aug 2026 a... [09:03:19] !log blake@cumin1003 START - Cookbook sre.hosts.reboot-single for host rdb1013.eqiad.wmnet [09:03:31] (03PS4) 10Arnaudb: puppetserver: pull puppet via discovery record [puppet] - 10https://gerrit.wikimedia.org/r/1296495 (https://phabricator.wikimedia.org/T420184) [09:04:34] (03CR) 10Klausman: [C:03+1] dse-k8s-eqiad: provision the airflow instance [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:04:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1046:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1046 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [09:04:49] (03CR) 10Klausman: [C:03+1] dse-k8s-eqiad: add the airflow-experiment-platform ns to the ceph tenant list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327552 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:05:07] (03CR) 10Klausman: [C:03+1] dse-k8s-eqiad: provision new airflow namespace [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327551 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:05:30] (03CR) 10Klausman: [C:03+1] deployment_server: define airflow config [puppet] - 10https://gerrit.wikimedia.org/r/1327550 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:05:51] (03PS5) 10Arnaudb: puppetserver: pull puppet via discovery record [puppet] - 10https://gerrit.wikimedia.org/r/1296495 (https://phabricator.wikimedia.org/T420184) [09:07:12] !log jmm@cumin2003 START - Cookbook sre.hosts.reimage for host urldownloader1005.wikimedia.org with OS trixie [09:07:24] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12240515 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin2003 for host urldownloader1005.wikimedia.org with OS trixie [09:09:17] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1013.eqiad.wmnet [09:09:31] FIRING: RedisReplicaDown: Redis replica down rdb1014:16379 redis_misc - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_misc - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=eqiad&var-job=redis_misc&var-instance=rdb1014:16379 - https://alerts.wikimedia.org/?q=alertname%3DRedisReplicaDown [09:10:02] (03CR) 10JMeybohm: [C:03+2] staging-codfw: Update to coredns 1.12 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327852 (https://phabricator.wikimedia.org/T428573) (owner: 10JMeybohm) [09:10:45] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet [09:11:15] !log blake@cumin1003 START - Cookbook sre.hosts.reboot-single for host rdb1015.eqiad.wmnet [09:11:25] !log jayme@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [09:11:45] !log jayme@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [09:12:37] !log jayme@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [09:13:36] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr[1-2]-eqiad,pfw1-eqiad with reason: upgrade pfw1a-eqiad and pfw1b-eqiad pair [09:14:31] RESOLVED: RedisReplicaDown: Redis replica down rdb1014:16379 redis_misc - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_misc - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=eqiad&var-job=redis_misc&var-instance=rdb1014:16379 - https://alerts.wikimedia.org/?q=alertname%3DRedisReplicaDown [09:15:59] RESOLVED: KafkaMirrorMakerConsumerMaxLag: Kafka MirrorMaker main-eqiad-to-main-codfw max lag in last 10 minutes - https://wikitech.wikimedia.org/wiki/Kafka/Administration#MirrorMaker - https://grafana.wikimedia.org/d/000000521/kafka-mirrormaker?var-mirror_name=main-eqiad-to-main-codfw - https://alerts.wikimedia.org/?q=alertname%3DKafkaMirrorMakerConsumerMaxLag [09:16:01] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet [09:16:02] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet [09:16:08] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet [09:16:32] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1015.eqiad.wmnet [09:17:25] FIRING: [4x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:17:55] (03PS6) 10Arnaudb: puppetserver: pull puppet via discovery record [puppet] - 10https://gerrit.wikimedia.org/r/1296495 (https://phabricator.wikimedia.org/T420184) [09:18:51] FIRING: [4x] SwitchCoreInterfaceDown: Switch core interface down - fasw2-e15a-eqiad:et-0/0/47 (Core: pfw1-eqiad:et-0/1/1) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [09:19:11] FIRING: PfwCoreBGPDown: Fundraising Firewall core BGP session down between pfw1-codfw and (null) (10.195.0.248) - group VPN - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=pfw1-codfw:9804&var-bgp_group=VPN&var-bgp_neighbor=(null) - https://alerts.wikimedia.org/?q=alertname%3DPfwCoreBGPDown [09:21:10] (03PS1) 10Hnowlan: graphite: disable uwsgi service for graphite web [puppet] - 10https://gerrit.wikimedia.org/r/1328153 (https://phabricator.wikimedia.org/T435341) [09:21:13] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet [09:21:35] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Packet loss on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12240542 (10cmooney) [09:22:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:22:31] !log jmm@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage [09:23:43] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Packet loss on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12240544 (10cmooney) [09:24:40] !log blake@cumin1003 START - Cookbook sre.hosts.reboot-single for host rdb2011.codfw.wmnet [09:25:58] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12240546 (10MoritzMuehlenhoff) [09:26:22] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet [09:26:23] (03PS2) 10Arnaudb: pontoon: exception handle in pontoonctl destroy-hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328150 (https://phabricator.wikimedia.org/T420184) [09:26:23] (03CR) 10Arnaudb: "something I stumbled upon working with pontoon 😊" [puppet] - 10https://gerrit.wikimedia.org/r/1328150 (https://phabricator.wikimedia.org/T420184) (owner: 10Arnaudb) [09:26:24] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet [09:26:30] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet [09:28:51] RESOLVED: [4x] SwitchCoreInterfaceDown: Switch core interface down - fasw2-e15a-eqiad:et-0/0/47 (Core: pfw1-eqiad:et-0/1/1) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [09:29:11] RESOLVED: PfwCoreBGPDown: Fundraising Firewall core BGP session down between pfw1-codfw and (null) (10.195.0.248) - group VPN - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=pfw1-codfw:9804&var-bgp_group=VPN&var-bgp_neighbor=(null) - https://alerts.wikimedia.org/?q=alertname%3DPfwCoreBGPDown [09:29:45] (03CR) 10Btullis: [C:03+1] deployment_server: define airflow config [puppet] - 10https://gerrit.wikimedia.org/r/1327550 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:29:49] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1005.wikimedia.org with reason: host reimage [09:30:06] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2011.codfw.wmnet [09:30:12] (03CR) 10Btullis: [C:03+1] dse-k8s-eqiad: provision new airflow namespace [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327551 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:30:33] (03CR) 10Atsuko: [C:03+2] deployment_server: define airflow config [puppet] - 10https://gerrit.wikimedia.org/r/1327550 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:31:07] (03CR) 10Btullis: [C:03+1] dse-k8s-eqiad: add the airflow-experiment-platform ns to the ceph tenant list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327552 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:32:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:32:40] (03CR) 10Btullis: [C:03+1] dse-k8s-eqiad: provision the airflow instance [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:32:55] !log blake@cumin1003 START - Cookbook sre.hosts.reboot-single for host rdb2013.codfw.wmnet [09:35:13] !log fnegri@cumin1003 START - Cookbook sre.wikireplicas.add-wiki for database minwikiquote (T429946) [09:35:18] T429946: [wikireplicas] Create views for new wiki minwikiquote - https://phabricator.wikimedia.org/T429946 [09:35:26] !log fnegri@cumin1003 END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database minwikiquote (T429946) [09:35:50] !log fnegri@cumin1003 START - Cookbook sre.wikireplicas.add-wiki for database bolwiki (T429954) [09:35:55] T429954: [wikireplicas] Create views for new wiki bolwiki - https://phabricator.wikimedia.org/T429954 [09:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:36:38] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet [09:38:22] !log blake@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2013.codfw.wmnet [09:38:36] 06SRE, 06Infrastructure-Foundations, 10netops: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506#12240576 (10cmooney) 05Open→03Resolved Closing, cr2-eqiad still seems happy. Though as discussed on irc with @robh it is somewhat worrying these both alarmed at the sam... [09:41:26] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet [09:41:28] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet [09:41:34] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet [09:45:02] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1005.wikimedia.org with OS trixie [09:45:14] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12240588 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin2003 for host urldownloader1005.wikimedia.org with OS trixie completed: - urldownloader1005 (**PASS**) - Dow... [09:46:03] !log jmm@cumin2003 START - Cookbook sre.hosts.reimage for host urldownloader1006.wikimedia.org with OS trixie [09:46:11] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12240603 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin2003 for host urldownloader1006.wikimedia.org with OS trixie [09:46:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:46:41] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet [09:47:06] 10ops-eqiad, 06SRE, 06DC-Ops: Recycle old MPC 3D 16x10G line cards in cr2-eqiad - https://phabricator.wikimedia.org/T435588 (10cmooney) 03NEW p:05Triage→03Low [09:51:32] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet [09:51:34] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet [09:51:39] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet [09:55:26] (03PS6) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [09:56:47] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet [09:57:34] !log jmm@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage [10:01:35] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet [10:01:37] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet [10:01:42] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet [10:03:39] (03CR) 10Filippo Giunchedi: [C:03+1] "Very nice, thank you !" [puppet] - 10https://gerrit.wikimedia.org/r/1328150 (https://phabricator.wikimedia.org/T420184) (owner: 10Arnaudb) [10:04:37] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1006.wikimedia.org with reason: host reimage [10:05:36] 06SRE, 06Infrastructure-Foundations, 10netops: Alert on packet loss to core sites from POPs - https://phabricator.wikimedia.org/T435592 (10cmooney) 03NEW p:05Triage→03Medium [10:06:49] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet [10:11:11] 06SRE, 06Infrastructure-Foundations, 10netops: Alert on packet loss to core sites from POPs - https://phabricator.wikimedia.org/T435592#12240731 (10cmooney) [10:12:33] 06SRE, 06Infrastructure-Foundations, 10netops: Alert on packet loss to core sites from POPs - https://phabricator.wikimedia.org/T435592#12240734 (10cmooney) [10:13:18] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet [10:13:20] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet [10:13:26] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet [10:13:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:18:33] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet [10:20:32] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1006.wikimedia.org with OS trixie [10:20:42] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12240739 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin2003 for host urldownloader1006.wikimedia.org with OS trixie completed: - urldownloader1006 (**PASS**) - Dow... [10:23:04] (03CR) 10Gmodena: [C:03+1] Update the Qlever image to include named subqueries. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327642 (https://phabricator.wikimedia.org/T435122) (owner: 10Lerickson) [10:24:54] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet [10:24:55] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet [10:25:01] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet [10:28:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:30:10] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet [10:33:00] !log jmm@cumin2003 START - Cookbook sre.hosts.reimage for host urldownloader2005.wikimedia.org with OS trixie [10:33:08] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12240761 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin2003 for host urldownloader2005.wikimedia.org with OS trixie [10:33:31] !log fnegri@cumin1003 END (PASS) - Cookbook sre.wikireplicas.add-wiki (exit_code=0) for database bolwiki (T429954) [10:33:36] T429954: [wikireplicas] Create views for new wiki bolwiki - https://phabricator.wikimedia.org/T429954 [10:34:32] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12240763 (10Marostegui) @VRiley-WMF Jaime is out of office for a few more weeks. What does reprovision implies here? [10:34:51] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet [10:34:53] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet [10:34:58] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet [10:35:38] (03PS10) 10Hashar: zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) [10:36:27] (03CR) 10CI reject: [V:04-1] zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [10:36:36] (03CR) 10Hashar: "This causes the log4j classes to be removed from the CLASSPATH when TLS is enabled. T435503 and fixed by Iffb0b9eda1b28db0f869aeb525fdd117" [puppet] - 10https://gerrit.wikimedia.org/r/1244927 (https://phabricator.wikimedia.org/T395938) (owner: 10Dzahn) [10:38:59] (03PS11) 10Hashar: zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) [10:39:47] (03CR) 10CI reject: [V:04-1] zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [10:40:07] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet [10:42:36] (03PS3) 10Slyngshede: P:tofurkey enable Tofurkey for MAGRU [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) [10:43:08] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet [10:43:10] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet [10:43:10] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw [10:43:57] (03CR) 10Hnowlan: [V:03+1] "PCC SUCCESS (DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9293/console" [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [10:45:15] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [10:47:49] (03PS12) 10Hashar: zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) [10:48:37] (03CR) 10CI reject: [V:04-1] zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [10:48:51] sigh [10:49:16] (03PS4) 10Slyngshede: P:tofurkey enable Tofurkey for MAGRU [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) [10:49:19] (03PS1) 10JMeybohm: coredns: Remove loop from commonplugins [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328163 (https://phabricator.wikimedia.org/T427864) [10:49:44] why do `.rspec` requires a SPDX header :\ [10:50:23] RESOLVED: CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:51:08] (03PS2) 10JMeybohm: coredns: Remove loop from commonplugins [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328163 (https://phabricator.wikimedia.org/T427864) [10:52:47] (03PS13) 10Hashar: zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) [10:53:03] !log jmm@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage [10:54:16] (03CR) 10Hashar: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [10:55:07] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [10:55:17] (03PS1) 10Marostegui: table-catalog.yaml: Remove wikifunctionsclient_usage [puppet] - 10https://gerrit.wikimedia.org/r/1328164 (https://phabricator.wikimedia.org/T434541) [10:57:40] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2005.wikimedia.org with reason: host reimage [11:00:04] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260821T0700) [11:00:05] jelto, arnoldokoth, mutante, and arnaudb: That opportune time for a GitLab version upgrades deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260821T1100). [11:01:53] (03CR) 10Hashar: "PCC https://puppet-compiler.wmflabs.org/output/1327569/7600/" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [11:04:35] (03CR) 10Slyngshede: P:tofurkey enable Tofurkey for MAGRU (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [11:10:47] PROBLEM - Postfix SMTP on crm2001 is CRITICAL: CRITICAL - Certificate crm2001.codfw.wmnet expires in 15 day(s) (Sun 06 Sep 2026 11:10:00 AM GMT +0000). https://wikitech.wikimedia.org/wiki/Mail%23Troubleshooting [11:11:32] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2005.wikimedia.org with OS trixie [11:11:37] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12240953 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin2003 for host urldownloader2005.wikimedia.org with OS trixie completed: - urldownloader2005 (**PASS**) - Dow... [11:13:35] (03CR) 10Hnowlan: beta-logs: add options required for logstash output plugin (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [11:15:17] (03CR) 10Blake: [C:03+1] trafficserver: Support testwiki pretrain routing in XWD [puppet] - 10https://gerrit.wikimedia.org/r/1327621 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [11:16:42] (03PS7) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [11:19:23] jmm@cumin2003 reimage (PID 3615763) is awaiting input [11:20:24] !log jmm@cumin2003 START - Cookbook sre.hosts.reimage for host urldownloader2006.wikimedia.org with OS trixie [11:20:37] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12240957 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin2003 for host urldownloader2006.wikimedia.org with OS trixie [11:21:22] (03PS1) 10Hnowlan: idp: remove graphite configuration [puppet] - 10https://gerrit.wikimedia.org/r/1328165 (https://phabricator.wikimedia.org/T435341) [11:21:28] (03PS1) 10Muehlenhoff: Re-apply the urldownloader role to the new Trixie hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328166 (https://phabricator.wikimedia.org/T427282) [11:22:06] (03PS8) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [11:23:03] (03CR) 10Muehlenhoff: [C:03+1] "Looks good. When this has been merged, can you please remove graphite from" [puppet] - 10https://gerrit.wikimedia.org/r/1328165 (https://phabricator.wikimedia.org/T435341) (owner: 10Hnowlan) [11:28:32] (03PS1) 10Kevin Bazira: ml-services: deploy outlink isvc with updated wiki list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328167 (https://phabricator.wikimedia.org/T435586) [11:29:16] (03PS1) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [11:31:46] (03CR) 10CI reject: [V:04-1] Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [11:32:05] (03CR) 10Bartosz Wójtowicz: [C:03+1] ml-services: deploy outlink isvc with updated wiki list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328167 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [11:33:48] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy outlink isvc with updated wiki list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328167 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [11:36:21] (03Merged) 10jenkins-bot: ml-services: deploy outlink isvc with updated wiki list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328167 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [11:38:01] !log jmm@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage [11:41:07] (03PS1) 10Muehlenhoff: jenkins:: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) [11:41:26] (03PS2) 10Muehlenhoff: jenkins:: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) [11:41:40] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [11:43:08] (03PS2) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [11:43:36] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2006.wikimedia.org with reason: host reimage [11:44:13] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [11:45:33] (03CR) 10CI reject: [V:04-1] Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [11:53:05] (03CR) 10Ladsgroup: [C:03+1] table-catalog.yaml: Remove wikifunctionsclient_usage [puppet] - 10https://gerrit.wikimedia.org/r/1328164 (https://phabricator.wikimedia.org/T434541) (owner: 10Marostegui) [11:53:23] (03PS3) 10Muehlenhoff: jenkins:: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) [11:53:54] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [11:53:58] (03PS1) 10Hnowlan: alertmanager: add irc output for critical netops alerts [puppet] - 10https://gerrit.wikimedia.org/r/1328172 (https://phabricator.wikimedia.org/T432376) [11:57:55] (03PS3) 10Hnowlan: team-sre: Add data-engineering tag [alerts] - 10https://gerrit.wikimedia.org/r/1319827 (https://phabricator.wikimedia.org/T432376) [11:57:55] (03PS3) 10Hnowlan: team-sre: add data-persistence [alerts] - 10https://gerrit.wikimedia.org/r/1319830 (https://phabricator.wikimedia.org/T432376) [11:57:55] (03PS3) 10Hnowlan: team-sre: add dcops tag [alerts] - 10https://gerrit.wikimedia.org/r/1319831 (https://phabricator.wikimedia.org/T432376) [11:57:56] (03PS3) 10Hnowlan: team-sre: add infrastructure-foundations tag [alerts] - 10https://gerrit.wikimedia.org/r/1319832 (https://phabricator.wikimedia.org/T432376) [11:57:57] (03PS2) 10Hnowlan: team-sre: add serviceops tag [alerts] - 10https://gerrit.wikimedia.org/r/1327076 (https://phabricator.wikimedia.org/T432376) [11:58:01] (03PS2) 10Hnowlan: team-netops: replace use of "team: sre" with "team: netops" [alerts] - 10https://gerrit.wikimedia.org/r/1327077 (https://phabricator.wikimedia.org/T432376) [11:58:15] (03PS2) 10Marostegui: installserver: Do not reimage db1283 [puppet] - 10https://gerrit.wikimedia.org/r/1327849 [11:58:16] (03PS1) 10Marostegui: filtered_tables.txt: Remove two columns [puppet] - 10https://gerrit.wikimedia.org/r/1328174 (https://phabricator.wikimedia.org/T435064) [11:58:27] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [11:59:08] (03CR) 10Marostegui: "FYI" [puppet] - 10https://gerrit.wikimedia.org/r/1328174 (https://phabricator.wikimedia.org/T435064) (owner: 10Marostegui) [11:59:25] (03CR) 10Marostegui: [C:03+2] filtered_tables.txt: Remove two columns [puppet] - 10https://gerrit.wikimedia.org/r/1328174 (https://phabricator.wikimedia.org/T435064) (owner: 10Marostegui) [11:59:44] (03CR) 10Marostegui: [V:03+2 C:03+2] installserver: Do not reimage db1283 [puppet] - 10https://gerrit.wikimedia.org/r/1327849 (owner: 10Marostegui) [12:00:17] !log kevinbazira@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [12:00:47] (03CR) 10CI reject: [V:04-1] team-sre: add infrastructure-foundations tag [alerts] - 10https://gerrit.wikimedia.org/r/1319832 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [12:01:28] (03CR) 10CI reject: [V:04-1] team-sre: Add data-engineering tag [alerts] - 10https://gerrit.wikimedia.org/r/1319827 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [12:01:33] (03CR) 10CI reject: [V:04-1] team-sre: add dcops tag [alerts] - 10https://gerrit.wikimedia.org/r/1319831 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [12:01:40] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2006.wikimedia.org with OS trixie [12:01:46] (03CR) 10CI reject: [V:04-1] team-sre: add data-persistence [alerts] - 10https://gerrit.wikimedia.org/r/1319830 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [12:01:52] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12241094 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin2003 for host urldownloader2006.wikimedia.org with OS trixie completed: - urldownloader2... [12:02:03] (03CR) 10CI reject: [V:04-1] team-sre: add serviceops tag [alerts] - 10https://gerrit.wikimedia.org/r/1327076 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [12:02:38] (03CR) 10CI reject: [V:04-1] team-netops: replace use of "team: sre" with "team: netops" [alerts] - 10https://gerrit.wikimedia.org/r/1327077 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [12:03:23] (03PS3) 10Hnowlan: team-sre: add serviceops tag [alerts] - 10https://gerrit.wikimedia.org/r/1327076 (https://phabricator.wikimedia.org/T432376) [12:03:38] (03PS3) 10Hnowlan: team-netops: replace use of "team: sre" with "team: netops" [alerts] - 10https://gerrit.wikimedia.org/r/1327077 (https://phabricator.wikimedia.org/T432376) [12:04:01] (03PS1) 10Filippo Giunchedi: Revert "site: get cloudvirts ready for E4 -> C8 move" [puppet] - 10https://gerrit.wikimedia.org/r/1328176 (https://phabricator.wikimedia.org/T431682) [12:04:46] (03CR) 10CI reject: [V:04-1] Revert "site: get cloudvirts ready for E4 -> C8 move" [puppet] - 10https://gerrit.wikimedia.org/r/1328176 (https://phabricator.wikimedia.org/T431682) (owner: 10Filippo Giunchedi) [12:06:14] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12241104 (10fgiunchedi) a:05VRiley-WMF→03fgiunchedi >>! In T431682#12239559, @VRiley-WMF wrote: > Okay, so... > > After troubleshooting it for a while, I thou... [12:06:42] (03PS3) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [12:08:47] (03CR) 10CI reject: [V:04-1] Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [12:10:23] (03PS2) 10Filippo Giunchedi: Revert "site: get cloudvirts ready for E4 -> C8 move" [puppet] - 10https://gerrit.wikimedia.org/r/1328176 (https://phabricator.wikimedia.org/T431682) [12:11:23] !log filippo@cumin1003 START - Cookbook sre.dns.netbox [12:13:00] (03PS4) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [12:14:04] (03PS1) 10Muehlenhoff: service::node: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328178 (https://phabricator.wikimedia.org/T429175) [12:14:56] (03CR) 10CI reject: [V:04-1] Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [12:16:30] !log filippo@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: T431682 - filippo@cumin1003" [12:16:34] !log filippo@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: T431682 - filippo@cumin1003" [12:16:34] !log filippo@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:16:35] T431682: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682 [12:22:12] (03PS5) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [12:22:59] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328178 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [12:26:12] FIRING: VarnishUnavailable: varnish-text has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/Varnish#Diagnosing_Varnish_alerts - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=3 - https://alerts.wikimedia.org/?q=alertname%3DVarnishUnavailable [12:26:13] FIRING: HaproxyUnavailable: HAProxy (cache_text) has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyUnavailable [12:28:31] (03CR) 10Marostegui: "I will wait to merge this once the table has been dropped (will do so next week, don't want to drop stuff on a Friday)" [puppet] - 10https://gerrit.wikimedia.org/r/1328164 (https://phabricator.wikimedia.org/T434541) (owner: 10Marostegui) [12:31:12] RESOLVED: VarnishUnavailable: varnish-text has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/Varnish#Diagnosing_Varnish_alerts - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=3 - https://alerts.wikimedia.org/?q=alertname%3DVarnishUnavailable [12:31:13] RESOLVED: HaproxyUnavailable: HAProxy (cache_text) has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyUnavailable [12:34:02] (03PS6) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [12:35:36] (03CR) 10Hashar: "Actually we can get the full diff for each hosts. Hosts with changes:" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [12:58:37] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:00:18] (03CR) 10Arnaudb: [C:03+2] pontoon: exception handle in pontoonctl destroy-hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328150 (https://phabricator.wikimedia.org/T420184) (owner: 10Arnaudb) [13:03:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:08:07] (03PS4) 10Hnowlan: team-sre: add data-persistence [alerts] - 10https://gerrit.wikimedia.org/r/1319830 (https://phabricator.wikimedia.org/T432376) [13:08:27] 06SRE, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): install1005 running out of disk due to squid log volume from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12241286 (10bking) [13:14:08] !log hashar@deploy1003 Started deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - T435392 [13:14:14] T435392: Publish PersonalDashboard documentation to doc.wikimedia.org - https://phabricator.wikimedia.org/T435392 [13:14:23] !log hashar@deploy1003 Finished deploy [integration/docroot@2d5ff9b]: opensource: add PersonalDashboard docs to MW components - T435392 (duration: 00m 15s) [13:19:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqord (208.80.154.198) - group Confed_eqord - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqord&var-bgp_neighbor=cr2-eqord - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:21:23] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12241305 (10bking) [13:23:59] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" [13:24:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqord (208.80.154.198) - group Confed_eqord - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqord&var-bgp_neighbor=cr2-eqord - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:24:43] (03PS1) 10Cathal Mooney: Remove eqord from configuration [puppet] - 10https://gerrit.wikimedia.org/r/1328183 (https://phabricator.wikimedia.org/T427050) [13:25:28] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "sync cr2-eqord router offline - cmooney@cumin1003" [13:32:40] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:36:21] (03CR) 10Hashar: "Thank you so much for the documentation update! Replacing an existing plugin is what upstream recommends (after they have moved Gerrit FE" [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327666 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [13:37:05] (03PS2) 10Hashar: plugins-router: set the MIME type of a served plugin [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327668 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [13:38:17] (03CR) 10Ssingh: [C:03+1] "We can reimage this today, or Monday, whatever you are comfortable with." [puppet] - 10https://gerrit.wikimedia.org/r/1327156 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:38:19] (03CR) 10Hashar: [C:03+2] "Nit: I have amended the commit message to rewrap it at 78 characters" [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327668 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [13:39:02] (03Merged) 10jenkins-bot: plugins-router: set the MIME type of a served plugin [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327668 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [13:46:43] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12241360 (10MoritzMuehlenhoff) [13:54:03] (03PS1) 10Cathal Mooney: Remove cr2-eqord and references to the site. [homer/public] - 10https://gerrit.wikimedia.org/r/1328185 (https://phabricator.wikimedia.org/T427050) [13:56:14] (03CR) 10Hashar: [C:03+1] "That is quite impressive and I am pretty sure that feature has been a very long ask from our developers." [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327669 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [13:56:47] (03CR) 10Jaime Nuche: [C:03+1] jenkins:: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:57:36] (03CR) 10Ssingh: [C:03+1] jenkins:: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:58:35] (03CR) 10Ssingh: [C:03+1] service::node: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328178 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [14:01:28] (03CR) 10Jaime Nuche: [C:03+1] "@mmuhlenhoff@wikimedia.org thank for you this patch. I've updated the corresponding Jenkins releases config: https://gitlab.wikimedia.org/" [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [14:01:47] (03PS2) 10CDanis: Remove eqord from configuration [puppet] - 10https://gerrit.wikimedia.org/r/1328183 (https://phabricator.wikimedia.org/T427050) (owner: 10Cathal Mooney) [14:01:48] (03CR) 10CDanis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328183 (https://phabricator.wikimedia.org/T427050) (owner: 10Cathal Mooney) [14:02:02] (03CR) 10Ssingh: [C:03+1] Re-apply the urldownloader role to the new Trixie hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328166 (https://phabricator.wikimedia.org/T427282) (owner: 10Muehlenhoff) [14:03:36] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12241428 (10bking) 05Open→03In progress a:03bking After consulting with @VRiley-WMF in IRC, I'm going to grab this ticket and try the firmware upd... [14:07:41] (03PS1) 10SomeRandomDeveloper: Profiler: Fix excimer component regex skipping the last frame in each stack [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328191 [14:08:51] (03CR) 10CDanis: [C:03+1] Remove eqord from configuration [puppet] - 10https://gerrit.wikimedia.org/r/1328183 (https://phabricator.wikimedia.org/T427050) (owner: 10Cathal Mooney) [14:08:58] (03PS2) 10SomeRandomDeveloper: Profiler: Fix excimer component regex skipping the last frame in each stack [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328191 [14:09:37] (03CR) 10Cathal Mooney: [C:03+2] Remove cr2-eqord and references to the site. [homer/public] - 10https://gerrit.wikimedia.org/r/1328185 (https://phabricator.wikimedia.org/T427050) (owner: 10Cathal Mooney) [14:09:47] (03CR) 10Cathal Mooney: [C:03+2] Remove eqord from configuration [puppet] - 10https://gerrit.wikimedia.org/r/1328183 (https://phabricator.wikimedia.org/T427050) (owner: 10Cathal Mooney) [14:11:01] (03Merged) 10jenkins-bot: Remove cr2-eqord and references to the site. [homer/public] - 10https://gerrit.wikimedia.org/r/1328185 (https://phabricator.wikimedia.org/T427050) (owner: 10Cathal Mooney) [14:11:22] !log imported openjdk 8u504-ga-1~deb12u1 for bookworm-wikimedia (backport of the latest Java 8 security fixes for bookworm) [14:11:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:11:51] (03CR) 10SomeRandomDeveloper: "I noticed this when implementing this functionality (sending excimer samples to statsd) based on this code in a MediaWiki extension and wr" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328191 (owner: 10SomeRandomDeveloper) [14:13:03] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in misc modules [puppet] - 10https://gerrit.wikimedia.org/r/1327643 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:13:54] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [14:15:38] (03CR) 10CDobbins: [C:03+2] hieradata: use pdns v5 cfg flag on dns5004 [puppet] - 10https://gerrit.wikimedia.org/r/1327156 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [14:18:53] (03PS7) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [14:19:44] cmooney@cumin1003 netbox (PID 2368479) is awaiting input [14:21:21] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" [14:21:25] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove entries for cr2-eqord - cmooney@cumin1003" [14:21:25] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:28:41] 10ops-eqiad, 06SRE, 06DC-Ops: Recycle old MPC 3D 16x10G line cards in cr2-eqiad - https://phabricator.wikimedia.org/T435588#12241461 (10RobH) I can ask but typically we stored the old blanks in the storage area, aren't they still there? [14:29:08] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=dns5004.* [reason: trixie upgrade] [14:30:19] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host dns5004.wikimedia.org with OS trixie [14:30:22] 10ops-eqiad, 06SRE, 06DC-Ops: Recycle old MPC 3D 16x10G line cards in cr2-eqiad - https://phabricator.wikimedia.org/T435588#12241477 (10cmooney) >>! In T435588#12241461, @RobH wrote: > I can ask but typically we stored the old blanks in the storage area, aren't they still there? I think Val said there was o... [14:33:21] 10ops-eqiad, 06SRE, 06DC-Ops: Recycle old MPC 3D 16x10G line cards in cr2-eqiad - https://phabricator.wikimedia.org/T435588#12241534 (10RobH) [14:33:52] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:35:10] FIRING: [2x] BFDdown: BFD session down between cr2-eqsin and 103.102.166.36 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:35:23] FIRING: GnmiInterfaceCountersDrop: ... [14:35:23] lsw1-c7-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-c7-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [14:36:30] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:36:42] RECOVERY - OSPF status on cr1-magru is OK: OSPFv2: 5/5 UP : OSPFv3: 5/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:40:10] FIRING: [4x] BFDdown: BFD session down between cr2-eqsin and 103.102.166.36 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:43:17] (03CR) 10Blake: [C:03+1] team-sre: add serviceops tag [alerts] - 10https://gerrit.wikimedia.org/r/1327076 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [14:45:23] RESOLVED: GnmiInterfaceCountersDrop: ... [14:45:23] lsw1-c7-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-c7-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [14:46:53] 06SRE: Add alerting for gnmic total series counters - https://phabricator.wikimedia.org/T435184#12241640 (10cmooney) So this fired today: ` 2026-08-21 14:34 lsw1-c7-eqiad - lsw1-c7-eqiad is exporting less than half the gNMI interface counters it had 24h ago ` I restarted the gnmic service on netflow1003 as a re... [14:53:49] !log andrew@cumin2003 START - Cookbook sre.hosts.reimage for host cloudcephosd1042.eqiad.wmnet with OS bookworm [14:56:42] 06SRE, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): install1005 running out of disk due to squid log volume from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12241656 (10Scott_French) @RKemper - Ah, that's great! Yes, if these workloads are able to... [15:05:48] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on dns5004.wikimedia.org with reason: host reimage [15:06:28] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploy v1.5.3 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328196 (https://phabricator.wikimedia.org/T397417) [15:09:20] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns5004.wikimedia.org with reason: host reimage [15:12:27] (03CR) 10Ahmon Dancy: [C:03+2] plugins: Add wm-wheres-my-code-running plugin [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327669 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [15:13:12] (03Merged) 10jenkins-bot: plugins: Add wm-wheres-my-code-running plugin [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1327669 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [15:13:46] !log andrew@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage [15:15:59] PROBLEM - Recursive DNS on 103.102.166.36 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [15:17:08] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploy v1.5.3 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328197 (https://phabricator.wikimedia.org/T397417) [15:17:39] !log dancy@deploy1003 Started deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 (T434726) [15:17:44] T434726: Where is my change running tool - https://phabricator.wikimedia.org/T434726 [15:17:54] !log dancy@deploy1003 Finished deploy [gerrit/gerrit@2cc11cc]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1327669 (T434726) (duration: 00m 14s) [15:18:02] (03PS2) 10Santiago Faci: Test Kitchen UI: Deploy v1.5.3 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328197 (https://phabricator.wikimedia.org/T397417) [15:18:22] !log andrew@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1042.eqiad.wmnet with reason: host reimage [15:21:00] PROBLEM - Recursive DNS on 2001:df2:e500:2:103:102:166:36 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [15:22:54] RECOVERY - Host krb1002 is UP: PING OK - Packet loss = 0%, RTA = 0.41 ms [15:23:25] (03PS7) 10Cwhite: beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) [15:23:25] (03PS4) 10Cwhite: logstash: add security-plugin required fields to output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) [15:23:39] FIRING: CoreBGPDown: Core BGP session down between cr2-eqdfw and cr2-magru (2a02:ec80:700:fe0b::2) - group Confed_magru - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-eqdfw:9804&var-bgp_group=Confed_magru&var-bgp_neighbor=cr2-magru - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:23:54] PROBLEM - SSH on krb1002 is CRITICAL: connect to address 10.64.32.69 and port 22: Connection refused https://wikitech.wikimedia.org/wiki/SSH/monitoring [15:25:10] FIRING: [5x] BFDdown: BFD session down between cr2-eqdfw and 2a02:ec80:700:fe0b::2 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:27:06] (03CR) 10Cwhite: beta-logs: add options required for logstash output plugin (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [15:28:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-eqdfw and cr2-magru (2a02:ec80:700:fe0b::2) - group Confed_magru - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-eqdfw:9804&var-bgp_group=Confed_magru&var-bgp_neighbor=cr2-magru - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:30:10] FIRING: [5x] BFDdown: BFD session down between cr2-eqdfw and 2a02:ec80:700:fe0b::2 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:35:42] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [15:36:56] RECOVERY - Recursive DNS on 2001:df2:e500:2:103:102:166:36 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [15:37:00] RECOVERY - Recursive DNS on 103.102.166.36 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [15:38:10] !log andrew@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1042.eqiad.wmnet with OS bookworm [15:38:58] 07sre-alert-triage, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Alert in need of triage: PuppetFailure (instance an-test-client1002:9100) - https://phabricator.wikimedia.org/T427399#12241864 (10BTullis) 05Open→03Resolved a:03BTullis Pupet is running cleanly. Closing. [15:40:10] RESOLVED: [5x] BFDdown: BFD session down between cr2-eqdfw and 2a02:ec80:700:fe0b::2 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:42:10] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12241881 (10VRiley-WMF) Both the BIOS and BMC have been updated with firmware. [15:43:42] (03CR) 10Scott French: [C:03+1] "Thanks, Ahmon. I think this seems reasonable given the relative rarity of non-Deployment workloads among mediawiki services (i.e., it's no" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326362 (https://phabricator.wikimedia.org/T375514) (owner: 10Ahmon Dancy) [15:51:51] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [15:51:51] !log cmooney@cumin1003 END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) [15:53:14] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [15:55:39] (03PS1) 10Btullis: Declare the webrequest.dumps.v1 stream in EventStreamConfig [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328201 (https://phabricator.wikimedia.org/T425087) [15:55:44] (03PS1) 10Btullis: Remove the webrequest.dumps.dev0 stream from EventStreamConfig [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328202 (https://phabricator.wikimedia.org/T425087) [15:55:58] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" [15:56:50] !log cdobbins@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host dns5004.wikimedia.org with OS trixie [15:57:53] (03PS1) 10Btullis: logstash: Consume the webrequest.dumps.v1 stream from Kafka [puppet] - 10https://gerrit.wikimedia.org/r/1328205 (https://phabricator.wikimedia.org/T425087) [15:57:59] (03PS1) 10Btullis: dumps: web: Produce access logs to the webrequest.dumps.v1 stream [puppet] - 10https://gerrit.wikimedia.org/r/1328206 (https://phabricator.wikimedia.org/T425087) [15:58:32] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:59:02] cmooney@cumin1003 netbox (PID 2382490) is awaiting input [15:59:12] (03CR) 10Aklapper: [V:03+2 C:03+2] "Thanks! Applies locally and language settings still render. :)" [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1326873 (owner: 10Pppery) [15:59:19] (03PS1) 10Btullis: logstash: Stop consuming the webrequest.dumps.dev0 stream from Kafka [puppet] - 10https://gerrit.wikimedia.org/r/1328208 (https://phabricator.wikimedia.org/T425087) [15:59:57] (03CR) 10Lerickson: [C:03+2] Update the Qlever image to include named subqueries. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327642 (https://phabricator.wikimedia.org/T435122) (owner: 10Lerickson) [16:00:24] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqord<->codfw arelion - cmooney@cumin1003" [16:00:24] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:00:24] (03PS1) 10Cathal Mooney: remove include statement for disused 2620:0:860:fe02::/64 [dns] - 10https://gerrit.wikimedia.org/r/1328209 [16:01:32] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12241958 (10bking) The host stalls during the boot process and asks for the root password. I'm logged in via console trying some kernel updates, but it... [16:01:42] (03CR) 10Hnowlan: [C:03+1] beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [16:01:50] (03CR) 10Cathal Mooney: [C:03+2] remove include statement for disused 2620:0:860:fe02::/64 [dns] - 10https://gerrit.wikimedia.org/r/1328209 (owner: 10Cathal Mooney) [16:02:10] 06SRE, 10SRE-Access-Requests, 06tools-infrastructure-team, 13Patch-For-Review: Requesting access to wmcs-roots for bliviero - https://phabricator.wikimedia.org/T435123#12241959 (10mark) >>! In T435123#12223551, @Andrew wrote: > @mark in theory Belinda is the approver of record for this group but we would l... [16:02:14] !log cmooney@dns3003 START - running authdns-update [16:02:25] (03CR) 10Hnowlan: [C:03+1] logstash: add security-plugin required fields to output plugin (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [16:02:29] (03Merged) 10jenkins-bot: Update the Qlever image to include named subqueries. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327642 (https://phabricator.wikimedia.org/T435122) (owner: 10Lerickson) [16:02:41] (03PS1) 10Ssingh: Revert "wikimedia.org: add TXT record for BIMI" [dns] - 10https://gerrit.wikimedia.org/r/1328210 [16:04:58] !log cmooney@dns3003 END - running authdns-update [16:06:22] (03CR) 10Ssingh: [C:03+2] Revert "wikimedia.org: add TXT record for BIMI" [dns] - 10https://gerrit.wikimedia.org/r/1328210 (owner: 10Ssingh) [16:06:39] !log sukhe@dns1004 START - running authdns-update [16:08:21] (03PS1) 10Ryan Kemper: cumin: drop duplicate mediabackup-storage alias [puppet] - 10https://gerrit.wikimedia.org/r/1328211 (https://phabricator.wikimedia.org/T419831) [16:08:34] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [16:08:55] !log sukhe@dns1004 END - running authdns-update [16:09:48] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1328211 (https://phabricator.wikimedia.org/T419831) (owner: 10Ryan Kemper) [16:11:56] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [16:12:25] (03PS1) 10Ladsgroup: turnilo: Expose X-analytics thumb_generated in webrequest_sampled_live [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) [16:17:49] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12242055 (10VRiley-WMF) After looking into this a bit, it seems we can set these to performance mode, but... [16:17:56] (03CR) 10Andrew Bogott: [C:03+2] "mark approved on the associated task" [puppet] - 10https://gerrit.wikimedia.org/r/1326343 (https://phabricator.wikimedia.org/T435123) (owner: 10Andrew Bogott) [16:24:40] PROBLEM - Host krb1002 is DOWN: PING CRITICAL - Packet loss = 100% [16:26:37] 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: BFD session fails from Anycast hosts over IPv6 on boot - https://phabricator.wikimedia.org/T434806#12242098 (10cmooney) So actually I may have spoken too soon here. We've had a couple of DNS hosts reimaged/rebooted in the past few days, and none of... [16:27:07] (03PS2) 10Milazg: Add configurable RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) [16:29:15] (03CR) 10RLazarus: [C:03+1] "Whoops. Good catch, thanks." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328163 (https://phabricator.wikimedia.org/T427864) (owner: 10JMeybohm) [16:29:39] 06SRE, 10Wikimedia-Mailing-lists: Create new mailing lists: wikidata-admins@lists.wikimedia.org - https://phabricator.wikimedia.org/T435638#12242105 (10Ladsgroup) The consensus is not really strong there. I write a message ask for objections. [16:31:33] (03PS3) 10Milazg: Add configurable RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) [16:31:49] (03PS4) 10Milazg: Add configurable RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) [16:33:10] 06SRE, 10Wikimedia-Mailing-lists: Create new mailing lists: wikidata-admins@lists.wikimedia.org - https://phabricator.wikimedia.org/T435638#12242111 (10Ladsgroup) 05Open→03Stalled https://www.wikidata.org/wiki/Wikidata:Administrators%27_noticeboard#Any_objections_to_creation_of_private_admins_mailing_list... [16:33:12] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [16:36:44] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:37:44] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" [16:37:55] (03PS1) 10Cathal Mooney: Remove include for IPv6 range used on old Telia cct to SG [dns] - 10https://gerrit.wikimedia.org/r/1328214 [16:39:10] FIRING: BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:40:48] cmooney@cumin1003 netbox (PID 2386340) is awaiting input [16:41:04] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on eqsin<->codfw arelion - cmooney@cumin1003" [16:41:04] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:41:05] (03CR) 10Cathal Mooney: [C:03+2] Remove include for IPv6 range used on old Telia cct to SG [dns] - 10https://gerrit.wikimedia.org/r/1328214 (owner: 10Cathal Mooney) [16:41:21] !log cmooney@dns3003 START - running authdns-update [16:41:59] (03CR) 10Xcollazo: [C:03+1] dumps: web: Produce access logs to the webrequest.dumps.v1 stream [puppet] - 10https://gerrit.wikimedia.org/r/1328206 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [16:44:11] !log cmooney@dns3003 END - running authdns-update [16:45:48] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12242143 (10bking) Update: the disk(s) don't seem to be bootable. I'll try a few more things before giving up. [16:47:05] (03PS1) 10Milazg: Remove mode from RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328215 (https://phabricator.wikimedia.org/T434267) [16:47:34] 06SRE, 06Infrastructure-Foundations, 10Puppet-Infrastructure, 13Patch-For-Review: Fix remaining scoped legacy fact usage - https://phabricator.wikimedia.org/T435225#12242148 (10jhathaway) [16:49:32] (03PS3) 10Cathal Mooney: Add magru HE OSPF ints and remove GTT ints for OSPF adjacency [homer/public] - 10https://gerrit.wikimedia.org/r/1327546 (https://phabricator.wikimedia.org/T424839) [16:52:07] !log cdobbins@cumin1003 START - Cookbook sre.hosts.remove-downtime for dns5004.wikimedia.org [16:52:08] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns5004.wikimedia.org [16:53:16] !log cdobbins@cumin1003 conftool action : set/pooled=yes; selector: name=dns5004.* [reason: trixie upgrade] [16:54:10] (03CR) 10Ssingh: P:tofurkey enable Tofurkey for MAGRU (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [16:54:18] (03CR) 10Cathal Mooney: [C:03+2] Add magru HE OSPF ints and remove GTT ints for OSPF adjacency [homer/public] - 10https://gerrit.wikimedia.org/r/1327546 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [16:55:30] (03CR) 10Milazg: "Updated the patch to have both and created this new one https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1328215 to remove `" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [16:55:52] (03Merged) 10jenkins-bot: Add magru HE OSPF ints and remove GTT ints for OSPF adjacency [homer/public] - 10https://gerrit.wikimedia.org/r/1327546 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [16:56:06] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [16:58:52] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:59:54] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [17:01:59] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [17:04:50] PROBLEM - check if authdns-update was run after a change was merged to operations/dns.git on dns5004 is CRITICAL: Local zone files are NOT in sync with operations/dns.git (SHA: local is f517af6dba91975fbdde1583a19ff94bb8414e6c, dns.git is 734c650cfbe66a4616f3098831114641e9607152) https://wikitech.wikimedia.org/wiki/DNS%23authdns_update_run [17:04:57] ah [17:05:06] !log sukhe@dns1004 START - running authdns-update [17:05:24] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [17:07:18] !log sukhe@dns1004 FAIL - running authdns-update [17:07:31] 10ops-eqiad, 06SRE, 06DC-Ops: Recycle old MPC 3D 16x10G line cards in cr2-eqiad - https://phabricator.wikimedia.org/T435588#12242184 (10VRiley-WMF) Correct, there is only one blank that we have onsite. Admittedly, I would recomment keeping the other two sitting in the slot because there are cables that are h... [17:08:00] !log sukhe@puppetserver1001 conftool action : set/pooled=no; selector: name=dns5004.wikimedia.org [reason: resolving authdns-update issues] [17:08:04] !log sukhe@dns1004 START - running authdns-update [17:09:46] RECOVERY - check if authdns-update was run after a change was merged to operations/dns.git on dns5004 is OK: The check was skipped as the host is not pooled for authdns-update https://wikitech.wikimedia.org/wiki/DNS%23authdns_update_run [17:10:16] !log sukhe@dns1004 END - running authdns-update [17:11:07] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12242202 (10VRiley-WMF) @MoritzMuehlenhoff We can do this on Monday the 24th if that works for you? [17:12:53] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [17:16:46] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [17:23:46] !log sukhe@cumin1003 START - Cookbook sre.dns.netbox [17:24:03] !log sukhe@cumin1003 END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) [17:24:10] RESOLVED: BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:24:12] !log sukhe@cumin1003 START - Cookbook sre.dns.netbox [17:28:00] !log sukhe@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to be26e30ae101 - sukhe@cumin1003" [17:28:04] !log sukhe@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to be26e30ae101 - sukhe@cumin1003" [17:28:04] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:28:37] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: service=authdns-update,name=dns5004.wikimedia.org [reason: resolving authdns-update issues] [17:28:40] !log sukhe@cumin1003 START - Cookbook sre.dns.netbox [17:29:39] (03CR) 10Pppery: "(Sidenote: "language setting still render" can't possibly be an issue for source string updates as opposed to translation updates - source" [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1326873 (owner: 10Pppery) [17:31:25] (03PS1) 10Ssingh: services: make urldownloader probe paging [puppet] - 10https://gerrit.wikimedia.org/r/1328218 (https://phabricator.wikimedia.org/T429175) [17:32:31] !log sukhe@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to be26e30ae101 - sukhe@cumin1003" [17:32:35] !log sukhe@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: force HEAD to be26e30ae101 - sukhe@cumin1003" [17:32:35] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:32:59] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=dns5004.wikimedia.org [reason: resolved authdns-update issues] [17:33:06] !log sukhe@dns1004 START - running authdns-update [17:35:18] !log sukhe@dns1004 END - running authdns-update [17:35:44] RECOVERY - OSPF status on cr1-magru is OK: OSPFv2: 5/5 UP : OSPFv3: 5/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:35:55] all DNS hosts maint done [17:37:32] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:38:04] (03CR) 10Cwhite: [C:03+2] beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [17:39:01] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 06Traffic, 13Patch-For-Review: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12242324 (10ssingh) 05Resolved→03Open [17:39:10] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 06Traffic, 13Patch-For-Review: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12242326 (10ssingh) [17:39:21] (03PS2) 10Ssingh: services: make urldownloader probe paging [puppet] - 10https://gerrit.wikimedia.org/r/1328218 (https://phabricator.wikimedia.org/T435648) [18:03:11] (03PS2) 10Lerickson: Bump wdqs-proxy 0.4.0 -> 0.6.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327630 (https://phabricator.wikimedia.org/T433935) [18:08:52] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:11:09] (03PS1) 10Cwhite: beta-logs: enable security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1328222 (https://phabricator.wikimedia.org/T350516) [18:15:02] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12242459 (10bking) The disks are detected in BIOS, but GRUB is not detected. I'll try a reimage, but I'm not hopeful. [18:16:01] (03CR) 10Lerickson: [C:03+2] "Discussed this with gmodena earlier today. I have done a home-dir deploy of this to staging, tested it, and it works perfectly: the eventg" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327630 (https://phabricator.wikimedia.org/T433935) (owner: 10Lerickson) [18:16:03] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host krb1002.eqiad.wmnet with OS bookworm [18:16:18] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12242461 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by bking@cumin2003 for host krb1002.eqiad.wmnet with OS bookworm [18:18:42] (03Merged) 10jenkins-bot: Bump wdqs-proxy 0.4.0 -> 0.6.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327630 (https://phabricator.wikimedia.org/T433935) (owner: 10Lerickson) [18:27:26] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [18:27:52] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [18:32:51] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12242483 (10BLiviero-WMF) HI! will let @Andrew weigh in on whether we can depool and make this change to... [18:35:26] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module wmflib [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [18:35:33] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [18:35:50] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [18:39:47] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12242494 (10Andrew) Thanks for the followup @VRiley-WMF -- indeed these will need to be drained before we... [18:50:44] PROBLEM - Swift https backend on ms-fe1015 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Swift [18:51:01] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [18:51:30] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [18:51:34] RECOVERY - Swift https backend on ms-fe1015 is OK: HTTP OK: HTTP/1.1 200 OK - 567 bytes in 0.063 second response time https://wikitech.wikimedia.org/wiki/Swift [18:58:52] FIRING: [2x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:59:56] FIRING: [3x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:59:56] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [19:00:31] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [19:01:10] 06SRE, 10SRE-Access-Requests, 06tools-infrastructure-team: Requesting access to wmcs-roots for bliviero - https://phabricator.wikimedia.org/T435123#12242511 (10Andrew) 05Stalled→03Resolved a:03Andrew she's in! [19:03:52] FIRING: [3x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:04:46] RECOVERY - Backup freshness on backup1014 is OK: Fresh: 140 jobs https://wikitech.wikimedia.org/wiki/Bacula%23Monitoring [19:12:56] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12242534 (10RobH) Updates: * Chatted with a rep today via 1-800-225-5345 * The shipment landed in France on August 7th - https://mydhl.express.dhl/us/en/tracking.html#/resu... [19:13:52] FIRING: [3x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:14:56] FIRING: [3x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:15:50] PROBLEM - Check unit status of statograph_post on alert1002 is CRITICAL: CRITICAL: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:16:05] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts rspec hiera config [puppet] - 10https://gerrit.wikimedia.org/r/1328229 (https://phabricator.wikimedia.org/T435225) [19:20:57] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12242550 (10RobH) DHL just called me back the case is open and they have a request into DHL france delivery directly. They should have an update by Monday August 24th no la... [19:21:03] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts rspec hiera config [puppet] - 10https://gerrit.wikimedia.org/r/1328229 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:25:50] RECOVERY - Check unit status of statograph_post on alert1002 is OK: OK: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:30:19] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in dcl hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328230 (https://phabricator.wikimedia.org/T435225) [19:32:23] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in pontoon hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328231 (https://phabricator.wikimedia.org/T435225) [19:32:56] (03CR) 10Cwhite: [C:03+2] beta-logs: enable security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1328222 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [19:33:56] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12242594 (10bking) Looking at console, the host ignored the cookbook's reboot command and remained stuck in the EFI shell. After discussing with @jhat... [19:35:07] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in production hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328232 (https://phabricator.wikimedia.org/T435225) [19:35:20] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328232 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:37:15] !log bking@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host krb1002.eqiad.wmnet with OS bookworm [19:37:22] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12242610 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by bking@cumin2003 for host krb1002.eqiad.wmnet with OS bookworm executed w... [19:38:04] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in dcl hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328230 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:38:10] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in pontoon hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328231 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:41:29] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in production hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328232 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:44:42] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in dev hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328233 (https://phabricator.wikimedia.org/T435225) [19:44:55] (03PS1) 10Cwhite: lookup_options: convert opensearch_output_password to sensitive [puppet] - 10https://gerrit.wikimedia.org/r/1328234 (https://phabricator.wikimedia.org/T350516) [19:46:52] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts in dev hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328233 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:55:27] (03CR) 10Cwhite: [C:03+2] lookup_options: convert opensearch_output_password to sensitive [puppet] - 10https://gerrit.wikimedia.org/r/1328234 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:04:49] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@eqiad in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [20:05:13] (03PS3) 10Cwhite: mediawiki: enable forward of fatal metrics to statsd exporter [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) [20:07:26] (03CR) 10CI reject: [V:04-1] mediawiki: enable forward of fatal metrics to statsd exporter [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) (owner: 10Cwhite) [20:09:46] (03CR) 10Krinkle: Profiler: Fix excimer component regex skipping the last frame in each stack (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328191 (owner: 10SomeRandomDeveloper) [20:09:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@eqiad in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [20:12:24] (03PS4) 10Cwhite: mediawiki: enable forward of fatal metrics to statsd exporter [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) [20:14:34] (03PS1) 10Cwhite: logstash: fix opensearch output template [puppet] - 10https://gerrit.wikimedia.org/r/1328235 (https://phabricator.wikimedia.org/T350516) [20:14:36] (03CR) 10CI reject: [V:04-1] mediawiki: enable forward of fatal metrics to statsd exporter [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) (owner: 10Cwhite) [20:14:56] (03CR) 10Cwhite: mediawiki: enable forward of fatal metrics to statsd exporter (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) (owner: 10Cwhite) [20:15:31] (03CR) 10Cwhite: [C:03+2] logstash: fix opensearch output template [puppet] - 10https://gerrit.wikimedia.org/r/1328235 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:22:00] (03PS2) 10JHathaway: Puppet 8: Replace legacy facts in dev hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328233 (https://phabricator.wikimedia.org/T435225) [20:25:49] !log jhathaway@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: T434750 [20:26:45] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in dev hiera file [puppet] - 10https://gerrit.wikimedia.org/r/1328233 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [20:33:12] !log jhathaway@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: T434750 [20:34:54] !log jhathaway@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: T434750 [20:35:47] (03PS1) 10Lerickson: Testing only: update index bucket. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328237 [20:36:03] !log jhathaway@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: T434750 [20:38:01] (03PS2) 10Lerickson: Testing only: update index bucket. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328237 [20:38:03] (03PS1) 10Cwhite: lookup_options: convert es_exporter password to sensitive [puppet] - 10https://gerrit.wikimedia.org/r/1328238 (https://phabricator.wikimedia.org/T350516) [20:38:15] 06SRE, 06Infrastructure-Foundations, 10Puppet-Infrastructure, 13Patch-For-Review: Fix remaining scoped legacy fact usage - https://phabricator.wikimedia.org/T435225#12242736 (10jhathaway) [20:38:24] (03PS2) 10Cwhite: lookup_options: convert es_exporter password to sensitive [puppet] - 10https://gerrit.wikimedia.org/r/1328238 (https://phabricator.wikimedia.org/T350516) [20:38:26] (03CR) 10CI reject: [V:04-1] lookup_options: convert es_exporter password to sensitive [puppet] - 10https://gerrit.wikimedia.org/r/1328238 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:40:16] (03CR) 10Cwhite: [C:03+2] lookup_options: convert es_exporter password to sensitive [puppet] - 10https://gerrit.wikimedia.org/r/1328238 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:41:40] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [20:42:27] (03PS1) 10Bernard Wang: Enable readinglist for phase0 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 [20:42:52] (03PS2) 10Bernard Wang: Enable readinglist for phase0 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) [20:44:46] (03PS9) 10Bernard Wang: Remove reading list experiment instrumentation [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [20:48:39] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12242773 (10RLazarus) Sorry for the slow response here. Structurally, looks good to me! Some questions: - You clearly thought abo... [20:56:06] (03PS1) 10Cwhite: profile: configure es_exporter class with pki intermediate [puppet] - 10https://gerrit.wikimedia.org/r/1328241 (https://phabricator.wikimedia.org/T350516) [20:57:10] (03CR) 10CI reject: [V:04-1] profile: configure es_exporter class with pki intermediate [puppet] - 10https://gerrit.wikimedia.org/r/1328241 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:59:04] (03PS2) 10Cwhite: profile: configure es_exporter class with pki intermediate [puppet] - 10https://gerrit.wikimedia.org/r/1328241 (https://phabricator.wikimedia.org/T350516) [21:09:59] (03PS3) 10Cwhite: profile: configure es_exporter class with pki intermediate [puppet] - 10https://gerrit.wikimedia.org/r/1328241 (https://phabricator.wikimedia.org/T350516) [21:12:03] (03PS3) 10Bernard Wang: Enable readinglist for phase0 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) [21:12:37] (03CR) 10Cwhite: [C:03+2] "PCC OK: https://puppet-compiler.wmflabs.org/output/1328241/9296/" [puppet] - 10https://gerrit.wikimedia.org/r/1328241 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [21:13:10] (03CR) 10Jdlrobson: [C:04-1] Enable readinglist for phase0 wikis (032 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) (owner: 10Bernard Wang) [21:23:23] (03CR) 10Bking: [C:03+1] rkemper-kafka: add README for pontoon stack [puppet] - 10https://gerrit.wikimedia.org/r/1327500 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [21:31:40] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [21:36:08] (03PS1) 10Cwhite: opensearch: configure curator based on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328244 (https://phabricator.wikimedia.org/T350516) [21:44:22] (03PS4) 10Bernard Wang: Enable readinglist for phase0 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) [21:44:23] (03CR) 10Bernard Wang: Enable readinglist for phase0 wikis (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) (owner: 10Bernard Wang) [21:48:13] (03CR) 10Cwhite: logstash: add security-plugin required fields to output plugin (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:09:56] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:11:12] (03PS1) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [22:11:53] (03CR) 10CI reject: [V:04-1] prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:13:33] (03PS2) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [22:13:51] (03PS3) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [22:14:48] (03PS4) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [22:15:46] (03PS5) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [22:23:39] (03CR) 10Cwhite: "PCC OK: https://puppet-compiler.wmflabs.org/output/1328246/9300/" [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:47:29] (03PS5) 10Lerickson: Update the index S3 bucket for staging only. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328237 (https://phabricator.wikimedia.org/T435674) [22:59:43] (03PS1) 10Cwhite: Revert "beta-logs: enable security plugin" [puppet] - 10https://gerrit.wikimedia.org/r/1328250 [23:00:38] (03CR) 10Cwhite: [C:03+2] Revert "beta-logs: enable security plugin" [puppet] - 10https://gerrit.wikimedia.org/r/1328250 (owner: 10Cwhite) [23:14:56] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [23:18:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:24:23] (03CR) 10Jdlrobson: [C:03+1] Enable readinglist for phase0 wikis (032 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) (owner: 10Bernard Wang) [23:41:37] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1328255 [23:41:37] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1328255 (owner: 10TrainBranchBot) [23:50:26] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1328255 (owner: 10TrainBranchBot)