[00:09:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:10:59] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12249680 (10Bethany) @Eevans , added to my office.wiki page. Thank you [00:12:49] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:13:45] FIRING: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [00:13:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:14:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:16:37] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:16:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:17:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:17:44] (03CR) 10BryanDavis: [C:03+1] Drop irc host from the Beta Cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328711 (https://phabricator.wikimedia.org/T396088) (owner: 10Majavah) [00:20:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:21:55] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:22:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:23:45] FIRING: [2x] WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [00:25:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:26:55] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:30:40] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:31:55] FIRING: [5x] BFDdown: BFD session down between cr1-magru and fe80::6687:8807:11f2:7018 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:35:40] RESOLVED: [5x] BFDdown: BFD session down between cr1-magru and fe80::6687:8807:11f2:7018 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:38:39] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:39:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:40:53] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:41:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.28% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:41:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:43:45] RESOLVED: [2x] WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [00:43:54] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [00:45:46] (03CR) 10Andrea Denisse: [C:03+1] "Nice, LGTM!!" [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [00:47:24] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [00:53:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:57:15] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1321669 (https://phabricator.wikimedia.org/T434690) (owner: 10Cwhite) [00:59:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:02:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:05:42] (03CR) 10Andrea Denisse: [C:03+2] page_form.html: be explicit about not entering private data [software/klaxon] - 10https://gerrit.wikimedia.org/r/1322809 (owner: 10Ssingh) [01:06:50] (03Merged) 10jenkins-bot: page_form.html: be explicit about not entering private data [software/klaxon] - 10https://gerrit.wikimedia.org/r/1322809 (owner: 10Ssingh) [01:06:53] (03CR) 10Andrea Denisse: [V:03+2 C:03+2] page_form.html: be explicit about not entering private data [software/klaxon] - 10https://gerrit.wikimedia.org/r/1322809 (owner: 10Ssingh) [01:07:55] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:10:26] (03PS1) 10TrainBranchBot: Branch commit for wmf/1.47.0-wmf.17 [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1328742 (https://phabricator.wikimedia.org/T430836) [01:10:28] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/1.47.0-wmf.17 [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1328742 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [01:11:16] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1328743 [01:11:16] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1328743 (owner: 10TrainBranchBot) [01:12:55] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:14:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:17:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:20:39] (03Merged) 10jenkins-bot: Branch commit for wmf/1.47.0-wmf.17 [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1328742 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [01:20:47] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1328743 (owner: 10TrainBranchBot) [01:24:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:27:55] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:29:40] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:35:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:38:55] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:40:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:42:49] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:43:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:44:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:48:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 25% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:49:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:51:49] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:53:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.35% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:54:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:59:55] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:00:04] Deploy window Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0200) [02:00:52] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:04:55] RESOLVED: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:07:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [02:08:20] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 28s) [02:08:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:10:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:12:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [02:19:37] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 7/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:21:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:23:53] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:24:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:27:39] FIRING: [2x] TransitBGPDown: Transit BGP session down between cr2-esams and Init7 (2001:1620:1000::85) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [02:28:39] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:28:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:29:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:29:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:29:49] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:30:49] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:34:10] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:36:55] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:37:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:39:10] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:44:10] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:49:10] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:52:37] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 7/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:53:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:54:10] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:55:04] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [02:56:29] PROBLEM - Host db1228 #page is DOWN: PING CRITICAL - Packet loss = 100% [02:56:41] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:56:49] o/ [02:56:54] o/ [02:57:15] checking to see if that's a db1228 problem or a network problem [02:57:17] not even ... pooled? [02:57:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:57:41] it's m5 master according to https://orchestrator.wikimedia.org/web/cluster/alias/m5 [02:57:56] ugh ... I was looking at the wrong sections [02:58:08] m5: Mailman, CXServer, WMCS services, and others [02:58:15] PROBLEM - MariaDB Replica IO: m5 on db1217 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db1228.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db1228.eqiad.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [02:58:32] ugh I *really* don't want to wake someone up again [02:58:35] PROBLEM - MariaDB Replica IO: m5 on db2235 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db1228.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db1228.eqiad.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [02:58:37] PROBLEM - mailman list info ssl expiry on lists1004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [02:59:05] PROBLEM - haproxy failover on dbproxy1029 is CRITICAL: CRITICAL check_failover servers up 1 down 1: https://wikitech.wikimedia.org/wiki/HAProxy [02:59:05] PROBLEM - haproxy failover on dbproxy1027 is CRITICAL: CRITICAL check_failover servers up 1 down 1: https://wikitech.wikimedia.org/wiki/HAProxy [03:00:04] Deploy window Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0300) [03:00:27] RECOVERY - mailman list info ssl expiry on lists1004 is OK: OK - Certificate lists.wikimedia.org will expire on Sat 03 Oct 2026 02:53:55 PM GMT +0000. https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [03:00:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:01:07] trying to figure out if there's any documented procedure we can pick up to resolve this on our own [03:01:16] I'm not finding anything [03:02:00] ... [03:02:01] (03PS1) 10TrainBranchBot: testwikis to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328748 (https://phabricator.wikimedia.org/T430836) [03:02:04] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by mwpresync@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328748 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [03:02:26] if we actually need to do mariadb work I don't think so, but I'm hoping to find out the machine is up and just network-partitioned in a way we can reverse or something [03:02:34] does not look great though [03:02:56] (03Merged) 10jenkins-bot: testwikis to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328748 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [03:03:17] !log mwpresync@deploy1003 Started scap sync-world: testwikis to 1.47.0-wmf.17 refs T430836 [03:03:23] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [03:04:25] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:04:34] swfrench-wmf: on racadm in lclog view, there's a log entry at 2026-08-25 02:55:42 that just says "Internal error has occurred check for additional logs." which doesn't sound like good news [03:04:51] okay I'm starting to think we do need an adult [03:05:05] PROBLEM - MariaDB Replica Lag: m5 on db1217 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 610.08 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:05:22] rzl: ah, I was just about to hop on the management console. yes, +1 - and apparently the best adult is actually oncall right now. [03:05:31] PROBLEM - MariaDB Replica Lag: m5 on db2160 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 634.45 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:05:35] PROBLEM - MariaDB Replica Lag: m5 on db2235 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 640.27 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:05:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:05:51] ughhh I owe this man an unhealthy number of drinks. you want to page him or shall I? [03:05:56] rzl: if you want to open a task with your findings, I can klaxon [03:06:01] deal [03:06:49] opened https://phabricator.wikimedia.org/T435891, NDA for now just due to cortobot defaults [03:08:08] <3 [03:08:15] Hey [03:08:15] klaxon'd [03:08:25] What's up [03:08:29] marostegui: good morning, and sorry =/ [03:08:56] db1228 seems to have crashed (m5 master) [03:09:25] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:09:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:10:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:14:57] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [03:15:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:16:37] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:16:51] RECOVERY - Host db1228 #page is UP: PING OK - Packet loss = 0%, RTA = 0.44 ms [03:17:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:17:35] RECOVERY - MariaDB Replica IO: m5 on db2235 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:18:05] RECOVERY - haproxy failover on dbproxy1029 is OK: OK check_failover servers up 2 down 0: https://wikitech.wikimedia.org/wiki/HAProxy [03:18:15] RECOVERY - MariaDB Replica IO: m5 on db1217 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:19:00] !incidents [03:19:00] 8291 (ACKED) Manual (paged) by Scott French (swfrench@wikimedia.org): Hot handoff: db1228 (m5 master) is down [03:19:00] 8290 (RESOLVED) Host db1228 (paged) [03:19:00] 8289 (RESOLVED) [3x] ATSBackendErrorsHigh cache_upload sre (swift.discovery.wmnet) [03:19:01] 8288 (RESOLVED) Host db2236 (paged) [03:19:05] RECOVERY - MariaDB Replica Lag: m5 on db1217 is OK: OK slave_sql_lag Replication lag: 0.10 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:19:09] !resolve 8291 [03:19:09] 8291 (RESOLVED) Manual (paged) by Scott French (swfrench@wikimedia.org): Hot handoff: db1228 (m5 master) is down [03:19:25] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:19:29] RECOVERY - MariaDB Replica Lag: m5 on db2160 is OK: OK slave_sql_lag Replication lag: 0.22 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:19:35] RECOVERY - MariaDB Replica Lag: m5 on db2235 is OK: OK slave_sql_lag Replication lag: 0.02 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:19:37] FIRING: GnmiInterfaceCountersDrop: ... [03:19:38] lsw1-e8-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-e8-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [03:22:27] !incidents [03:22:27] 8291 (RESOLVED) Manual (paged) by Scott French (swfrench@wikimedia.org): Hot handoff: db1228 (m5 master) is down [03:22:28] 8290 (RESOLVED) Host db1228 (paged) [03:22:28] 8289 (RESOLVED) [3x] ATSBackendErrorsHigh cache_upload sre (swift.discovery.wmnet) [03:22:28] 8288 (RESOLVED) Host db2236 (paged) [03:23:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:26:51] (03PS1) 10Marostegui: instances.yaml: Remove db1243 [puppet] - 10https://gerrit.wikimedia.org/r/1328749 (https://phabricator.wikimedia.org/T435892) [03:27:49] (03CR) 10Marostegui: [C:03+2] instances.yaml: Remove db1243 [puppet] - 10https://gerrit.wikimedia.org/r/1328749 (https://phabricator.wikimedia.org/T435892) (owner: 10Marostegui) [03:27:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:27:56] (03CR) 10Marostegui: [C:03+2] check_private_data_report: Add db1269 [puppet] - 10https://gerrit.wikimedia.org/r/1328566 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [03:29:36] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1243].eqiad.wmnet with reason: Cloning [03:29:53] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove db1243 from dbctl T435892', diff saved to https://phabricator.wikimedia.org/P96240 and previous config saved to /var/cache/conftool/dbconfig/20260825-032952-marostegui.json [03:29:57] T435892: Move db1243 to m5 - https://phabricator.wikimedia.org/T435892 [03:30:35] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:31:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:32:07] (03PS1) 10Marostegui: mariadb: Move db1243 to m5 [puppet] - 10https://gerrit.wikimedia.org/r/1328751 (https://phabricator.wikimedia.org/T435892) [03:33:06] (03PS2) 10Marostegui: mariadb: Move db1243 to m5 [puppet] - 10https://gerrit.wikimedia.org/r/1328751 (https://phabricator.wikimedia.org/T435892) [03:34:37] (03CR) 10Marostegui: [C:03+2] mariadb: Move db1243 to m5 [puppet] - 10https://gerrit.wikimedia.org/r/1328751 (https://phabricator.wikimedia.org/T435892) (owner: 10Marostegui) [03:36:35] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:38:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:38:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:40:05] PROBLEM - haproxy failover on dbproxy1029 is CRITICAL: CRITICAL check_failover servers up 1 down 1: https://wikitech.wikimedia.org/wiki/HAProxy [03:40:21] !log mwpresync@deploy1003 Finished scap sync-world: testwikis to 1.47.0-wmf.17 refs T430836 (duration: 37m 04s) [03:40:26] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [03:40:55] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:42:55] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:47:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:48:40] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:52:55] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:53:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:54:57] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr2-esams:xe-0/1/6 (Transit: Init7 (N/A)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-esams:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [03:57:39] RESOLVED: [2x] TransitBGPDown: Transit BGP session down between cr2-esams and Init7 (2001:1620:1000::85) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [03:58:07] (03PS1) 10Tim Starling: Grant produnto-update to groups that have editprotected [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328752 (https://phabricator.wikimedia.org/T421436) [04:00:04] Deploy window Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0400) [04:00:05] RECOVERY - haproxy failover on dbproxy1029 is OK: OK check_failover servers up 2 down 0: https://wikitech.wikimedia.org/wiki/HAProxy [04:00:05] RECOVERY - haproxy failover on dbproxy1027 is OK: OK check_failover servers up 2 down 0: https://wikitech.wikimedia.org/wiki/HAProxy [04:02:30] !log mwpresync@deploy1003 Pruned MediaWiki: 1.47.0-wmf.14 (duration: 02m 24s) [04:03:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:06:47] PROBLEM - Backup freshness on backup1014 is CRITICAL: All failures: 1 (krb1003), Fresh: 140 jobs https://wikitech.wikimedia.org/wiki/Bacula%23Monitoring [04:07:55] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:14:37] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:15:37] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:19:55] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:24:55] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:25:40] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:26:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:29:55] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:31:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:32:39] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 7/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:33:39] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:34:55] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:41:39] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:42:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:43:39] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:43:49] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:47:37] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1217,1228,1243].eqiad.wmnet with reason: m5 master switch T432967 [04:47:42] T432967: Switchover m5 master db1164 -> db1228 - https://phabricator.wikimedia.org/T432967 [04:51:49] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:52:07] (03PS1) 10Marostegui: mariadb: Promote db1243 to m5 master [puppet] - 10https://gerrit.wikimedia.org/r/1328755 (https://phabricator.wikimedia.org/T435895) [04:52:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:53:16] (03CR) 10Marostegui: [C:03+2] mariadb: Promote db1243 to m5 master [puppet] - 10https://gerrit.wikimedia.org/r/1328755 (https://phabricator.wikimedia.org/T435895) (owner: 10Marostegui) [04:53:49] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:53:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:56:51] !log Failover m5 from db1228 to db1243 - T435895 [04:56:57] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [04:56:57] T435895: Switchover m5 master db1228 -> db1243 - https://phabricator.wikimedia.org/T435895 [04:57:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:01:46] (03PS1) 10Marostegui: db1228: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1328756 (https://phabricator.wikimedia.org/T435892) [05:03:24] (03CR) 10Marostegui: [C:03+2] db1228: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1328756 (https://phabricator.wikimedia.org/T435892) (owner: 10Marostegui) [05:06:48] RECOVERY - Backup freshness on backup1014 is OK: Fresh: 141 jobs https://wikitech.wikimedia.org/wiki/Bacula%23Monitoring [05:08:41] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12249982 (10Marostegui) @Jhancock.wm @wiki_willy is there a way to search hosts with config E with Intel Xeon Gold 5317 in netbox? I am not finding how to. [05:14:05] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12249986 (10andrea.denisse) [05:14:27] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12249990 (10Marostegui) >>! In T435271#12233861, @Jhancock.wm wrote: > I have some updates and a quick ask. > ask first, is it okay for me to do some firmware updates on this server right now? > > u... [05:15:40] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:17:40] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:18:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:18:57] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12249994 (10andrea.denisse) [05:22:37] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12249997 (10Cyberpower678) >>! In T435743#12246795, @TheDJ wrote: > Maybe @cyberpower678 can help find the right contacts I’ve relayed this to the engineers. [05:22:50] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:23:10] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:23:50] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:27:44] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:28:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:28:42] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:34:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:34:55] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:37:56] PROBLEM - Blazegraph Port for wdqs-blazegraph on wdqs1021 is CRITICAL: connect to address 127.0.0.1 and port 9999: Connection refused https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [05:38:56] RECOVERY - Blazegraph Port for wdqs-blazegraph on wdqs1021 is OK: TCP OK - 0.000 second response time on 127.0.0.1 port 9999 https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [05:39:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:39:50] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:40:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:40:55] FIRING: [3x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:45:40] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:45:50] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:50:40] RESOLVED: [9x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:50:42] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12250012 (10andrea.denisse) [05:55:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:55:55] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0600) [06:00:04] marostegui, Amir1, and federico3: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for Primary database switchover deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0600). [06:00:40] RESOLVED: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:00:55] FIRING: [9x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:05:40] FIRING: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:06:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [06:06:38] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:07:42] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:10:40] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:11:15] RESOLVED: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [06:15:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:20:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:20:42] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:21:38] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:23:52] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:24:40] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:24:52] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:25:38] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:25:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:32:02] jouncebot: refresh [06:32:03] I refreshed my knowledge about deployments. [06:32:06] jouncebot: nowandnext [06:32:06] For the next 0 hour(s) and 27 minute(s): MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0600) [06:32:06] In 0 hour(s) and 27 minute(s): UTC morning backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0700) [06:32:29] hmm [06:33:11] jouncebot is broken somehow [06:33:42] the second message should have pointed to the train [06:35:01] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12250028 (10andrea.denisse) [06:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 9.708% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:36:28] 06SRE, 10Wikimedia-Mailing-lists: Postorius: "Message could not be found" error preventing mailing list moderation - https://phabricator.wikimedia.org/T435893#12250030 (10JJMC89) I discarded the messages that were pending moderation. [06:36:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:36:52] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:37:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [06:37:50] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:40:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:41:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 0% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:41:52] !ack [06:41:53] 8292 (ACKED) PHPFPMTooBusy sre (mw-api-ext main eqiad) [06:44:42] (03PS1) 10Muehlenhoff: Remove obsolete Cumin alias for parsoid-testing [puppet] - 10https://gerrit.wikimedia.org/r/1329157 [06:45:38] (03PS1) 10Kevin Bazira: ml-services: deploy qwen36-27b to use qwen3_coder parser for tool calling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329158 (https://phabricator.wikimedia.org/T434274) [06:46:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 3.584% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:46:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 3.584% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:46:55] FIRING: [3x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:47:15] RESOLVED: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [06:47:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:51:55] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:52:40] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:55:04] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [06:56:55] RESOLVED: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:58:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:00:04] Amir1, urbanecm, and awight: Time to do the UTC morning backport window deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0700). [07:00:04] No Gerrit patches in the queue for this window AFAICS. [07:00:32] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12250063 (10andrea.denisse) >>! In T435271#12229855, @CWilliams-WMF wrote: > @Marostegui for reference db2207 had a switchover (T434565) recently - 11 August - for the kernel update. Making queries w... [07:01:45] (03PS2) 10Muehlenhoff: Remove obsolete Cumin alias for parsoid-testing [puppet] - 10https://gerrit.wikimedia.org/r/1329157 [07:01:55] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:03:40] FIRING: [9x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:04:50] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply [07:05:12] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply [07:06:55] RESOLVED: [9x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:07:01] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [07:07:47] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [07:08:35] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [07:08:41] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply [07:09:10] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply [07:09:18] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply [07:09:28] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [07:09:48] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply [07:10:27] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply [07:10:47] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12250077 (10Marostegui) >>! In T435271#12250063, @andrea.denisse wrote: >>>! In T435271#12229855, @CWilliams-WMF wrote: >> @Marostegui for reference db2207 had a switchover (T434565) recently - 11 Aug... [07:10:51] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply [07:11:00] 10ops-eqiad, 06SRE, 06DC-Ops, 10Kafka-Infrastructure, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12250080 (10brouberol) [07:11:02] 10ops-eqiad, 06SRE, 06DC-Ops, 10Kafka-Infrastructure, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12250081 (10brouberol) 05Open→03In progress [07:11:55] FIRING: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:12:01] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply [07:12:35] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply [07:12:40] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply [07:13:04] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply [07:13:18] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply [07:13:40] FIRING: [11x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:13:51] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply [07:13:59] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply [07:14:21] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply [07:14:47] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply [07:14:57] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:15:30] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply [07:15:50] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply [07:16:22] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply [07:16:55] FIRING: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:17:32] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply [07:17:47] (03CR) 10Bartosz Wójtowicz: [C:03+1] ml-services: deploy qwen36-27b to use qwen3_coder parser for tool calling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329158 (https://phabricator.wikimedia.org/T434274) (owner: 10Kevin Bazira) [07:17:54] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply [07:18:02] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply [07:18:36] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12250099 (10andrea.denisse) >>! In T435271#12250077, @Marostegui wrote: >>>! In T435271#12250063, @andrea.denisse wrote: >>>>! In T435271#12229855, @CWilliams-WMF wrote: >>> @Marostegui for reference... [07:18:39] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply [07:18:40] RESOLVED: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:19:11] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply [07:19:37] FIRING: GnmiInterfaceCountersDrop: ... [07:19:38] lsw1-e8-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-e8-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [07:19:50] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply [07:21:55] FIRING: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:23:40] FIRING: [11x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:24:47] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply [07:24:56] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply [07:25:07] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply [07:26:10] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply [07:26:11] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply [07:26:55] FIRING: [11x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:27:15] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply [07:27:16] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply [07:28:28] (03CR) 10Filippo Giunchedi: [C:03+2] Revert "site: get cloudvirts ready for E4 -> C8 move" [puppet] - 10https://gerrit.wikimedia.org/r/1328176 (https://phabricator.wikimedia.org/T431682) (owner: 10Filippo Giunchedi) [07:28:39] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply [07:28:40] FIRING: [12x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:28:41] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply [07:28:57] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12250126 (10MoritzMuehlenhoff) This needs to wait until we consistently have 10G NICs in eqiad/codfw (eqiad will be ready once the servers ordered... [07:29:08] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply [07:29:09] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply [07:30:15] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply [07:30:16] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply [07:31:26] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply [07:31:55] RESOLVED: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:32:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by taavi@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328711 (https://phabricator.wikimedia.org/T396088) (owner: 10Majavah) [07:32:03] (03CR) 10Nikerabbit: [V:03+2] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1328579 (owner: 10L10n-bot) [07:33:25] (03Merged) 10jenkins-bot: Drop irc host from the Beta Cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328711 (https://phabricator.wikimedia.org/T396088) (owner: 10Majavah) [07:35:08] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12250165 (10MoritzMuehlenhoff) >>! In T435354#12248640, @jhathaway wrote: > I booted the box into SystemRescue running 12.02 running Linux kernel 6.12.4... [07:36:20] (03PS1) 10Brouberol: blunderbuss: use the kerberos servers defined in the global values [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329165 (https://phabricator.wikimedia.org/T435903) [07:36:22] (03PS1) 10Brouberol: dse-k8s-eqiad: define a meta helmfile for kerberized apps [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329166 (https://phabricator.wikimedia.org/T435903) [07:36:50] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:37:26] (03CR) 10Muehlenhoff: [C:03+1] Record that a principal was created for user mkrolik-wmf [puppet] - 10https://gerrit.wikimedia.org/r/1328736 (https://phabricator.wikimedia.org/T434877) (owner: 10Eevans) [07:37:52] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:39:01] (03CR) 10CI reject: [V:04-1] dse-k8s-eqiad: define a meta helmfile for kerberized apps [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329166 (https://phabricator.wikimedia.org/T435903) (owner: 10Brouberol) [07:39:40] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:40:22] (03PS2) 10Brouberol: dse-k8s-eqiad: define a meta helmfile for kerberized apps [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329166 (https://phabricator.wikimedia.org/T435903) [07:41:55] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:42:42] (03CR) 10CI reject: [V:04-1] dse-k8s-eqiad: define a meta helmfile for kerberized apps [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329166 (https://phabricator.wikimedia.org/T435903) (owner: 10Brouberol) [07:43:25] (03CR) 10Slyngshede: [C:03+1] services: make urldownloader probe paging [puppet] - 10https://gerrit.wikimedia.org/r/1328218 (https://phabricator.wikimedia.org/T435648) (owner: 10Ssingh) [07:46:35] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudvirt1049.eqiad.wmnet [07:48:39] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1901.eqiad.wmnet with reason: Cloning [07:50:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:51:46] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12250219 (10cmooney) ` 2026-08-25 06:46:48 GMT - Due to unforeseen circumstances, the Lumen NOC has advised that this Urgent Maintenance Network Event could not be completed during the s... [07:52:19] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [07:52:55] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:53:15] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [07:55:40] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:57:55] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:59:19] 06SRE, 10Thumbor, 06Traffic, 07affects-Kiwix-and-openZIM: MWoffliner scrapes slowed down by Thumbor failure throttling 429s - https://phabricator.wikimedia.org/T304814#12250223 (10SLyngshede-WMF) The infrastructure around thumbnails have been significantly reworked since 2022. While the 429 remains in plac... [07:59:45] I am rolling the train [08:00:04] (03PS1) 10Muehlenhoff: Record LDAP access for noahmvf [puppet] - 10https://gerrit.wikimedia.org/r/1329209 [08:00:05] hashar and andre: Deploy window MediaWiki train - Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0800) [08:00:10] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329210 (https://phabricator.wikimedia.org/T430836) [08:00:13] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by hashar@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329210 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [08:00:16] o/ [08:00:40] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:00:42] 🌊 [08:01:11] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329210 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [08:02:11] (03CR) 10Muehlenhoff: [C:03+2] Record LDAP access for noahmvf [puppet] - 10https://gerrit.wikimedia.org/r/1329209 (owner: 10Muehlenhoff) [08:02:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:05:12] (03CR) 10Jelto: [C:03+2] wikimedia.org: add etherpad-next and lower TTL [dns] - 10https://gerrit.wikimedia.org/r/1328350 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [08:05:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:07:00] !log jelto@dns1004 START - running authdns-update [08:07:55] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:08:10] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12250256 (10AlexisJazz) [08:09:20] !log jelto@dns1004 END - running authdns-update [08:14:54] !log hashar@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.17 refs T430836 [08:15:00] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [08:15:53] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12250269 (10Fabfur) a:03Fabfur [08:17:28] !log filippo@cumin1003 END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1049.eqiad.wmnet [08:17:56] !log put cloudvirts back in service T431682 [08:17:59] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:18:00] T431682: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682 [08:19:32] (03PS1) 10Marostegui: db1276: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1329220 (https://phabricator.wikimedia.org/T407942) [08:19:56] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [08:19:56] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [08:20:56] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [08:20:56] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [08:21:33] (03CR) 10Marostegui: [C:03+2] db1276: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1329220 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [08:22:55] !log filippo@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1049.eqiad.wmnet with OS trixie [08:23:13] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12250304 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt1049.eqiad.wmnet with OS trixie [08:23:17] (03PS1) 10Marostegui: instances.yaml: Add db1276 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1329221 (https://phabricator.wikimedia.org/T407942) [08:24:05] !log filippo@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1051.eqiad.wmnet with OS trixie [08:27:32] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1276 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1329221 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [08:28:17] (03CR) 10Tiziano Fogli: [C:03+2] prometheus/jobunavailable: double the alert firing time [alerts] - 10https://gerrit.wikimedia.org/r/1328587 (owner: 10Tiziano Fogli) [08:29:23] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1276 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96242 and previous config saved to /var/cache/conftool/dbconfig/20260825-082922-marostegui.json [08:29:24] !log installing Expat security updates [08:29:29] T407942: Productionize db12[65-90] - https://phabricator.wikimedia.org/T407942 [08:29:31] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:29:43] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1276: Pool back [08:31:46] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [08:34:24] 14SRE-Sprint-Week-Sustainability-March2023, 06Traffic, 07Sustainability (Incident Followup): Rate limiting for hotlinked images - https://phabricator.wikimedia.org/T317799#12250331 (10SLyngshede-WMF) @RLazarus / @CDanis do we want to keep this task open, or do we now view this as "mostly under control"? [08:34:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [08:37:45] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Manage X-Forwarded-For header in HAProxy logic [puppet] - 10https://gerrit.wikimedia.org/r/1328642 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [08:38:41] (03CR) 10Atsuko: [C:03+1] blunderbuss: use the kerberos servers defined in the global values [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329165 (https://phabricator.wikimedia.org/T435903) (owner: 10Brouberol) [08:38:45] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [08:39:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:39:55] !log filippo@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage [08:40:40] !log filippo@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage [08:40:40] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:40:58] (03CR) 10Atsuko: [C:03+1] "I'm okay with this but linter is complaining" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329166 (https://phabricator.wikimedia.org/T435903) (owner: 10Brouberol) [08:41:14] (03PS1) 10Mszwarc: enwikivoyage: Fix the applychangetags assignment to '*' [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329222 [08:42:27] I'd like to deploy the above config patch as soon as possible. hashar, andre: what's the train deployment status? [08:43:09] (03CR) 10Atsuko: [C:03+2] provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1328576 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [08:43:44] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1051.eqiad.wmnet with reason: host reimage [08:43:54] !log atsuko@dns1004 START - running authdns-update [08:43:57] (03PS2) 10Mszwarc: enwikivoyage: Fix the applychangetags assignment to '*' [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329222 (https://phabricator.wikimedia.org/T435907) [08:44:25] !log atsuko@dns1004 START - running authdns-update [08:44:49] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [08:45:18] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Manage X-Forwarded-For header in HAProxy logic [puppet] - 10https://gerrit.wikimedia.org/r/1328642 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [08:46:22] (03CR) 10Dreamy Jazz: [C:04-1] "Looks like moved to the wrong group override (I.e. wrong wiki)?" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329222 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [08:46:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:46:43] !log atsuko@dns1004 END - running authdns-update [08:46:52] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [08:47:08] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1049.eqiad.wmnet with reason: host reimage [08:47:31] (03PS3) 10Mszwarc: enwikivoyage: Fix the applychangetags assignment to '*' [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329222 (https://phabricator.wikimedia.org/T435907) [08:48:14] Msz2001: I rolled the train yes [08:48:16] (03CR) 10Mszwarc: "Oh, you're right. Fixed, thanks!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329222 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [08:48:39] So, I'll proceed with this deployment in a few minutes, is that okay? [08:49:01] Msz2001: it is quiet apparently, so I guess you can deploy yes! [08:49:06] Msz2001, train is done per https://sal.toolforge.org/log/YLT8N6ABDAZZyZXnlIoK and logs look okay to me but in the end hashar decides :) [08:49:11] Thanks! [08:49:51] the one you fix is for an error that is currently happening on wikivoyage isn't it? (I see `Call to a member function isNamed() on null` in the UserChangeHooks [08:49:54] 06SRE, 06Infrastructure-Foundations, 10netops: Missing series for BGP session_state from eqiad CRs since upgrade to 23.4R2-S8.7 - https://phabricator.wikimedia.org/T435909 (10cmooney) 03NEW p:05Triage→03Medium [08:49:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:50:05] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329222 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [08:50:13] (03CR) 10Hashar: [C:03+1] enwikivoyage: Fix the applychangetags assignment to '*' [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329222 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [08:50:21] hashar: it's a little worse than that, I'd see the task [08:50:29] (03CR) 10Brouberol: [C:03+2] blunderbuss: use the kerberos servers defined in the global values [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329165 (https://phabricator.wikimedia.org/T435903) (owner: 10Brouberol) [08:50:39] (03PS1) 10Slyngshede: R:cache::upload remove unused profile netconsole [puppet] - 10https://gerrit.wikimedia.org/r/1329224 (https://phabricator.wikimedia.org/T427646) [08:51:06] (03CR) 10Atsuko: [C:03+2] Add analytics-experiment system user and groups [puppet] - 10https://gerrit.wikimedia.org/r/1328582 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [08:51:16] ack [08:51:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:51:41] (03Merged) 10jenkins-bot: enwikivoyage: Fix the applychangetags assignment to '*' [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329222 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [08:51:54] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply [08:52:02] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1329222|enwikivoyage: Fix the applychangetags assignment to '*' (T435907)]] [08:52:11] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329224 (https://phabricator.wikimedia.org/T427646) (owner: 10Slyngshede) [08:52:11] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply [08:52:20] 06SRE, 06Infrastructure-Foundations, 10netops: Missing series for BGP session_state from eqiad CRs since upgrade to 23.4R2-S8.7 - https://phabricator.wikimedia.org/T435909#12250416 (10cmooney) [08:53:07] 06SRE, 06Infrastructure-Foundations, 10netops: Missing series for BGP session_state from eqiad CRs since upgrade to 23.4R2-S8.7 - https://phabricator.wikimedia.org/T435909#12250417 (10cmooney) [08:53:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:53:54] (03PS8) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [08:54:12] (03CR) 10Filippo Giunchedi: "Approach/idea LGTM, though please don't bundle together unrelated changes (e.g. changing SPDX license check requirements)" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [08:54:21] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply [08:54:22] !log mszwarc@deploy1003 mszwarc: Backport for [[gerrit:1329222|enwikivoyage: Fix the applychangetags assignment to '*' (T435907)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [08:54:48] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply [08:54:55] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:55:11] !log mszwarc@deploy1003 mszwarc: Continuing with deployment [08:55:20] (03CR) 10Hashar: "Good point, I'll split it, and probably add another change which introduces the rspec tests. Thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [08:55:27] Msz2001: see A09's comment, might be worth doing a full review if you've got time [08:55:52] (03CR) 10Jelto: [V:03+1 C:03+2] "the correct cookbook command is" [puppet] - 10https://gerrit.wikimedia.org/r/1328348 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [08:56:13] (03CR) 10Cmelo: [C:03+2] Stop setting wgCampaignEventsEnableWorklists [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326400 (https://phabricator.wikimedia.org/T429510) (owner: 10Daimona Eaytoy) [08:56:36] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw [08:56:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:56:57] I can't see any other issues though with a quick ctrl f [08:57:40] (03PS9) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [08:59:16] (03PS10) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [08:59:31] RhinosF1: Thanks for looking, me neither [08:59:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:00:37] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329222|enwikivoyage: Fix the applychangetags assignment to '*' (T435907)]] (duration: 08m 34s) [09:00:45] Finished deploying [09:01:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:04:00] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:04:07] (03PS11) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [09:04:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:06:00] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:07:34] (03CR) 10Slyngshede: "This removes the sysfs package. That should be fine as all the parameters are already set as required." [puppet] - 10https://gerrit.wikimedia.org/r/1329224 (https://phabricator.wikimedia.org/T427646) (owner: 10Slyngshede) [09:08:54] jouncebot: nowandnext [09:08:54] For the next 0 hour(s) and 51 minute(s): MediaWiki train - Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T0800) [09:08:54] In 0 hour(s) and 51 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1000) [09:10:45] (03CR) 10Majavah: "I believe this is finally ready for review." [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [09:12:50] (03PS3) 10Atsuko: Grant sudo privileges for the analytics-experiment-users group [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) [09:14:47] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1276: Pool back [09:20:13] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [09:20:25] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db-test1001.eqiad.wmnet [09:21:39] (03CR) 10Jelto: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9306/co" [puppet] - 10https://gerrit.wikimedia.org/r/1328349 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [09:23:58] 06SRE, 06Traffic, 13Patch-For-Review: netconsole being used for cache hosts? - https://phabricator.wikimedia.org/T427646#12250507 (10MoritzMuehlenhoff) >>! In T427646#11967900, @ssingh wrote: > I meant we set `profile::netconsole::client::ensure: absent` in `hieradata/role/common/cache/upload.yaml` so it sho... [09:24:11] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1329224 (https://phabricator.wikimedia.org/T427646) (owner: 10Slyngshede) [09:24:21] (03CR) 10Hnowlan: [C:03+1] logstash: add security-plugin required fields to output plugin (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [09:24:48] (03CR) 10Hnowlan: [C:03+1] prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [09:24:52] (03PS1) 10Hashar: spdx: ignore `.rspec` file [puppet] - 10https://gerrit.wikimedia.org/r/1329227 [09:25:26] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [09:26:52] (03PS12) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [09:26:52] (03PS1) 10Majavah: P:wmcs::novaproxy: Stop reading proxy data from Redis in Nginx [puppet] - 10https://gerrit.wikimedia.org/r/1329229 (https://phabricator.wikimedia.org/T429930) [09:26:55] (03PS1) 10Majavah: P:wmcs::novaproxy: Stop writing data to Redis [puppet] - 10https://gerrit.wikimedia.org/r/1329230 (https://phabricator.wikimedia.org/T429930) [09:26:57] (03PS1) 10Majavah: P:wmcs::novaproxy: Undeploy Redis instance [puppet] - 10https://gerrit.wikimedia.org/r/1329231 (https://phabricator.wikimedia.org/T429930) [09:29:04] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:29:56] (03CR) 10Muehlenhoff: "Can you update the patch to clarify what these will be used for? It's not obvious from the task what the eventual goal is. Will there even" [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [09:30:04] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:30:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:30:42] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [09:30:47] fceratto@cumin1003 decommission (PID 3461476) is awaiting input [09:30:47] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Stop writing data to Redis [puppet] - 10https://gerrit.wikimedia.org/r/1329230 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [09:31:30] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Undeploy Redis instance [puppet] - 10https://gerrit.wikimedia.org/r/1329231 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [09:32:08] (03PS2) 10Majavah: P:wmcs::novaproxy: Stop writing data to Redis [puppet] - 10https://gerrit.wikimedia.org/r/1329230 (https://phabricator.wikimedia.org/T429930) [09:32:08] (03PS2) 10Majavah: P:wmcs::novaproxy: Undeploy Redis instance [puppet] - 10https://gerrit.wikimedia.org/r/1329231 (https://phabricator.wikimedia.org/T429930) [09:32:38] (03PS1) 10JavierMonton: stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329232 (https://phabricator.wikimedia.org/T431555) [09:32:39] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Metrics: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855#12250560 (10cmooney) [09:34:34] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy qwen36-27b to use qwen3_coder parser for tool calling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329158 (https://phabricator.wikimedia.org/T434274) (owner: 10Kevin Bazira) [09:34:57] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [09:35:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:35:17] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [09:35:17] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:35:18] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db-test1001.eqiad.wmnet [09:35:57] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Metrics: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855#12250572 (10cmooney) [09:36:01] (03PS1) 10Hashar: zookeeper: add spec tests [puppet] - 10https://gerrit.wikimedia.org/r/1329228 (https://phabricator.wikimedia.org/T435503) [09:36:02] (03CR) 10Hashar: "I have changed the default for `zookeeper::hosts` from NumericValue 1 to String '1', else the rspec fails to compile due to the unsupporte" [puppet] - 10https://gerrit.wikimedia.org/r/1329228 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [09:36:17] cmooney@cumin1003 netbox (PID 3467393) is awaiting input [09:36:25] (03PS1) 10Marostegui: instances.yaml: Add db1176,db2230 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1329233 (https://phabricator.wikimedia.org/T427059) [09:36:31] !log cmooney@cumin1003 END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) [09:36:42] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [09:36:47] (03Merged) 10jenkins-bot: ml-services: deploy qwen36-27b to use qwen3_coder parser for tool calling [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329158 (https://phabricator.wikimedia.org/T434274) (owner: 10Kevin Bazira) [09:38:25] fceratto@cumin1003 decommission (PID 3461476) is awaiting input [09:38:53] (03PS1) 10Federico Ceratto: site.pp, db-test1001.yaml: Decommission db-test1001 [puppet] - 10https://gerrit.wikimedia.org/r/1329234 (https://phabricator.wikimedia.org/T435912) [09:39:12] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:39:23] 10ops-eqiad, 06SRE, 06DC-Ops: Recycle old MPC 3D 16x10G line cards in cr2-eqiad - https://phabricator.wikimedia.org/T435588#12250586 (10VRiley-WMF) 05Open→03Resolved [09:39:41] (03PS1) 10Gkyziridis: ml-services: Deploy latest langid model version under llm ns on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329236 (https://phabricator.wikimedia.org/T429675) [09:39:56] jelto@cumin1003 migrate-service-ipip (PID 3441647) is awaiting input [09:40:13] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [09:40:22] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1176,db2230 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1329233 (https://phabricator.wikimedia.org/T427059) (owner: 10Marostegui) [09:41:28] RECOVERY - Host ganeti3005 is UP: PING OK - Packet loss = 0%, RTA = 77.28 ms [09:41:44] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2901.codfw.wmnet with reason: Cloning [09:42:04] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:42:26] PROBLEM - ganeti-wconfd running on ganeti3005 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 110 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [09:42:47] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [09:42:48] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw [09:43:04] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:43:10] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12250594 (10Fabfur) I've noticed that archiving a new page from Commons, Web Archive completes the task successfully: it fetches the page and the image (and thumbnails) as expe... [09:43:14] ^ ganeti3005 seems to be smart hards succeeded in fixing the backplane, the followup is fine since the server is drained from active duty [09:43:46] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1176 and db2230 to dbctl T427059', diff saved to https://phabricator.wikimedia.org/P96248 and previous config saved to /var/cache/conftool/dbconfig/20260825-094345-marostegui.json [09:43:51] T427059: Discussion for adding support of test-s4 in dbctl - https://phabricator.wikimedia.org/T427059 [09:43:53] RESOLVED: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:44:27] (03PS1) 10Gkyziridis: ml-services: Deploy latest langid model version under llm ns on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329237 (https://phabricator.wikimedia.org/T429675) [09:45:36] (03PS1) 10Federico Ceratto: site.pp,db2903.yaml,preseed.yaml: add db2903 [puppet] - 10https://gerrit.wikimedia.org/r/1329239 (https://phabricator.wikimedia.org/T435059) [09:46:00] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:46:08] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add test-s4 to dbctl T427059', diff saved to https://phabricator.wikimedia.org/P96249 and previous config saved to /var/cache/conftool/dbconfig/20260825-094607-marostegui.json [09:46:40] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:46:55] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:47:48] (03CR) 10Marostegui: [C:04-1] site.pp,db2903.yaml,preseed.yaml: add db2903 (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329239 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [09:48:00] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:49:34] (03CR) 10Filippo Giunchedi: [C:03+1] "LGTM, nice" [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [09:50:37] (03CR) 10Filippo Giunchedi: [C:03+1] spdx: ignore `.rspec` file [puppet] - 10https://gerrit.wikimedia.org/r/1329227 (owner: 10Hashar) [09:50:56] (03CR) 10Filippo Giunchedi: [C:03+1] zookeeper: add spec tests [puppet] - 10https://gerrit.wikimedia.org/r/1329228 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [09:50:59] (03PS1) 10Mvolz: zotero: update version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329240 (https://phabricator.wikimedia.org/T435179) [09:51:05] (03PS1) 10Gkyziridis: ml-services: Deploy latest articlequality model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329241 (https://phabricator.wikimedia.org/T429675) [09:51:40] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:52:00] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:52:42] FIRING: JobUnavailable: Reduced availability for job mysql-test in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:53:46] (03PS1) 10Gkyziridis: ml-services: Deploy latest articlequality model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329243 (https://phabricator.wikimedia.org/T429675) [09:54:00] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [09:54:38] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1051.eqiad.wmnet with OS trixie [09:58:14] !log filippo@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 [09:58:43] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1049.eqiad.wmnet with OS trixie [09:58:45] !log filippo@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 [09:58:59] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12250657 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1049.eqiad.wmnet with OS trixie completed: - cl... [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1000) [10:02:22] !log filippo@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 [10:02:36] (03PS2) 10Federico Ceratto: site.pp,db2903.yaml,preseed.yaml: add db2903 [puppet] - 10https://gerrit.wikimedia.org/r/1329239 (https://phabricator.wikimedia.org/T435059) [10:02:51] !log filippo@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 [10:02:59] (03CR) 10Federico Ceratto: site.pp,db2903.yaml,preseed.yaml: add db2903 (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329239 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [10:06:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:08:59] !log filippo@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1051 [10:09:38] !log filippo@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1051 [10:11:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:12:03] (03CR) 10MSantos: [C:03+1] Enable Produnto on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324965 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [10:12:05] !log upgrading apus eqiad cluster to Reef 18.2.8 [10:12:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:12:44] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudvirt1054.eqiad.wmnet [10:15:17] (03CR) 10Hnowlan: [C:03+1] profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [10:19:02] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1054.eqiad.wmnet [10:21:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:22:20] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs1019 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [10:22:43] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudvirt1055.eqiad.wmnet [10:23:01] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudvirt1056.eqiad.wmnet [10:23:20] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudvirt1057.eqiad.wmnet [10:26:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:27:32] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs1020 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [10:29:02] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1055.eqiad.wmnet [10:29:27] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1057.eqiad.wmnet [10:29:29] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1056.eqiad.wmnet [10:29:56] (03PS1) 10Santiago Faci: Added a property to define instake URL for beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329246 (https://phabricator.wikimedia.org/T433711) [10:30:47] (03CR) 10CI reject: [V:04-1] Added a property to define instake URL for beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329246 (https://phabricator.wikimedia.org/T433711) (owner: 10Santiago Faci) [10:31:32] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest articlequality model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329243 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [10:32:16] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12250780 (10fgiunchedi) 05In progress→03Resolved This is done, hosts have been relocated and are back in service [10:35:16] (03PS1) 10Muehlenhoff: proton: Bump image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329249 [10:36:00] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12250792 (10VRiley-WMF) 05Open→03Resolved This has been resolved [10:37:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:37:55] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:40:13] (03PS2) 10JavierMonton: stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329232 (https://phabricator.wikimedia.org/T431555) [10:40:38] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12250815 (10fgiunchedi) 05Resolved→03Open Unfortunately it hasn't, we still have a few rack moves to do [10:40:47] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12250818 (10fgiunchedi) a:05VRiley-WMF→03fgiunchedi [10:41:06] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [10:41:12] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12250820 (10VRiley-WMF) I apologize, I thought that was all we had. [10:44:57] (03CR) 10Blake: [C:03+2] rest-gateway: Upgrade to envoy 1.39.0-1 in production. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328549 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:45:08] (03CR) 10Hnowlan: [C:04-1] profile: configure apache to optionally connect to OpenSearch with TLS (035 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [10:45:43] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921 (10fgiunchedi) 03NEW [10:46:44] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Physical move of servers - https://phabricator.wikimedia.org/T435922 (10VRiley-WMF) 03NEW [10:46:45] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Move cloudvirt1075 from F4 to E4 - https://phabricator.wikimedia.org/T435923 (10fgiunchedi) 03NEW [10:46:54] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Physical move of servers - https://phabricator.wikimedia.org/T435922#12250901 (10VRiley-WMF) 05Open→03Resolved a:03VRiley-WMF This is completed [10:47:23] (03Merged) 10jenkins-bot: rest-gateway: Upgrade to envoy 1.39.0-1 in production. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328549 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:47:34] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Ensure cloudvirt capacity is more evenly spread out among racks - https://phabricator.wikimedia.org/T424658#12250906 (10fgiunchedi) >>! In T424658#12250820, @VRiley-WMF wrote: > I apologize, I thought that was all we had. No problem at all, I filed... [10:48:02] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest langid model version under llm ns on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329236 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [10:49:56] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/rest-gateway: apply [10:50:25] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply [10:53:32] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [10:53:51] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [10:54:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [10:54:49] (03CR) 10Marostegui: [C:03+1] site.pp, db-test1001.yaml: Decommission db-test1001 [puppet] - 10https://gerrit.wikimedia.org/r/1329234 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [10:55:15] (03CR) 10Marostegui: [C:03+1] site.pp,db2903.yaml,preseed.yaml: add db2903 [puppet] - 10https://gerrit.wikimedia.org/r/1329239 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [10:55:30] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/rest-gateway: apply [10:55:43] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply [10:56:14] (03PS12) 10Brouberol: dse-k8s-eqiad: define a meta helmfile for kerberized apps [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329166 (https://phabricator.wikimedia.org/T435903) [10:56:57] FIRING: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:57:29] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest articlequality model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329241 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [10:57:37] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest langid model version under llm ns on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329237 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [10:59:01] I'm seeing a bunch of errors related to not finding a test-s4 sectiin [10:59:03] *section [10:59:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.11% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:01:05] (03PS1) 10Majavah: P:wmcs::novaproxy: Fix extracting IPv6 addresses from URLs [puppet] - 10https://gerrit.wikimedia.org/r/1329253 (https://phabricator.wikimedia.org/T429930) [11:01:35] (03CR) 10Federico Ceratto: [C:03+2] site.pp, db-test1001.yaml: Decommission db-test1001 [puppet] - 10https://gerrit.wikimedia.org/r/1329234 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [11:01:55] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:01:57] RESOLVED: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:02:24] (03CR) 10Federico Ceratto: [C:03+2] site.pp,db2903.yaml,preseed.yaml: add db2903 [puppet] - 10https://gerrit.wikimedia.org/r/1329239 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [11:03:40] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:07:02] !log fceratto@cumin1003 Removing db-test1001 from zarcillo T435912 [11:07:07] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) [11:07:08] T435912: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912 [11:07:16] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12251037 (10ops-monitoring-bot) db-test1001 has been deleted from zarcillo [11:07:19] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12251039 (10ops-monitoring-bot) db-test1001 has been decommissioned by Data Persistence [11:07:22] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12251041 (10ops-monitoring-bot) This host is ready for DC-Ops to decommission [11:08:44] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2901.codfw.wmnet with reason: Cloning [11:13:55] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:15:31] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [11:15:40] RESOLVED: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:16:06] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:16:40] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db2903.codfw.wmnet [11:17:04] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:17:42] RESOLVED: JobUnavailable: Reduced availability for job mysql-test in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:18:20] (03CR) 10Muehlenhoff: [C:03+2] proton: Bump image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329249 (owner: 10Muehlenhoff) [11:19:11] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [11:19:38] FIRING: GnmiInterfaceCountersDrop: ... [11:19:38] lsw1-e8-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-e8-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [11:20:06] !log jmm@deploy1003 helmfile [staging] START helmfile.d/services/proton: apply [11:21:01] fceratto@cumin1003 netbox (PID 3543742) is awaiting input [11:21:17] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [11:22:01] !log jmm@deploy1003 helmfile [staging] DONE helmfile.d/services/proton: apply [11:24:12] !log jmm@deploy1003 helmfile [codfw] START helmfile.d/services/proton: apply [11:25:00] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest langid model version under llm ns on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329236 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [11:25:37] !log jmm@deploy1003 helmfile [codfw] DONE helmfile.d/services/proton: apply [11:25:49] !log jmm@deploy1003 helmfile [eqiad] START helmfile.d/services/proton: apply [11:26:51] fceratto@cumin1003 makevm (PID 3544444) is awaiting input [11:27:03] !log jmm@deploy1003 helmfile [eqiad] DONE helmfile.d/services/proton: apply [11:27:07] (03Merged) 10jenkins-bot: ml-services: Deploy latest langid model version under llm ns on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329236 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [11:27:25] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) [11:27:31] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [11:28:18] !log gkyziridis@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . [11:28:30] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2901 ipv6 addr - fceratto@cumin1003" [11:28:34] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db2901 ipv6 addr - fceratto@cumin1003" [11:28:34] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:29:29] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest langid model version under llm ns on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329237 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [11:31:33] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Fix extracting IPv6 addresses from URLs [puppet] - 10https://gerrit.wikimedia.org/r/1329253 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [11:31:43] (03Merged) 10jenkins-bot: ml-services: Deploy latest langid model version under llm ns on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329237 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [11:31:46] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2903.codfw.wmnet - fceratto@cumin1003" [11:31:50] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2903.codfw.wmnet - fceratto@cumin1003" [11:31:50] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:31:50] !log fceratto@cumin1003 START - Cookbook sre.dns.wipe-cache db2903.codfw.wmnet on all recursors [11:31:53] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2903.codfw.wmnet on all recursors [11:32:25] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2903.codfw.wmnet - fceratto@cumin1003" [11:32:30] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2903.codfw.wmnet - fceratto@cumin1003" [11:33:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:34:03] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [11:34:14] !log gkyziridis@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' . [11:35:06] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 25 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324031 (https://phabricator.wikimedia.org/T429122) (owner: 10Abijeet Patro) [11:35:23] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest articlequality model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329241 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [11:35:30] fceratto@cumin1003 makevm (PID 3544444) is awaiting input [11:37:36] (03Merged) 10jenkins-bot: ml-services: Deploy latest articlequality model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329241 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [11:38:16] !log fceratto@cumin1003 START - Cookbook sre.hosts.reimage for host db2903.codfw.wmnet with OS trixie [11:38:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:38:50] !log gkyziridis@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . [11:40:19] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest articlequality model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329243 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [11:40:40] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Fix extracting IPv6 addresses from URLs [puppet] - 10https://gerrit.wikimedia.org/r/1329253 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [11:42:13] (03PS1) 10Ladsgroup: etcd: Ignore test-s4 in externalLoads too [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329255 (https://phabricator.wikimedia.org/T435059) [11:42:28] (03Merged) 10jenkins-bot: ml-services: Deploy latest articlequality model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329243 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [11:43:06] (03PS1) 10Cathal Mooney: Enable RPM probes for HE circuits to magru [homer/public] - 10https://gerrit.wikimedia.org/r/1329256 (https://phabricator.wikimedia.org/T435855) [11:43:17] jouncebot: nowandnext [11:43:17] No deployments scheduled for the next 0 hour(s) and 16 minute(s) [11:43:17] In 0 hour(s) and 16 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1200) [11:43:40] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:45:01] (03PS2) 10Cathal Mooney: Enable RPM probes for HE circuits to magru [homer/public] - 10https://gerrit.wikimedia.org/r/1329256 (https://phabricator.wikimedia.org/T435855) [11:46:28] (03CR) 10Marostegui: [C:03+1] etcd: Ignore test-s4 in externalLoads too [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329255 (https://phabricator.wikimedia.org/T435059) (owner: 10Ladsgroup) [11:46:31] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329255 (https://phabricator.wikimedia.org/T435059) (owner: 10Ladsgroup) [11:47:29] (03Merged) 10jenkins-bot: etcd: Ignore test-s4 in externalLoads too [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329255 (https://phabricator.wikimedia.org/T435059) (owner: 10Ladsgroup) [11:47:53] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1329255|etcd: Ignore test-s4 in externalLoads too (T435059)]] [11:47:57] T435059: Deploy test-s4 with new hostnames - https://phabricator.wikimedia.org/T435059 [11:48:23] (03PS1) 10Marostegui: dbconfig.schema: Add test-s4 to main [puppet] - 10https://gerrit.wikimedia.org/r/1329257 (https://phabricator.wikimedia.org/T427059) [11:48:52] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12251206 (10Cyberpower678) The Wayback Machine engineers have responded with the following: The hostname en.wikipedia.org was classified as "dead host". This happens when we ge... [11:49:44] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' . [11:49:52] !log gkyziridis@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' . [11:50:08] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1329255|etcd: Ignore test-s4 in externalLoads too (T435059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:50:30] (03CR) 10CI reject: [V:04-1] dbconfig.schema: Add test-s4 to main [puppet] - 10https://gerrit.wikimedia.org/r/1329257 (https://phabricator.wikimedia.org/T427059) (owner: 10Marostegui) [11:50:42] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [11:53:01] !log fceratto@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on db2903.codfw.wmnet with reason: host reimage [11:53:01] (03PS1) 10Gkyziridis: ml-services: Deploy latest readability model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329258 (https://phabricator.wikimedia.org/T429675) [11:54:04] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:54:10] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:54:31] (03PS1) 10Gkyziridis: ml-services: Deploy latest readability model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329259 (https://phabricator.wikimedia.org/T429675) [11:55:04] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:55:06] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:55:06] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329255|etcd: Ignore test-s4 in externalLoads too (T435059)]] (duration: 07m 13s) [11:55:12] T435059: Deploy test-s4 with new hostnames - https://phabricator.wikimedia.org/T435059 [11:57:48] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12251241 (10Cyberpower678) The takeaway here is the live web checker seemed to have gotten a large number of 503s or 999s. [11:58:45] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2903.codfw.wmnet with reason: host reimage [11:59:25] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1200) [12:00:47] (03CR) 10Brouberol: [C:03+2] dse-k8s-eqiad: define a meta helmfile for kerberized apps [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329166 (https://phabricator.wikimedia.org/T435903) (owner: 10Brouberol) [12:03:12] (03PS3) 10Cathal Mooney: Enable RPM probes for HE circuits to magru [homer/public] - 10https://gerrit.wikimedia.org/r/1329256 (https://phabricator.wikimedia.org/T435855) [12:04:25] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:07:09] (03PS1) 10Gkyziridis: ml-services: Deploy latest reference_quality image in reference-need and reference-risk model servers under the revision-models ns. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329264 (https://phabricator.wikimedia.org/T429675) [12:08:52] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12251269 (10MoritzMuehlenhoff) >>! In T435354#12250165, @MoritzMuehlenhoff wrote: >> Perhaps there was a regression with the kernel update? Would the re... [12:09:33] (03PS2) 10Marostegui: dbconfig.schema: Add test-s4 to main [puppet] - 10https://gerrit.wikimedia.org/r/1329257 (https://phabricator.wikimedia.org/T427059) [12:12:34] (03PS1) 10Brouberol: Remove file that should not have been committed [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329265 [12:12:59] (03PS1) 10Gkyziridis: ml-services: Deploy latest article-descriptions version. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329266 (https://phabricator.wikimedia.org/T429675) [12:13:29] (03CR) 10Cathal Mooney: [C:03+2] Enable RPM probes for HE circuits to magru [homer/public] - 10https://gerrit.wikimedia.org/r/1329256 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [12:13:30] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [12:13:48] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [12:13:51] (03PS1) 10Cathal Mooney: gnmic metrics: add subscription to Juniper rpm probe paths [puppet] - 10https://gerrit.wikimedia.org/r/1329267 (https://phabricator.wikimedia.org/T435855) [12:14:57] (03Merged) 10jenkins-bot: Enable RPM probes for HE circuits to magru [homer/public] - 10https://gerrit.wikimedia.org/r/1329256 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [12:15:23] PROBLEM - Host wikikube-worker1135 is DOWN: PING CRITICAL - Packet loss = 33%, RTA = 2181.39 ms [12:15:27] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Metrics, 13Patch-For-Review: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855#12251284 (10cmooney) [12:15:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:15:55] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:16:07] RECOVERY - Host wikikube-worker1135 is UP: PING OK - Packet loss = 0%, RTA = 0.55 ms [12:20:38] (03CR) 10Ladsgroup: "One question" [puppet] - 10https://gerrit.wikimedia.org/r/1329257 (https://phabricator.wikimedia.org/T427059) (owner: 10Marostegui) [12:22:55] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:23:46] (03CR) 10KartikMistry: [C:03+1] ULS: Remove wgULSLanguageSelectorV2Enabled config [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324031 (https://phabricator.wikimedia.org/T429122) (owner: 10Abijeet Patro) [12:24:02] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [12:24:05] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:24:15] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db-test1002.eqiad.wmnet [12:25:03] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:26:15] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest article-descriptions version. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329266 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [12:26:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:27:33] (03PS2) 10Federico Ceratto: site.pp, db-test1002.yaml: Decommission db-test1002 [puppet] - 10https://gerrit.wikimedia.org/r/1329269 (https://phabricator.wikimedia.org/T435912) [12:27:40] (03CR) 10Ladsgroup: [C:03+2] tests: add test to verify rsvg use of language codes are as expected [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1327566 (https://phabricator.wikimedia.org/T337139) (owner: 10Hnowlan) [12:28:49] (03CR) 10Marostegui: dbconfig.schema: Add test-s4 to main (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329257 (https://phabricator.wikimedia.org/T427059) (owner: 10Marostegui) [12:28:58] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2903.codfw.wmnet with OS trixie [12:28:59] !log fceratto@cumin1003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db2903.codfw.wmnet [12:29:10] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [12:29:23] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Metrics, 13Patch-For-Review: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855#12251323 (10cmooney) So thinking about this further, after rubber-ducking it with Hugh earlier, I... [12:29:45] (03CR) 10Marostegui: [C:03+1] site.pp, db-test1002.yaml: Decommission db-test1002 [puppet] - 10https://gerrit.wikimedia.org/r/1329269 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [12:31:26] (03CR) 10Cathal Mooney: "I'm self-merging this one as we need these stats and there is nobody else about who is familiar to review. The config has been tested in " [puppet] - 10https://gerrit.wikimedia.org/r/1329267 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [12:31:31] (03CR) 10Cathal Mooney: [C:03+2] gnmic metrics: add subscription to Juniper rpm probe paths [puppet] - 10https://gerrit.wikimedia.org/r/1329267 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [12:31:56] (03PS5) 10Blake: memcache: add a prometheus node-textfile for cert expiry. [puppet] - 10https://gerrit.wikimedia.org/r/1329247 (https://phabricator.wikimedia.org/T353511) [12:32:05] (03PS2) 10Santiago Faci: Added a property to define instake URL for beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329246 (https://phabricator.wikimedia.org/T433711) [12:32:59] !log mvernon@cumin2003 START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe [12:33:15] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test1002.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [12:33:47] (03Merged) 10jenkins-bot: tests: add test to verify rsvg use of language codes are as expected [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1327566 (https://phabricator.wikimedia.org/T337139) (owner: 10Hnowlan) [12:34:12] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:34:14] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:34:25] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [12:34:38] !log mvernon@cumin2003 START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling reboot on A:thanos-fe [12:34:47] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test1002.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [12:34:47] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:34:48] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db-test1002.eqiad.wmnet [12:35:03] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops, and 2 others: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12251332 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by fceratto@cumin1003 for hosts: `db-test1002.eqiad.wmnet` - db-test1002.eqiad.wmnet (**PASS**) - Downtimed... [12:35:31] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest reference_quality image in reference-need and reference-risk model servers under the revision-models ns. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329264 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [12:35:44] RECOVERY - Check whether ferm is active by checking the default input chain on wikikube-worker1068 is OK: OK ferm input default policy is set https://wikitech.wikimedia.org/wiki/Monitoring/check_ferm [12:35:44] RECOVERY - Check whether ferm is active by checking the default input chain on wikikube-worker1262 is OK: OK ferm input default policy is set https://wikitech.wikimedia.org/wiki/Monitoring/check_ferm [12:35:50] (03PS1) 10Majavah: P:wmcs::novaproxy: Increment file descriptor limit [puppet] - 10https://gerrit.wikimedia.org/r/1329271 (https://phabricator.wikimedia.org/T429930) [12:37:48] fceratto@cumin1003 decommission (PID 3595970) is awaiting input [12:37:55] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:38:08] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:38:09] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Increment file descriptor limit [puppet] - 10https://gerrit.wikimedia.org/r/1329271 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:38:10] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:38:22] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest reference_quality image in reference-need and reference-risk model servers under the revision-models ns. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329264 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [12:38:38] (03CR) 10Federico Ceratto: [C:03+2] site.pp, db-test1002.yaml: Decommission db-test1002 [puppet] - 10https://gerrit.wikimedia.org/r/1329269 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [12:40:25] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete ipv6 addr - fceratto@cumin1003" [12:40:29] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete ipv6 addr - fceratto@cumin1003" [12:40:29] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:40:49] (03Merged) 10jenkins-bot: ml-services: Deploy latest reference_quality image in reference-need and reference-risk model servers under the revision-models ns. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329264 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [12:42:40] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:42:57] !log fceratto@cumin1003 Removing db-test1002 from zarcillo T435912 [12:42:59] !log gkyziridis@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . [12:43:01] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) [12:43:02] T435912: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912 [12:43:11] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12251360 (10ops-monitoring-bot) db-test1002 has been deleted from zarcillo [12:43:13] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12251362 (10ops-monitoring-bot) db-test1002 has been decommissioned by Data Persistence [12:43:15] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12251363 (10ops-monitoring-bot) This host is ready for DC-Ops to decommission [12:44:15] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [12:44:36] (03CR) 10Jelto: [V:03+1 C:03+2] service::catalog: Set ipip_encapsulation for apertium in eqiad. [puppet] - 10https://gerrit.wikimedia.org/r/1328349 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [12:44:47] (03PS2) 10Jelto: service::catalog: Set ipip_encapsulation for apertium in eqiad. [puppet] - 10https://gerrit.wikimedia.org/r/1328349 (https://phabricator.wikimedia.org/T420436) [12:44:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:45:27] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' . [12:45:37] !log gkyziridis@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' . [12:45:40] PROBLEM - orchestrator resolve cache non-FQDNs on dborch1002 is CRITICAL: CRITICAL: 1 non-FQDN entries in orchestrator resolve cache: https://wikitech.wikimedia.org/wiki/Orchestrator [12:47:19] fceratto@cumin1003 decommission (PID 3609191) is awaiting input [12:49:03] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db-test1003.eqiad.wmnet [12:49:06] (03PS3) 10Santiago Faci: Added a property to define instake URL for beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329246 (https://phabricator.wikimedia.org/T433711) [12:49:17] (03PS1) 10Federico Ceratto: site.pp, db-test1003.yaml: Decommission db-test1003 [puppet] - 10https://gerrit.wikimedia.org/r/1329273 (https://phabricator.wikimedia.org/T435912) [12:51:48] !log mvernon@cumin2003 END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe [12:52:12] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:52:14] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:52:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:52:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:53:59] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [12:54:12] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest article-descriptions version. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329266 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [12:54:12] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:54:12] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:54:44] (03CR) 10Ladsgroup: [C:03+1] dbconfig.schema: Add test-s4 to main [puppet] - 10https://gerrit.wikimedia.org/r/1329257 (https://phabricator.wikimedia.org/T427059) (owner: 10Marostegui) [12:54:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:56:36] (03Merged) 10jenkins-bot: ml-services: Deploy latest article-descriptions version. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329266 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [12:57:17] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:58:12] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [12:58:19] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:59:23] RESOLVED: GnmiInterfaceCountersDrop: ... [12:59:23] lsw1-e8-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-e8-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:00:04] Lucas_WMDE, urbanecm, and TheresNoTime: #bothumor I � Unicode. All rise for UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1300). [13:00:04] No Gerrit patches in the queue for this window AFAICS. [13:00:15] (03PS1) 10Hashar: rake_modules: stop rspec testing on Debian Buster (10) [puppet] - 10https://gerrit.wikimedia.org/r/1329250 (https://phabricator.wikimedia.org/T435917) [13:00:15] (03CR) 10Hashar: "I will address that in a follow up." [puppet] - 10https://gerrit.wikimedia.org/r/1329250 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [13:01:16] fceratto@cumin1003 decommission (PID 3609191) is awaiting input [13:01:32] !log gkyziridis@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . [13:02:15] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:03:15] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:05:57] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [13:05:57] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:05:58] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db-test1003.eqiad.wmnet [13:06:10] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware, 13Patch-For-Review: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12251490 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by fceratto@cumin1003 for hosts: `db-test1003.eqiad.wmnet` - db-test1003.eqiad.wmnet... [13:08:58] fceratto@cumin1003 decommission (PID 3609191) is awaiting input [13:12:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:13:35] (03CR) 10Brouberol: [C:03+2] datahub-next: remove duplicate key [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321520 (owner: 10Mvolz) [13:13:53] FIRING: [4x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:14:36] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . [13:14:44] !log gkyziridis@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' . [13:14:55] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:15:30] (03CR) 10Brouberol: "@hnowlan@wikimedia.org I think you need to update the `team` tag in the associated tests as well" [alerts] - 10https://gerrit.wikimedia.org/r/1319827 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [13:15:50] (03CR) 10Ssingh: [C:03+1] "Thanks for cleaning it up! (Out of an abundance of caution, let's disable Puppet on A:cp-upload, test on one host and then roll it out, th" [puppet] - 10https://gerrit.wikimedia.org/r/1329224 (https://phabricator.wikimedia.org/T427646) (owner: 10Slyngshede) [13:16:15] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:16:17] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:16:29] (03CR) 10Brouberol: [C:03+1] Declare the webrequest.dumps.v1 stream in EventStreamConfig [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328201 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [13:16:45] (03CR) 10Brouberol: [C:03+1] Remove the webrequest.dumps.dev0 stream from EventStreamConfig [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328202 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [13:17:15] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:17:15] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:17:40] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:17:54] (03CR) 10Atsuko: [C:03+1] Remove file that should not have been committed [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329265 (owner: 10Brouberol) [13:17:54] (03CR) 10Bking: [C:03+2] Remove file that should not have been committed [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329265 (owner: 10Brouberol) [13:19:13] 10ops-esams, 06SRE, 06Commons, 06DC-Ops, and 3 others: ESAMS and others serving older revisions of overwritten files - https://phabricator.wikimedia.org/T425216#12251545 (10Krinkle) The report date (2 May) aligns well with after the first rollout of this to Commons: >>! In T414338#11880509, on 1 May 2026:... [13:19:21] 10ops-esams, 06SRE, 06Commons, 06DC-Ops, and 3 others: ESAMS and others serving older revisions of overwritten files - https://phabricator.wikimedia.org/T425216#12251548 (10Krinkle) [13:19:24] !log installing python3.13 security updates [13:19:27] (03CR) 10Elukey: [C:03+1] docker::baseimages: Switch to /srv/debuerreotype as the build directory [puppet] - 10https://gerrit.wikimedia.org/r/1328551 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [13:19:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:20:14] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip_encapsulation for apertium in eqiad. [puppet] - 10https://gerrit.wikimedia.org/r/1328349 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [13:20:23] (03CR) 10Elukey: [C:03+2] docker_registry: add /v2/repos/releng to the Releng's location [puppet] - 10https://gerrit.wikimedia.org/r/1328666 (https://phabricator.wikimedia.org/T432829) (owner: 10Elukey) [13:20:36] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, 06MediaWiki-Core-Platform-Team (Kanban): Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12251556 (10Krinkle) [13:20:46] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, 06MediaWiki-Core-Platform-Team (Kanban): Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12251559 (10Krinkle) p:05Triage→03High a:03Krinkle [13:20:48] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [13:21:10] (03CR) 10Brouberol: [C:03+1] logstash: Consume the webrequest.dumps.v1 stream from Kafka [puppet] - 10https://gerrit.wikimedia.org/r/1328205 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [13:21:28] (03CR) 10Brouberol: [C:03+1] dumps: web: Produce access logs to the webrequest.dumps.v1 stream [puppet] - 10https://gerrit.wikimedia.org/r/1328206 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [13:21:39] (03CR) 10Brouberol: [C:03+1] logstash: Stop consuming the webrequest.dumps.dev0 stream from Kafka [puppet] - 10https://gerrit.wikimedia.org/r/1328208 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [13:23:14] (03CR) 10Slyngshede: [C:03+2] R:cache::upload remove unused profile netconsole [puppet] - 10https://gerrit.wikimedia.org/r/1329224 (https://phabricator.wikimedia.org/T427646) (owner: 10Slyngshede) [13:23:16] !log failover dumps-nfs from clouddumps1002 to clouddumps1001 - T411248 [13:23:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:23:23] T411248: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248 [13:23:43] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, 06MediaWiki-Core-Platform-Team (Kanban): Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12251581 (10Krinkle) Varnish strips query parameters for GET requests to upl... [13:24:15] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:24:17] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:24:17] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [13:24:32] (03CR) 10Ssingh: [C:03+2] conftool-data: switch urldownloader backends to trixie hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328635 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [13:24:37] RECOVERY - Check if Pybal has been restarted after pybal.conf was changed on lvs1019 is OK: OK: pybal.service was restarted after /etc/pybal/pybal.conf was changed. https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [13:24:55] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:25:11] RECOVERY - Check if Pybal has been restarted after pybal.conf was changed on lvs1020 is OK: OK: pybal.service was restarted after /etc/pybal/pybal.conf was changed. https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [13:25:15] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:25:17] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:25:23] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [13:25:23] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad [13:25:43] (03PS1) 10Majavah: P:wmcs::novaproxy: Override HAProxy global max connection limit [puppet] - 10https://gerrit.wikimedia.org/r/1329280 (https://phabricator.wikimedia.org/T429930) [13:26:38] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9307/console" [puppet] - 10https://gerrit.wikimedia.org/r/1329280 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:27:52] (03PS2) 10Majavah: P:wmcs::novaproxy: Override HAProxy global max connection limit [puppet] - 10https://gerrit.wikimedia.org/r/1329280 (https://phabricator.wikimedia.org/T429930) [13:28:07] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: cluster=urldownloader [13:28:31] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9308/co" [puppet] - 10https://gerrit.wikimedia.org/r/1329280 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:28:40] FIRING: [9x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:28:53] FIRING: [4x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:29:55] FIRING: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:30:02] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Override HAProxy global max connection limit [puppet] - 10https://gerrit.wikimedia.org/r/1329280 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:30:19] (03CR) 10Majavah: [V:03+1 C:03+2] P:wmcs::novaproxy: Override HAProxy global max connection limit [puppet] - 10https://gerrit.wikimedia.org/r/1329280 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:31:01] !log mvernon@cumin2003 END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling reboot on A:thanos-fe [13:31:29] !log mvernon@cumin2003 START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling reboot on A:thanos-fe [13:33:40] RESOLVED: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:34:48] (03CR) 10Ssingh: [C:03+1] hieradata: add pdns v5 flag for dns2005 [puppet] - 10https://gerrit.wikimedia.org/r/1327155 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:35:16] (03CR) 10CDobbins: [V:03+1 C:03+2] hieradata: add pdns v5 flag for dns2005 [puppet] - 10https://gerrit.wikimedia.org/r/1327155 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:37:32] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=dns2005.* [reason: trixie upgrade] [13:38:05] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host dns2005.wikimedia.org with OS trixie [13:38:40] FIRING: [11x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:39:22] FIRING: GnmiInterfaceCountersDrop: ... [13:39:23] cr1-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=cr1-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:39:32] 06SRE, 10Infrastructure Security, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 07SecTeam-Processed, and 2 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12251686 (10Gehel) a:05bking→03None Unassigning @bking as this is a tracking task where mu... [13:39:55] FIRING: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:40:50] (03PS2) 10Majavah: P:wmcs::novaproxy: Stop reading proxy data from Redis in Nginx [puppet] - 10https://gerrit.wikimedia.org/r/1329229 (https://phabricator.wikimedia.org/T429930) [13:40:50] (03PS3) 10Majavah: P:wmcs::novaproxy: Stop writing data to Redis [puppet] - 10https://gerrit.wikimedia.org/r/1329230 (https://phabricator.wikimedia.org/T429930) [13:40:50] (03PS3) 10Majavah: P:wmcs::novaproxy: Undeploy Redis instance [puppet] - 10https://gerrit.wikimedia.org/r/1329231 (https://phabricator.wikimedia.org/T429930) [13:41:00] (03PS1) 10Hashar: rake_modules: support Debian 13 (Trixie) facts [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) [13:43:40] FIRING: [14x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:44:17] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:44:17] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:44:41] (03CR) 10Hashar: "This one is hopefully straight forward: drop Debian Buster from rspec testing :] I need that to then upgrade gems and add support for Debi" [puppet] - 10https://gerrit.wikimedia.org/r/1329250 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [13:44:55] FIRING: [15x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:46:17] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:46:17] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:47:22] (03PS2) 10Hashar: rake_modules: default to test on Bullseye & Bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1329251 (https://phabricator.wikimedia.org/T435917) [13:47:22] (03CR) 10Hashar: "Adding Bookworm to the default list of OS tested when using `WMFConfig.test_on` would certainly causes a few of our specs to suddenly fail" [puppet] - 10https://gerrit.wikimedia.org/r/1329251 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [13:47:44] (03PS4) 10Hnowlan: team-sre: Add data-engineering tag [alerts] - 10https://gerrit.wikimedia.org/r/1319827 (https://phabricator.wikimedia.org/T432376) [13:48:40] FIRING: [13x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:49:03] 06SRE, 06Traffic: netconsole being used for cache hosts? - https://phabricator.wikimedia.org/T427646#12251765 (10SLyngshede-WMF) 05Open→03Resolved p:05Triage→03Low a:03SLyngshede-WMF [13:49:06] (03CR) 10Hashar: zookeeper: add spec tests (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329228 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [13:49:21] (03PS14) 10Hashar: zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) [13:49:22] FIRING: [3x] GnmiInterfaceCountersDrop: cr1-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:51:22] (03CR) 10Hnowlan: "My bad, stacked patch fail. Updated!" [alerts] - 10https://gerrit.wikimedia.org/r/1319827 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [13:51:46] (03PS1) 10Andrew Bogott: backy2: reduce ceph bandwidth throttle by 50% [puppet] - 10https://gerrit.wikimedia.org/r/1329286 (https://phabricator.wikimedia.org/T429387) [13:52:00] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest readability model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329258 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [13:52:07] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12251774 (10Fabfur) Thanks for the update! Can you still confirm the archive.org is remained the same and/or no changes has been made regard to the network from where these req... [13:52:14] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest readability model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329259 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [13:53:24] (03PS1) 10Zabe: Close bswiktionary [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329287 (https://phabricator.wikimedia.org/T435954) [13:54:00] !log migrated the Docker Registry's /v2/repos/releng to the Releng's S3 bucket - T432829 [13:54:04] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:54:05] T432829: Move Docker images under the /v2/releng prefix to S3 - https://phabricator.wikimedia.org/T432829 [13:54:07] (03CR) 10Andrew Bogott: [C:03+2] backy2: reduce ceph bandwidth throttle by 50% [puppet] - 10https://gerrit.wikimedia.org/r/1329286 (https://phabricator.wikimedia.org/T429387) (owner: 10Andrew Bogott) [13:54:22] FIRING: [5x] GnmiInterfaceCountersDrop: cr1-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:54:33] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations, 13Patch-For-Review: Move Docker images under the /v2/releng prefix to S3 - https://phabricator.wikimedia.org/T432829#12251795 (10elukey) All migrated, let's wait a couple of days again to see if anything comes up! [13:54:48] 06SRE, 10SRE-swift-storage, 06Traffic-Icebox, 07affects-Kiwix-and-openZIM, 07Wikimedia-Performance-recommendation: Swift response sends invalid ETag header (missing double quotes) - https://phabricator.wikimedia.org/T256217#12251797 (10Krinkle) [13:54:55] FIRING: [11x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:56:45] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12251805 (10Cyberpower678) I'm not aware of any network changes. I'll double check. [13:58:40] FIRING: [9x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:58:51] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on dns2005.wikimedia.org with reason: host reimage [13:59:55] FIRING: [9x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:00:05] Deploy window Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1400) [14:02:35] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2005.wikimedia.org with reason: host reimage [14:03:40] FIRING: [8x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:04:18] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:04:22] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [14:05:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:06:18] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:06:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:06:43] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499#12251891 (10elukey) After a lot of filtering for non useful logs I found something good-enough: ` elukey@registry2004:~$ sudo jo... [14:06:44] 06SRE, 06Traffic: Investigate / Fix upload.wikimedia.org lack of Cache-Control headers - https://phabricator.wikimedia.org/T431621#12251894 (10Krinkle) > We do emit `ETag` […] See also {T295556} and {T256217}, which together make it hard or impossible for clients to avoid re-downloading images when they reval... [14:06:46] PROBLEM - Recursive DNS on 208.80.153.74 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [14:08:40] FIRING: [7x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:08:43] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic: Thumbnail urls should be versioned and sent with Cache-Control max-age headers - https://phabricator.wikimedia.org/T19577#12251911 (10Krinkle) [14:08:50] (03PS1) 10Gkyziridis: ml-services: Deploy latest logo-detection model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329292 (https://phabricator.wikimedia.org/T435946) [14:09:55] FIRING: [7x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:10:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:10:39] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest readability model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329258 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [14:11:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:13:04] (03Merged) 10jenkins-bot: ml-services: Deploy latest readability model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329258 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [14:13:17] (03CR) 10Brouberol: [C:03+1] team-sre: Add data-engineering tag [alerts] - 10https://gerrit.wikimedia.org/r/1319827 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [14:13:40] FIRING: [10x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:14:04] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12251944 (10MoritzMuehlenhoff) [14:14:50] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12251947 (10hashar) Thank you for the subtask! On the parent task I emitted the hypotheses that maybe the... [14:14:55] FIRING: [10x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:16:52] !log gkyziridis@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . [14:17:38] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12251965 (10Jhancock.wm) @Marostegui the cpu info isn't in netbox. i had to go through our procurement docs to figure out what each order of db servers had what in it. [14:18:40] FIRING: [10x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:18:54] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12251979 (10Marostegui) >>! In T435271#12251965, @Jhancock.wm wrote: > @Marostegui the cpu info isn't in netbox. i had to go through our procurement docs to figure out what each order of db servers ha... [14:19:19] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest readability model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329259 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [14:19:55] FIRING: [10x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:20:12] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db2903.codfw.wmnet with reason: Cloning [14:20:27] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12251986 (10Jhancock.wm) these servers have that cpu db1206-1263 db2185-2247 [14:20:43] (03PS4) 10Scott French: P:services_proxy::envoy: Drop support for split and introduce splits [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) [14:20:44] RECOVERY - Recursive DNS on 208.80.153.74 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [14:20:48] jouncebot: nowandnext [14:20:48] For the next 0 hour(s) and 9 minute(s): Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1400) [14:20:48] In 0 hour(s) and 9 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1430) [14:21:03] (03CR) 10Zabe: [C:03+2] Close bswiktionary [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329287 (https://phabricator.wikimedia.org/T435954) (owner: 10Zabe) [14:21:47] (03Merged) 10jenkins-bot: ml-services: Deploy latest readability model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329259 (https://phabricator.wikimedia.org/T429675) (owner: 10Gkyziridis) [14:21:58] (03Merged) 10jenkins-bot: Close bswiktionary [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329287 (https://phabricator.wikimedia.org/T435954) (owner: 10Zabe) [14:22:56] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:23:09] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:23:40] FIRING: [11x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:23:44] !log mvernon@cumin2003 END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling reboot on A:thanos-fe [14:24:01] (03PS1) 10Zabe: Disable CU on bswiktionary [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329295 (https://phabricator.wikimedia.org/T435954) [14:24:17] (03PS5) 10Scott French: P:services_proxy::envoy: Drop support for split and introduce splits [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) [14:24:48] (03CR) 10Zabe: [C:03+2] Disable CU on bswiktionary [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329295 (https://phabricator.wikimedia.org/T435954) (owner: 10Zabe) [14:24:55] FIRING: [12x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:25:13] !log jmm@cumin2003 START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on A:config-master [14:25:31] !log jmm@cumin2003 START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors [14:25:34] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors [14:26:04] (03Merged) 10jenkins-bot: Disable CU on bswiktionary [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329295 (https://phabricator.wikimedia.org/T435954) (owner: 10Zabe) [14:27:02] !log zabe@deploy1003 Started scap sync-world: Backport for [[gerrit:1329295|Disable CU on bswiktionary (T435954)]], [[gerrit:1329287|Close bswiktionary (T435954)]] [14:27:10] T435954: Close bs.wiktionary - https://phabricator.wikimedia.org/T435954 [14:27:18] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:27:35] (03PS1) 10Brouberol: Remove skein certificate expiry alert [alerts] - 10https://gerrit.wikimedia.org/r/1329296 [14:27:43] (03PS5) 10Cwhite: profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) [14:28:18] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:28:35] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' . [14:28:40] FIRING: [13x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:28:45] !log gkyziridis@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' . [14:29:17] !log zabe@deploy1003 zabe: Backport for [[gerrit:1329295|Disable CU on bswiktionary (T435954)]], [[gerrit:1329287|Close bswiktionary (T435954)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:29:48] !log zabe@deploy1003 zabe: Continuing with deployment [14:29:53] !log jmm@cumin2003 START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors [14:29:55] FIRING: [14x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:29:56] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1430) [14:33:17] (03CR) 10Elukey: "To keep archives happy:" [puppet] - 10https://gerrit.wikimedia.org/r/1328709 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [14:33:18] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:33:40] RESOLVED: [13x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:34:13] !log zabe@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329295|Disable CU on bswiktionary (T435954)]], [[gerrit:1329287|Close bswiktionary (T435954)]] (duration: 07m 11s) [14:34:18] T435954: Close bs.wiktionary - https://phabricator.wikimedia.org/T435954 [14:35:14] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12252130 (10bking) @MoritzMuehlenhoff It seems like EQIAD only has 10G connected hosts at the moment based on the PromQL query ` count( max by... [14:36:11] (03PS1) 10Elukey: profile::kerberos: fix replicate_krb_database script [puppet] - 10https://gerrit.wikimedia.org/r/1329297 (https://phabricator.wikimedia.org/T435854) [14:36:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:36:28] !log jmm@cumin2003 END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on A:config-master [14:36:38] (03CR) 10Elukey: "https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329297" [puppet] - 10https://gerrit.wikimedia.org/r/1328709 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [14:38:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 19.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:38:18] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:38:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:38:21] (03CR) 10Bking: [C:03+1] profile::kerberos: fix replicate_krb_database script [puppet] - 10https://gerrit.wikimedia.org/r/1329297 (https://phabricator.wikimedia.org/T435854) (owner: 10Elukey) [14:38:51] (03CR) 10Cwhite: profile: configure apache to optionally connect to OpenSearch with TLS (035 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [14:40:01] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329300 (https://phabricator.wikimedia.org/T430836) [14:40:04] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by zabe@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329300 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [14:40:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 979.9ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [14:40:41] zabe: Are you on train duty this week? :-) [14:41:09] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329300 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [14:41:12] no, I just need to fix the wikiversion for bswiktionary since it moved from group1 to group0 as part of its closure [14:41:25] !log elukey@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [14:42:12] !log elukey@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [14:43:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 21.64% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:44:18] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:44:40] FIRING: [10x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:44:54] RECOVERY - orchestrator resolve cache non-FQDNs on dborch1002 is OK: OK: all orchestrator resolve cache entries are FQDNs https://wikitech.wikimedia.org/wiki/Orchestrator [14:44:58] (03CR) 10Elukey: [C:03+2] profile::kerberos: fix replicate_krb_database script [puppet] - 10https://gerrit.wikimedia.org/r/1329297 (https://phabricator.wikimedia.org/T435854) (owner: 10Elukey) [14:45:01] zabe: Ah, interesting. Thanks! [14:45:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 19.51% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:45:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 979.9ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [14:46:18] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:46:25] zabe: Please ping me when you're done. I have a new release of scap to deploy. [14:46:40] !log xcollazo@deploy1003 Started deploy [analytics/refinery@cd8542f] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@cd8542f3] [14:47:18] !log xcollazo@deploy1003 Finished deploy [analytics/refinery@cd8542f] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@cd8542f3] (duration: 00m 37s) [14:47:24] (03CR) 10Scott French: [C:03+1] "Thanks, Blake!" [puppet] - 10https://gerrit.wikimedia.org/r/1329247 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:47:38] !log xcollazo@deploy1003 Started deploy [analytics/refinery@cd8542f]: Regular analytics weekly train [analytics/refinery@cd8542f3] [14:49:18] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:49:40] FIRING: [10x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:49:49] (03CR) 10Filippo Giunchedi: [C:03+1] "\o/" [puppet] - 10https://gerrit.wikimedia.org/r/1329229 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [14:49:55] FIRING: [9x] BFDdown: BFD session down between cr1-codfw and 208.80.153.74 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:50:04] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2005.wikimedia.org with OS trixie [14:50:18] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:50:35] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet [14:51:32] !log zabe@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.17 refs T430836 [14:51:38] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [14:51:49] dancy: over to you:) [14:51:58] zabe: ty [14:52:18] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Stop reading proxy data from Redis in Nginx [puppet] - 10https://gerrit.wikimedia.org/r/1329229 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [14:52:49] (03PS6) 10Blake: memcache: add a prometheus node-textfile for cert expiry. [puppet] - 10https://gerrit.wikimedia.org/r/1329247 (https://phabricator.wikimedia.org/T353511) [14:52:58] (03CR) 10Blake: memcache: add a prometheus node-textfile for cert expiry. (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329247 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:52:58] !log xcollazo@deploy1003 Finished deploy [analytics/refinery@cd8542f]: Regular analytics weekly train [analytics/refinery@cd8542f3] (duration: 05m 20s) [14:53:14] !log dancy@deploy1003 Installing scap version "4.284.0" for 3 host(s) [14:53:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:54:31] !log jelto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet with reason: Phabricator deploy [14:54:37] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1329297 (https://phabricator.wikimedia.org/T435854) (owner: 10Elukey) [14:54:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:54:49] !log xcollazo@deploy1003 Started deploy [analytics/refinery@cd8542f] (thin): Regular analytics weekly train THIN [analytics/refinery@cd8542f3] [14:54:54] !log jelto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab1004.eqiad.wmnet with reason: Phabricator deploy [14:55:12] !log dancy@deploy1003 Installation of scap version "4.284.0" completed for 3 hosts [14:55:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 17.47% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:55:28] !log jelto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab1005.eqiad.wmnet with reason: Phabricator deploy [14:56:46] (03CR) 10Scott French: [C:03+1] memcache: add a prometheus node-textfile for cert expiry. (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329247 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:57:05] !log xcollazo@deploy1003 Finished deploy [analytics/refinery@cd8542f] (thin): Regular analytics weekly train THIN [analytics/refinery@cd8542f3] (duration: 02m 16s) [14:57:18] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:57:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:57:30] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet [14:57:37] (03CR) 10Muehlenhoff: "That seems fine to me, but shouldn't be merged until we have a proper runbook for on callers" [puppet] - 10https://gerrit.wikimedia.org/r/1328218 (https://phabricator.wikimedia.org/T435648) (owner: 10Ssingh) [14:57:55] (03CR) 10Blake: [C:03+2] memcache: add a prometheus node-textfile for cert expiry. [puppet] - 10https://gerrit.wikimedia.org/r/1329247 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:58:20] RECOVERY - Check unit status of replicate-krb-database on krb2002 is OK: OK: Status of the systemd unit replicate-krb-database https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [14:58:21] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet [14:58:34] (03CR) 10Ssingh: "That's fair and I was thinking about it but I was wondering what should the runbook specify for a service? I am happy to work on it but do" [puppet] - 10https://gerrit.wikimedia.org/r/1328218 (https://phabricator.wikimedia.org/T435648) (owner: 10Ssingh) [14:59:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:00:05] jelto, arnoldokoth, mutante, and arnaudb: It is that lovely time of the day again! You are hereby commanded to deploy SRE Collaboration Services office hours. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1500). [15:00:08] !log cdobbins@cumin1003 START - Cookbook sre.hosts.remove-downtime for dns2005.wikimedia.org [15:00:09] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2005.wikimedia.org [15:00:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:00:25] (03CR) 10Muehlenhoff: "I can write up an earlier version in the next days and then we can refine this" [puppet] - 10https://gerrit.wikimedia.org/r/1328218 (https://phabricator.wikimedia.org/T435648) (owner: 10Ssingh) [15:00:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:01:26] !log brennen@deploy1003 Started deploy [phabricator/deployment@7c592da]: deploy phab2003 for T435958 [15:01:31] T435958: Deploy Phab/Phorge 2026-08-25 - https://phabricator.wikimedia.org/T435958 [15:01:39] (03CR) 10Ssingh: "Thanks, I am happy to wait until then; please also let me know if I can help? Removing the votes." [puppet] - 10https://gerrit.wikimedia.org/r/1328218 (https://phabricator.wikimedia.org/T435648) (owner: 10Ssingh) [15:01:55] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:02:18] !log brennen@deploy1003 Finished deploy [phabricator/deployment@7c592da]: deploy phab2003 for T435958 (duration: 00m 51s) [15:02:18] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:02:38] !log brennen@deploy1003 Started deploy [phabricator/deployment@7c592da]: deploy phab1004 for T435958 [15:02:45] !log cdobbins@cumin1003 conftool action : set/pooled=yes; selector: name=dns2005.* [reason: trixie upgrade] [15:03:12] !log brennen@deploy1003 Finished deploy [phabricator/deployment@7c592da]: deploy phab1004 for T435958 (duration: 00m 33s) [15:03:17] (03CR) 10Muehlenhoff: [C:03+1] "Looks great, one nit inline (but we can also simply polish that up later)" [puppet] - 10https://gerrit.wikimedia.org/r/1304784 (https://phabricator.wikimedia.org/T277841) (owner: 10Slyngshede) [15:03:26] !log brennen@deploy1003 Started deploy [phabricator/deployment@7c592da]: deploy phab1005 for T435958 [15:04:07] !log brennen@deploy1003 Finished deploy [phabricator/deployment@7c592da]: deploy phab1005 for T435958 (duration: 00m 41s) [15:04:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:05:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:05:20] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet [15:07:10] RESOLVED: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:07:22] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:08:05] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1325526 (https://phabricator.wikimedia.org/T434809) (owner: 10Ahmon Dancy) [15:08:23] (03CR) 10Muehlenhoff: [C:03+2] sssd: Make responder_idle_timeout configurable [puppet] - 10https://gerrit.wikimedia.org/r/1325526 (https://phabricator.wikimedia.org/T434809) (owner: 10Ahmon Dancy) [15:10:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 20.03% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:10:18] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:10:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:12:55] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:13:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 1.084s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [15:16:49] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12252434 (10MoritzMuehlenhoff) [15:16:51] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet [15:17:55] RESOLVED: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:18:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 1.011s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [15:20:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 20.61% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:20:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:22:18] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:23:55] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:23:59] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware, 13Patch-For-Review: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12252467 (10VRiley-WMF) a:03VRiley-WMF [15:24:05] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet [15:25:20] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:25:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:25:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:26:07] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet [15:28:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:29:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:30:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:32:58] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet [15:35:10] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:35:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 19.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:35:20] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:35:22] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:38:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:38:39] (03CR) 10Muehlenhoff: [C:03+2] docker::baseimages: Switch to /srv/debuerreotype as the build directory [puppet] - 10https://gerrit.wikimedia.org/r/1328551 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [15:39:28] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [15:40:10] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:40:20] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:41:45] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 867.7ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [15:42:17] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12252602 (10MoritzMuehlenhoff) @RobH It seems this already happened? The server became reachable earlier the day. I'll go reimage it tomorrow unless we expect the smart hands to do more on it? [15:44:20] (03PS1) 10Lerickson: Update memory settings for wdqs-next. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329313 (https://phabricator.wikimedia.org/T434854) [15:45:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 24.26% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:45:20] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:45:29] (03CR) 10Kamila Součková: "I can (and should, thanks!) do that spot check, but that needs to happen during the deployment, can't really test it beforehand. So I'd ne" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [15:46:45] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 867.7ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [15:48:13] (03CR) 10ArielGlenn: "Sounds great." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [15:48:20] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:48:53] (03PS7) 10Tiziano Fogli: kafka-logging: add kafka-logging100[7-8] to eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) [15:48:53] (03CR) 10Tiziano Fogli: "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [15:48:55] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [15:49:06] (03PS4) 10Tiziano Fogli: kafka-logging: stop kafka services on kafka-logging1003 [puppet] - 10https://gerrit.wikimedia.org/r/1329299 (https://phabricator.wikimedia.org/T432444) [15:49:16] (03PS4) 10Tiziano Fogli: kafka-logging: bring up kafka-logging1006 with node id 1006 [puppet] - 10https://gerrit.wikimedia.org/r/1329302 (https://phabricator.wikimedia.org/T432444) [15:49:26] (03PS3) 10Tiziano Fogli: kafka-logging: add kafka-logging1003 with node ids 1003 [puppet] - 10https://gerrit.wikimedia.org/r/1329304 (https://phabricator.wikimedia.org/T432444) [15:50:40] FIRING: [9x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:55:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:55:55] RESOLVED: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:59:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 21.82% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:00:05] jhathaway and rzl: How many deployers does it take to do Puppet request window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1600). [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:01:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 985.3ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [16:05:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:06:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 841.8ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [16:09:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 20.74% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:09:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:09:26] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 1657416864 and 134 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [16:10:20] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:10:29] (03CR) 10Clare Ming: [C:03+2] Test Kitchen UI: Deploy v1.5.3 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328196 (https://phabricator.wikimedia.org/T397417) (owner: 10Santiago Faci) [16:10:38] (03CR) 10Clare Ming: [C:03+2] Test Kitchen UI: Deploy v1.5.3 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328197 (https://phabricator.wikimedia.org/T397417) (owner: 10Santiago Faci) [16:10:39] FIRING: CoreBGPDown: Core BGP session down between cr1-codfw and cr1-eqiad (208.80.153.220) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [16:10:40] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:10:55] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:11:20] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:11:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:11:40] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:12:47] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.5.3 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328196 (https://phabricator.wikimedia.org/T397417) (owner: 10Santiago Faci) [16:12:50] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.5.3 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328197 (https://phabricator.wikimedia.org/T397417) (owner: 10Santiago Faci) [16:13:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [16:15:55] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:16:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:17:52] PROBLEM - MariaDB Replica Lag: pc6 on pc2016 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 467.21 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:18:06] PROBLEM - MariaDB Replica Lag: pc3 on pc2023 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 480.41 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:18:14] PROBLEM - MariaDB Replica Lag: pc1 on pc2021 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 490.03 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:18:14] PROBLEM - MariaDB Replica Lag: pc5 on pc2015 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 482.05 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:20:06] RECOVERY - MariaDB Replica Lag: pc3 on pc2023 is OK: OK slave_sql_lag Replication lag: 0.33 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:20:14] RECOVERY - MariaDB Replica Lag: pc1 on pc2021 is OK: OK slave_sql_lag Replication lag: 0.09 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:20:14] RECOVERY - MariaDB Replica Lag: pc5 on pc2015 is OK: OK slave_sql_lag Replication lag: 0.15 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:20:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-codfw and cr1-eqiad (208.80.153.220) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [16:20:41] (03PS1) 10Ladsgroup: Allow linking to thumb.wikimedia.org in TemplateStyles [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329317 (https://phabricator.wikimedia.org/T435979) [16:20:43] (03CR) 10Hnowlan: "nice! Sorry, one more thought" [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [16:20:52] RECOVERY - MariaDB Replica Lag: pc6 on pc2016 is OK: OK slave_sql_lag Replication lag: 0.17 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:20:55] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:21:26] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 51472 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [16:23:00] (03PS1) 10Aleksandar Mastilovic: Fine-tune Presto memory spilling [puppet] - 10https://gerrit.wikimedia.org/r/1329318 (https://phabricator.wikimedia.org/T435862) [16:23:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [16:24:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 18.95% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:24:20] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:24:22] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:25:17] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329318 (https://phabricator.wikimedia.org/T435862) (owner: 10Aleksandar Mastilovic) [16:25:20] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:25:22] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:25:27] (03PS1) 10Cathal Mooney: RPM Probes: add for all other transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1329319 (https://phabricator.wikimedia.org/T435855) [16:26:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 945.3ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [16:26:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:26:41] (03CR) 10CI reject: [V:04-1] RPM Probes: add for all other transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1329319 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [16:28:11] (03PS2) 10Cathal Mooney: RPM Probes: add for all other transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1329319 (https://phabricator.wikimedia.org/T435855) [16:28:22] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:28:54] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12252894 (10BLiviero-WMF) consulting with Claude, there should be no reason for intel_pstate to not recogn... [16:29:20] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [16:30:52] (03CR) 10Jforrester: [C:03+1] Allow linking to thumb.wikimedia.org in TemplateStyles [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329317 (https://phabricator.wikimedia.org/T435979) (owner: 10Ladsgroup) [16:30:55] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:31:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 908.5ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [16:34:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 20.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:35:55] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:36:18] (03CR) 10Majavah: [C:03+2] cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [16:36:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:38:54] (03CR) 10Majavah: [C:03+1] cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [16:40:55] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:41:24] (03CR) 10Atsuko: [C:03+1] Remove skein certificate expiry alert [alerts] - 10https://gerrit.wikimedia.org/r/1329296 (owner: 10Brouberol) [16:44:04] (03CR) 10Ssingh: [C:03+1] hieradata: add pdns v5 flag for dns2004 [puppet] - 10https://gerrit.wikimedia.org/r/1327154 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [16:44:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [16:45:22] (03CR) 10CDobbins: [V:03+1 C:03+2] hieradata: add pdns v5 flag for dns2004 [puppet] - 10https://gerrit.wikimedia.org/r/1327154 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [16:45:55] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:46:25] 06SRE, 10DNS, 10Domains, 06Traffic-Icebox, 07HTTPS: Merge Wikipedia subdomains into one, to discourage censorship - https://phabricator.wikimedia.org/T215071#12253004 (10BCornwall) 05Open→03Declined Thanks for the clarification. Closing. [16:47:56] (03PS2) 10Abijeet Patro: ArticleGuidance: Configure feedback links to local talk pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329314 (https://phabricator.wikimedia.org/T433483) [16:48:29] (03PS15) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [16:48:41] 06SRE, 10Wikimedia-Mailing-lists: Create new mailing lists: wikidata-admins@lists.wikimedia.org - https://phabricator.wikimedia.org/T435638#12253007 (10Saroj_Uprety) [16:49:07] (03CR) 10CI reject: [V:04-1] cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [16:49:54] (03PS16) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [16:50:47] (03CR) 10Bking: [C:03+2] Fine-tune Presto memory spilling [puppet] - 10https://gerrit.wikimedia.org/r/1329318 (https://phabricator.wikimedia.org/T435862) (owner: 10Aleksandar Mastilovic) [16:51:03] (03CR) 10Xcollazo: [C:03+1] "Suggest we reference ticket in code. Otherwise LGTM." [puppet] - 10https://gerrit.wikimedia.org/r/1329318 (https://phabricator.wikimedia.org/T435862) (owner: 10Aleksandar Mastilovic) [16:51:10] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [16:51:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 19.56% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:52:30] (03CR) 10Hnowlan: kafka-logging: add kafka-logging100[7-8] to eqiad cluster (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [16:52:31] !log dancy@deploy1003 Installing scap version "4.285.1" for 3 host(s) [16:53:21] (03PS6) 10Cwhite: profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) [16:53:33] 06SRE, 10Wikimedia-Mailing-lists: Create new mailing lists: wikidata-admins@lists.wikimedia.org - https://phabricator.wikimedia.org/T435638#12253020 (10Saroj_Uprety) I can be the second admin and have added my email address. [16:53:34] (03CR) 10Cwhite: profile: configure apache to optionally connect to OpenSearch with TLS (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [16:53:35] (03CR) 10Aleksandar Mastilovic: [V:03+1 C:03+1] Fine-tune Presto memory spilling (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1329318 (https://phabricator.wikimedia.org/T435862) (owner: 10Aleksandar Mastilovic) [16:53:45] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 871.4ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [16:54:26] !log dancy@deploy1003 Installation of scap version "4.285.1" completed for 3 hosts [16:54:44] the PHPFPMTooBusy and MediaWikiLatencyExceeded look kinda odd [16:55:04] Looks like there's some DB circuit breaking happening too [16:55:07] .. okay I think I know what this might be [16:55:15] dancy: ack, thanks! [16:56:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 21.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:57:50] (03CR) 10Hnowlan: "fwiw, this will not actually stop kafka on the host, just drop it from configurations." [puppet] - 10https://gerrit.wikimedia.org/r/1329299 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [16:58:01] (03PS2) 10Blake: mw-*: switch to envoy 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329325 (https://phabricator.wikimedia.org/T421418) [16:58:33] (03PS17) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [16:58:40] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [16:58:45] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 893.6ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [17:00:03] (03CR) 10Hnowlan: kafka-logging: bring up kafka-logging1006 with node id 1006 (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329302 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T1700) [17:00:31] o/ [17:00:39] o/ [17:01:05] FYI, dancy and I have work planned for the Infra window. please don't start new MediaWiki deployments in the interim. [17:01:22] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=dns2004.* [reason: trixie upgrade] [17:01:39] (03CR) 10Hnowlan: [C:03+1] profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [17:01:47] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host dns2004.wikimedia.org with OS trixie [17:02:00] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12253074 (10cmooney) They came back with the same update again this evening. Sigh. ` Hi Tim, With all due respect that is the update from 11 hours ago! Our service has been down for o... [17:02:01] dancy: following up about the mw-api-int issue - I'll hopefully be ready to get started in 5-10m if that's okay with you [17:02:09] ok [17:05:06] !log bking@cumin2003 START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons. [17:06:04] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Metrics, 13Patch-For-Review: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855#12253094 (10cmooney) p:05High→03Medium The stats for the magru HE circuits are working well,... [17:06:20] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:06:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:07:20] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:07:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:07:26] !log bking@cumin2003 END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons. [17:07:34] !log bking@cumin2003 START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons. [17:07:40] FIRING: [6x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:07:45] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 21.1% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:07:49] dancy: back! maybe let's start with a "plain" `scap sync-world` to kick the tires the non-failing case, then we can try the artificially induced rollback? [17:08:09] Sure thing. I'll let you take the controls. I'll be around for assistance. [17:08:26] (03CR) 10Elliottetzkorn: [C:03+1] "Lgtm let's merge!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328722 (owner: 10Jdlrobson) [17:08:52] dancy: actually, it would be helpful if you could do the scap-driving. I'm happy to wrangle all the other bits. [17:09:05] ok. Starting a scap sync-world now [17:09:20] awesome - thanks! I'll let you know when we're ready for the second one. [17:09:40] (03PS18) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [17:09:41] !log dancy@deploy1003 Started scap sync-world: testing https://gitlab.wikimedia.org/repos/releng/scap/-/merge_requests/1271 [17:09:46] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [17:10:07] !log bking@cumin2003 END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons. [17:11:32] !log bking@cumin2003 START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons. [17:12:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.7% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:13:42] !log dancy@deploy1003 Finished scap sync-world: testing https://gitlab.wikimedia.org/repos/releng/scap/-/merge_requests/1271 (duration: 04m 01s) [17:13:55] dancy: alright, my config edits are ready. [17:14:09] ok.. re-running sync [17:14:21] !log dancy@deploy1003 Started scap sync-world: testing https://gitlab.wikimedia.org/repos/releng/scap/-/merge_requests/1271 [17:14:24] awesome, thanks! [17:15:01] dancy: so, this will fail at the production stage - I've temporarily configured /bin/false as a command check for pretrain at that stage [17:16:34] (03PS3) 10Cathal Mooney: RPM Probes: add for all other transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1329319 (https://phabricator.wikimedia.org/T435855) [17:16:38] (03PS19) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [17:16:50] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [17:17:20] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:17:20] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:17:40] FIRING: [7x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:17:55] FIRING: [7x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:18:20] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:18:20] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [17:20:01] !log dancy@deploy1003 Rolling back deployment [17:20:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:20:17] swfrench-wmf: hit the /bin/false. Retried once.. now rolling back. [17:20:31] dancy: awesome [17:21:18] It is interesting that the rollback is taking non-zero time since it was a no-op deployment. [17:22:23] !log dancy@deploy1003 Finished scap sync-world: testing https://gitlab.wikimedia.org/repos/releng/scap/-/merge_requests/1271 (duration: 08m 02s) [17:22:29] interesting ... I guess there's still a lot of bookkeeping that gets done. [17:22:40] FIRING: [8x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:22:43] https://www.irccloud.com/pastebin/GAyyiGPT/ [17:22:49] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12253202 (10bking) [17:22:51] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on dns2004.wikimedia.org with reason: host reimage [17:22:54] (03CR) 10Sbisson: [C:04-1] ArticleGuidance: Configure feedback links to local talk pages (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329314 (https://phabricator.wikimedia.org/T433483) (owner: 10Abijeet Patro) [17:23:06] (03PS20) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [17:24:30] ... and that would be both bookkeeping scap is doing (e.g., various helm-status calls) and helm is doing when rolling back [17:25:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:26:15] swfrench-wmf: Anything else to do? [17:26:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1158:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1158 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [17:26:43] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [17:27:16] dancy: unless anything seemed off from your end that you'd like to test further, I think we're good! all that's left is for me to clean up my hacks :) [17:27:32] FIRING: SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [17:28:02] swfrench-wmf: I have pretty high confidence (tested thoroughly in train-dev beforehand) [17:28:27] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns2004.wikimedia.org with reason: host reimage [17:29:08] dancy: and thank you for doing so, as well as testing in production with me right now :) cool, I'll clean up my hacks and we can adjourn. thanks again! [17:29:37] Awesome. Stepping out for a break. Hit me up if anything interesting happens. [17:29:50] * swfrench-wmf thumbs up [17:30:34] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T435989 (10bking) 03NEW [17:31:31] PROBLEM - Recursive DNS on 208.80.153.48 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [17:31:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1158:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1158 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [17:32:32] RESOLVED: SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [17:32:40] FIRING: [8x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:34:32] FIRING: SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [17:35:02] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12253247 (10bking) [17:36:27] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12253254 (10bking) p:05Triage→03Medium [17:36:31] PROBLEM - Recursive DNS on 2620:0:860:2:208:80:153:48 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [17:36:51] (03PS1) 10Bking: Kerberos: Prepare a VM for deployment as krb replica [puppet] - 10https://gerrit.wikimedia.org/r/1329337 (https://phabricator.wikimedia.org/T435873) [17:37:05] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329337 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [17:37:40] FIRING: [7x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:38:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1158:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1158 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [17:40:51] (03PS1) 10Hamish: zhwiki: allow TAV remove themselves and grant abusefilter-revert to sysop and abusefilter [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329339 (https://phabricator.wikimedia.org/T435887) [17:41:22] !log bking@cumin2003 END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons. [17:41:38] (03CR) 10CI reject: [V:04-1] zhwiki: allow TAV remove themselves and grant abusefilter-revert to sysop and abusefilter [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329339 (https://phabricator.wikimedia.org/T435887) (owner: 10Hamish) [17:42:40] FIRING: [6x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:44:28] (03CR) 10Brouberol: [C:03+2] Remove skein certificate expiry alert [alerts] - 10https://gerrit.wikimedia.org/r/1329296 (owner: 10Brouberol) [17:44:42] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12253292 (10bking) Update based on the team standup: - We are OK with trying to run Kerberos in a VM. - [[ https://wikimedia... [17:47:06] (03CR) 10Bking: "The PCC failure is because krb1003 is too new. I already tried (and failed) to update Puppet facts via the manual process described at htt" [puppet] - 10https://gerrit.wikimedia.org/r/1329337 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [17:49:04] (03PS2) 10Hamish: zhwiki: allow TAV remove themselves and grant abusefilter-revert to sysop and abusefilter [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329339 (https://phabricator.wikimedia.org/T435887) [17:49:29] RECOVERY - Recursive DNS on 208.80.153.48 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [17:49:29] RECOVERY - Recursive DNS on 2620:0:860:2:208:80:153:48 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [17:49:32] RESOLVED: SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [17:51:09] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 25 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329339 (https://phabricator.wikimedia.org/T435887) (owner: 10Hamish) [17:59:01] (03PS1) 10Jasmine: site.pp, preseed.yaml: add conf101[0-2] [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) [18:02:34] !log aokoth@cumin1003 START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet [18:03:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1158:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1158 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [18:03:50] (03CR) 10Scott French: [C:03+1] "Thanks, Jasmine!" [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) (owner: 10Jasmine) [18:04:22] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [18:04:30] !log aokoth@cumin1003 END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet [18:07:40] RESOLVED: [5x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:07:55] FIRING: [5x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:08:40] RESOLVED: [2x] BFDdown: BFD session down between cr1-codfw and 208.80.153.48 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:08:44] (03PS2) 10Jasmine: site.pp, preseed.yaml: add conf101[0-2] [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) [18:08:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:10:00] (03CR) 10Jasmine: site.pp, preseed.yaml: add conf101[0-2] (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) (owner: 10Jasmine) [18:10:52] (03CR) 10Jasmine: "Ty!" [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) (owner: 10Jasmine) [18:11:57] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [18:12:22] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [18:12:36] Hello? Anyone can check T435995? [18:12:37] T435995: ci: gate-and-submit not triggering - https://phabricator.wikimedia.org/T435995 [18:15:28] (03CR) 10RLazarus: [C:03+1] "You might include mw-misc and mw-wikifunctions here too -- easy to overlook because they're both so small that they don't have a canary de" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329325 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [18:17:43] (03CR) 10Scott French: site.pp, preseed.yaml: add conf101[0-2] (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) (owner: 10Jasmine) [18:17:54] (03PS3) 10Blake: mw-*: switch to envoy 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329325 (https://phabricator.wikimedia.org/T421418) [18:18:19] (03CR) 10Blake: "Ah, right, thanks - done!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329325 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [18:19:10] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply [18:19:11] (03CR) 10RLazarus: [C:03+1] mw-*: switch to envoy 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329325 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [18:19:21] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply [18:20:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 0% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:20:17] Les4353: That change won't merge until https://gerrit.wikimedia.org/r/c/mediawiki/extensions/AbuseFilter/+/1325562/3 is +2'd [18:20:29] wikitech seems down? [18:20:39] Dreamy_Jazz: What?? My bad!! [18:20:57] FIRING: [7x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:21:01] PROBLEM - MariaDB Replica IO: pc8 on pc1018 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:21:05] PROBLEM - MariaDB Replica IO: pc1 on pc2021 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 1040, Errmsg: error connecting to master repl2024@pc1021.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Too many connections https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:21:06] Dolly Parton died, some reports of slow loading & errors from editors [18:21:19] PROBLEM - MariaDB Event Scheduler pc1 on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:21:21] PROBLEM - MariaDB Event Scheduler pc4 on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:21:22] PROBLEM - MariaDB read only pc1 #page on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:21:23] PROBLEM - MariaDB Event Scheduler pc8 on pc1018 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:21:23] PROBLEM - MariaDB read only pc8 on pc1018 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:21:26] PROBLEM - MariaDB read only pc5 #page on pc1015 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:21:33] PROBLEM - MariaDB Replica SQL: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:22:07] PROBLEM - MariaDB Replica IO: pc4 on pc2024 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 1040, Errmsg: error connecting to master repl2024@pc1024.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Too many connections https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:22:07] PROBLEM - MariaDB Replica IO: pc5 on pc2015 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 1040, Errmsg: error connecting to master repl2024@pc1015.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Too many connections https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:22:21] es.wikipedia is down [18:22:22] PROBLEM - MariaDB read only pc4 #page on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:22:25] PROBLEM - MariaDB Event Scheduler pc5 on pc1015 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:22:34] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 0% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:23:01] PROBLEM - MariaDB Replica IO: pc8 on pc1018 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:23:12] FIRING: VarnishUnavailable: varnish-text has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/Varnish#Diagnosing_Varnish_alerts - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=3 - https://alerts.wikimedia.org/?q=alertname%3DVarnishUnavailable [18:23:13] FIRING: HaproxyUnavailable: HAProxy (cache_text) has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyUnavailable [18:23:21] PROBLEM - MariaDB Event Scheduler pc4 on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:23:33] PROBLEM - MariaDB Replica SQL: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:23:36] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns2004.wikimedia.org with OS trixie [18:24:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 2.5s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [18:25:37] https://www.wikimediastatus.net/incidents/9mg9mtxfxpjl [18:25:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1158:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1158 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [18:25:57] FIRING: [8x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:26:25] PROBLEM - MariaDB Event Scheduler pc5 on pc1015 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:26:57] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:28:01] PROBLEM - MariaDB Replica Lag: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_lag could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:29:15] PROBLEM - MariaDB Replica Lag: pc1 on pc2021 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 626.74 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:29:15] FIRING: [2x] MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 1.748s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [18:29:24] PROBLEM - MariaDB read only pc4 #page on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:29:33] PROBLEM - MariaDB Replica SQL: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:30:15] PROBLEM - MariaDB Replica Lag: pc5 on pc2015 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 658.07 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:30:25] PROBLEM - MariaDB read only pc8 on pc1018 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:30:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1158:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1158 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [18:30:57] FIRING: [8x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:31:57] FIRING: [3x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:34:51] FIRING: [8x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from mw-web-ro.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [18:35:23] PROBLEM - MariaDB Event Scheduler pc4 on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:35:23] PROBLEM - MariaDB read only pc1 #page on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:35:57] FIRING: [8x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:36:25] PROBLEM - MariaDB read only pc8 on pc1018 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:36:46] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool pc1022: Pooling [18:36:46] !log fceratto@cumin1003 START - Cookbook sre.mysql.parsercache [18:36:53] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [18:36:53] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Pooling [18:37:03] FIRING: MediaWikiLoginFailures: Elevated MediaWiki centrallogin failures (centralauth_error_badtoken) - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?viewPanel=3 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLoginFailures [18:37:22] PROBLEM - MariaDB read only pc4 #page on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:37:37] PROBLEM - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:38:01] PROBLEM - MariaDB Replica Lag: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_lag could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:38:01] PROBLEM - MariaDB Replica IO: pc8 on pc1018 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:38:19] PROBLEM - MariaDB Event Scheduler pc1 on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:39:19] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [18:39:25] PROBLEM - MariaDB Event Scheduler pc8 on pc1018 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:39:25] PROBLEM - MariaDB Event Scheduler pc5 on pc1015 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:39:51] FIRING: [10x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from mw-web-ro.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [18:40:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:40:57] FIRING: [8x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:41:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [18:41:25] PROBLEM - MariaDB read only pc8 on pc1018 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:42:07] RECOVERY - MariaDB Replica IO: pc5 on pc2015 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:42:08] RESOLVED: MediaWikiLoginFailures: Elevated MediaWiki centrallogin failures (centralauth_error_badtoken) - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?viewPanel=3 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLoginFailures [18:42:24] PROBLEM - MariaDB read only pc4 #page on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:42:25] PROBLEM - MariaDB read only pc1 #page on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:42:33] PROBLEM - MariaDB Replica SQL: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:44:51] FIRING: [11x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from mw-web-ro.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [18:45:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:45:15] RECOVERY - MariaDB Replica Lag: pc5 on pc2015 is OK: OK slave_sql_lag Replication lag: 1.99 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:45:33] PROBLEM - MariaDB Replica SQL: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:45:57] FIRING: [8x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:46:01] PROBLEM - MariaDB Replica Lag: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_lag could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:46:19] PROBLEM - MariaDB Event Scheduler pc1 on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:46:24] PROBLEM - MariaDB read only pc1 #page on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:46:28] PROBLEM - MariaDB read only pc5 #page on pc1015 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:46:28] PROBLEM - MariaDB Event Scheduler pc5 on pc1015 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:46:57] FIRING: [3x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip6) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:47:22] PROBLEM - MariaDB read only pc4 #page on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:48:23] PROBLEM - MariaDB Event Scheduler pc8 on pc1018 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:48:56] !log fceratto@cumin1003 START - Cookbook sre.mysql.depool depool pc1022: depooling [18:48:56] !log fceratto@cumin1003 START - Cookbook sre.mysql.parsercache [18:48:57] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [18:48:57] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: depooling [18:49:23] PROBLEM - MariaDB read only pc8 on pc1018 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:49:27] RECOVERY - MariaDB Event Scheduler pc5 on pc1015 is OK: Version 10.11.16-MariaDB-log, Uptime 4273093s, read_only: False, event_scheduler: True, 8234.73 QPS, connection latency: 0.024482s, query latency: 0.001245s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:49:30] RECOVERY - MariaDB read only pc5 #page on pc1015 is OK: Version 10.11.16-MariaDB-log, Uptime 4273093s, read_only: False, event_scheduler: True, 8182.56 QPS, connection latency: 0.031152s, query latency: 0.000803s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:49:51] FIRING: [12x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from mw-web-ro.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [18:50:57] FIRING: [8x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:51:01] PROBLEM - MariaDB Replica Lag: pc8 on pc1018 is CRITICAL: CRITICAL slave_sql_lag could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:51:01] PROBLEM - MariaDB Replica IO: pc8 on pc1018 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:51:23] PROBLEM - MariaDB Event Scheduler pc4 on pc1024 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:51:57] FIRING: [4x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip6) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:53:12] RESOLVED: VarnishUnavailable: varnish-text has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/Varnish#Diagnosing_Varnish_alerts - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=3 - https://alerts.wikimedia.org/?q=alertname%3DVarnishUnavailable [18:53:13] RESOLVED: HaproxyUnavailable: HAProxy (cache_text) has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyUnavailable [18:54:01] RECOVERY - MariaDB Replica IO: pc8 on pc1018 is OK: OK slave_io_state not a slave https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:54:01] RECOVERY - MariaDB Replica Lag: pc8 on pc1018 is OK: OK slave_sql_lag not a slave https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:54:15] FIRING: [2x] MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 1.4s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [18:54:25] RECOVERY - MariaDB Event Scheduler pc8 on pc1018 is OK: Version 10.11.16-MariaDB-log, Uptime 3569715s, read_only: False, event_scheduler: True, 7985.68 QPS, connection latency: 0.021737s, query latency: 0.000784s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:54:25] RECOVERY - MariaDB read only pc8 on pc1018 is OK: Version 10.11.16-MariaDB-log, Uptime 3569715s, read_only: False, event_scheduler: True, 8006.29 QPS, connection latency: 0.027645s, query latency: 0.001608s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:54:33] RECOVERY - MariaDB Replica SQL: pc8 on pc1018 is OK: OK slave_sql_state not a slave https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [18:54:51] FIRING: [12x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from mw-web-ro.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [18:55:15] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 8.06% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:55:24] PROBLEM - MariaDB read only pc1 #page on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [18:55:57] RESOLVED: [8x] ProbeDown: Service mw-web:4450 has failed probes (http_mw-web_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:56:39] !log cdobbins@cumin1003 conftool action : set/pooled=yes; selector: name=dns2004.* [reason: trixie upgrade] [18:56:57] FIRING: [9x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip6) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:57:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 0% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:57:17] !log cdobbins@cumin1003 START - Cookbook sre.hosts.remove-downtime for dns2004.wikimedia.org [18:57:18] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns2004.wikimedia.org [18:57:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [18:58:19] PROBLEM - MariaDB Event Scheduler pc1 on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [18:58:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [18:59:51] RESOLVED: [9x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from mw-web-ro.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [19:00:24] RECOVERY - MariaDB read only pc4 #page on pc1024 is OK: Version 10.11.16-MariaDB-log, Uptime 4277094s, read_only: False, event_scheduler: True, 8799.33 QPS, connection latency: 0.012050s, query latency: 0.000749s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [19:00:24] RECOVERY - MariaDB Event Scheduler pc4 on pc1024 is OK: Version 10.11.16-MariaDB-log, Uptime 4277094s, read_only: False, event_scheduler: True, 8856.14 QPS, connection latency: 0.013469s, query latency: 0.000568s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [19:01:05] RECOVERY - MariaDB Replica IO: pc4 on pc2024 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:01:57] RESOLVED: [8x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip6) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:02:45] 06SRE, 06Traffic: Wikipedia and Wiktionary pages returning "upstream connect error or disconnect/reset before headers" - https://phabricator.wikimedia.org/T436004#12253722 (10-sche) [19:03:23] PROBLEM - MariaDB read only pc1 #page on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [19:04:09] !incidents [19:04:09] !incidents [19:04:10] 8295 (ACKED) pc1021 (paged)/MariaDB read only pc1 (paged) [19:04:10] 8302 (ACKED) Large wiki outage [19:04:10] 8298 (RESOLVED) pc1024 (paged)/MariaDB read only pc4 (paged) [19:04:10] 8301 (RESOLVED) [8x] ATSBackendErrorsHigh cache_text sre () [19:04:10] 8297 (RESOLVED) PHPFPMTooBusy sre (mw-web main eqiad) [19:04:11] 8294 (RESOLVED) [7x] ProbeDown sre (probes/service) [19:04:11] 8300 (RESOLVED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [19:04:11] 8299 (RESOLVED) VarnishUnavailable global sre (varnish-text thanos-rule@main) [19:04:11] 8296 (RESOLVED) pc1015 (paged)/MariaDB read only pc5 (paged) [19:04:12] 8292 (RESOLVED) PHPFPMTooBusy sre (mw-api-ext main eqiad) [19:04:12] 8291 (RESOLVED) Manual (paged) by Scott French (swfrench@wikimedia.org): Hot handoff: db1228 (m5 master) is down [19:04:13] 8290 (RESOLVED) Host db1228 (paged) [19:04:13] 8295 (ACKED) pc1021 (paged)/MariaDB read only pc1 (paged) [19:04:14] 8302 (ACKED) Large wiki outage [19:04:14] 8298 (RESOLVED) pc1024 (paged)/MariaDB read only pc4 (paged) [19:04:15] 8301 (RESOLVED) [8x] ATSBackendErrorsHigh cache_text sre () [19:04:15] 8297 (RESOLVED) PHPFPMTooBusy sre (mw-web main eqiad) [19:04:15] RESOLVED: [2x] MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-int releases routed via main (k8s) 1.442s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [19:04:16] 8294 (RESOLVED) [7x] ProbeDown sre (probes/service) [19:04:16] 8300 (RESOLVED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [19:04:17] 8299 (RESOLVED) VarnishUnavailable global sre (varnish-text thanos-rule@main) [19:04:17] 8296 (RESOLVED) pc1015 (paged)/MariaDB read only pc5 (paged) [19:04:18] 8292 (RESOLVED) PHPFPMTooBusy sre (mw-api-ext main eqiad) [19:04:18] 8291 (RESOLVED) Manual (paged) by Scott French (swfrench@wikimedia.org): Hot handoff: db1228 (m5 master) is down [19:04:19] 8290 (RESOLVED) Host db1228 (paged) [19:04:38] 06SRE, 06Traffic, 07Wikimedia-Incident: Wikipedia and Wiktionary pages returning "upstream connect error or disconnect/reset before headers" - https://phabricator.wikimedia.org/T436004#12253732 (10AntiCompositeNumber) This is a known incident being actively responded to. https://www.wikimediastatus.net/incid... [19:04:59]  !incidentsd [19:05:02] !incidents [19:05:03] You're not allowed to perform this action. [19:05:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:05:15] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 20.77% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:07:21] FIRING: SLOBudgetBurn: Search update lag is below 95% target in eqiad - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [19:10:10] RESOLVED: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:10:19] PROBLEM - MariaDB Event Scheduler pc1 on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [19:10:36] 06SRE, 06Traffic, 07Wikimedia-Incident: Wikipedia and Wiktionary pages returning "upstream connect error or disconnect/reset before headers" - https://phabricator.wikimedia.org/T436004#12253759 (10Jdforrester-WMF) [19:13:24] PROBLEM - MariaDB read only pc1 #page on pc1021 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [19:13:35] !incidents [19:13:35] 8295 (ACKED) pc1021 (paged)/MariaDB read only pc1 (paged) [19:13:35] 8302 (ACKED) Large wiki outage [19:13:36] 8298 (RESOLVED) pc1024 (paged)/MariaDB read only pc4 (paged) [19:13:36] 8301 (RESOLVED) [8x] ATSBackendErrorsHigh cache_text sre () [19:13:36] 8297 (RESOLVED) PHPFPMTooBusy sre (mw-web main eqiad) [19:13:36] 8294 (RESOLVED) [7x] ProbeDown sre (probes/service) [19:13:36] 8300 (RESOLVED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [19:13:37] 8299 (RESOLVED) VarnishUnavailable global sre (varnish-text thanos-rule@main) [19:13:37] 8296 (RESOLVED) pc1015 (paged)/MariaDB read only pc5 (paged) [19:13:38] 8292 (RESOLVED) PHPFPMTooBusy sre (mw-api-ext main eqiad) [19:13:38] 8291 (RESOLVED) Manual (paged) by Scott French (swfrench@wikimedia.org): Hot handoff: db1228 (m5 master) is down [19:13:39] 8290 (RESOLVED) Host db1228 (paged) [19:13:54] 06SRE, 06Traffic, 07Wikimedia-Incident: Wikipedia and Wiktionary pages returning "upstream connect error or disconnect/reset before headers" - https://phabricator.wikimedia.org/T436004#12253790 (10RhinosF1) [19:14:21] RECOVERY - MariaDB Event Scheduler pc1 on pc1021 is OK: Version 10.11.18-MariaDB-log, Uptime 25s, read_only: False, event_scheduler: True, 2769.75 QPS, connection latency: 0.009930s, query latency: 0.000234s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [19:14:24] RECOVERY - MariaDB read only pc1 #page on pc1021 is OK: Version 10.11.18-MariaDB-log, Uptime 28s, read_only: False, event_scheduler: True, 2508.29 QPS, connection latency: 0.009283s, query latency: 0.000238s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [19:14:54] !log bking@cumin2003 START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons. [19:15:05] RECOVERY - MariaDB Replica IO: pc1 on pc2021 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:15:15] RECOVERY - MariaDB Replica Lag: pc1 on pc2021 is OK: OK slave_sql_lag Replication lag: 0.01 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [19:15:15] RESOLVED: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 20.77% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:15:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:20:25] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:20:41] (03PS1) 10Kimberly Sarabia: Add reader experiments exclusion to CNB [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329358 (https://phabricator.wikimedia.org/T435658) [19:20:47] !incidents [19:20:48] 8302 (RESOLVED) Large wiki outage [19:20:48] 8295 (RESOLVED) pc1021 (paged)/MariaDB read only pc1 (paged) [19:20:48] 8298 (RESOLVED) pc1024 (paged)/MariaDB read only pc4 (paged) [19:20:48] 8301 (RESOLVED) [8x] ATSBackendErrorsHigh cache_text sre () [19:20:48] 8297 (RESOLVED) PHPFPMTooBusy sre (mw-web main eqiad) [19:20:49] 8294 (RESOLVED) [7x] ProbeDown sre (probes/service) [19:20:49] 8300 (RESOLVED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [19:20:49] 8299 (RESOLVED) VarnishUnavailable global sre (varnish-text thanos-rule@main) [19:20:50] 8296 (RESOLVED) pc1015 (paged)/MariaDB read only pc5 (paged) [19:20:50] 8292 (RESOLVED) PHPFPMTooBusy sre (mw-api-ext main eqiad) [19:20:51] 8291 (RESOLVED) Manual (paged) by Scott French (swfrench@wikimedia.org): Hot handoff: db1228 (m5 master) is down [19:20:51] 8290 (RESOLVED) Host db1228 (paged) [19:21:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.55% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:22:21] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [19:30:25] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:31:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.25% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:32:40] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:33:47] (03PS1) 10Dreamy Jazz: EventStreamConfig: Register the abuse_review_interaction stream [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329360 (https://phabricator.wikimedia.org/T435517) [19:37:16] (03CR) 10Andrew Bogott: [C:03+2] cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [19:37:37] RECOVERY - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:42:55] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:43:48] (03PS3) 10Jasmine: site.pp, preseed.yaml: add conf101[0-2] [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) [19:44:34] !log bking@cumin2003 END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons. [19:44:49] (03PS1) 10Xcollazo: Scale down mw-content-history-reconcile-enrich [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329361 [19:45:11] (03CR) 10Jasmine: site.pp, preseed.yaml: add conf101[0-2] (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) (owner: 10Jasmine) [19:45:30] (03PS6) 10Arlolra: prv: Enable parsoid rendering for 9 wikisources [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 (https://phabricator.wikimedia.org/T435830) (owner: 10Jgiannelos) [19:46:49] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 25 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 (https://phabricator.wikimedia.org/T435830) (owner: 10Jgiannelos) [19:46:56] !log bking@apt1002 reprepro --component thirdparty/opensearch3 update trixie-wikimedia T420230 [19:47:01] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:47:02] T420230: Migrate to OpenSearch 3.x - https://phabricator.wikimedia.org/T420230 [19:47:05] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [19:47:55] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:48:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:52:55] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:57:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:57:21] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:58:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:58:21] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:59:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.41% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:00:04] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T2000). [20:00:05] bwang, hamishcz, and arlolra: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:22] hi:) [20:00:31] o/ [20:00:47] great [20:00:57] impact from the earlier incident is resolved, you're all clear to proceed from SRE's point of view [20:03:06] does anyone need a deployer? otherwise feel free to self-service per queue order [20:03:15] i think i need [20:03:49] hamishcz: i got you [20:03:54] Thank you clare [20:04:03] bwang: do you want me to deploy for you? [20:04:09] And I consider I have no permission to self-service right? otherwise I can try to do tha [20:04:11] Yes pls [20:04:12] that* [20:04:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:04:36] (03PS10) 10Bernard Wang: Remove reading list experiment instrumentation [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [20:04:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [20:05:05] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [20:06:30] (03Merged) 10jenkins-bot: Remove reading list experiment instrumentation [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [20:06:53] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1259251|Remove reading list experiment instrumentation (T421939)]] [20:06:58] T421939: [Reading List] Experiment cleanup - https://phabricator.wikimedia.org/T421939 [20:07:05] (03PS5) 10Cwhite: logstash: add security-plugin required fields to output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) [20:08:55] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [20:09:29] !log cjming@deploy1003 lmora, cjming: Backport for [[gerrit:1259251|Remove reading list experiment instrumentation (T421939)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:09:37] bwang: ok to sync? [20:10:04] (03CR) 10Cwhite: [C:03+2] profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:11:06] Lemme check ! [20:11:11] Which test survey> [20:11:19] mwdebug [20:11:21] ty [20:11:38] i did a quick codesearch - doesn't seem to have any refs [20:13:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [20:14:39] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12254073 (10RobH) DHL closed our case stating: > Hello again Rob > > The delivery station has sent the message below. Please reach out to your shipper for further help.... [20:16:01] (03PS1) 10Arlolra: Set a default for wgUseParsoidMessages in prod [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329369 (https://phabricator.wikimedia.org/T432760) [20:16:55] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [20:17:39] bwang: gtg? [20:18:30] yes [20:18:37] !log cjming@deploy1003 lmora, cjming: Continuing with deployment [20:18:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [20:18:55] (03CR) 10Cwhite: [C:03+2] logstash: add security-plugin required fields to output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:19:09] (03PS5) 10Bernard Wang: Enable readinglist for phase0 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) [20:20:05] (03CR) 10Cwhite: [C:03+2] prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:21:22] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 25 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329369 (https://phabricator.wikimedia.org/T432760) (owner: 10Arlolra) [20:22:21] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [20:22:56] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1259251|Remove reading list experiment instrumentation (T421939)]] (duration: 16m 03s) [20:23:01] T421939: [Reading List] Experiment cleanup - https://phabricator.wikimedia.org/T421939 [20:23:24] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) (owner: 10Bernard Wang) [20:24:15] (03Merged) 10jenkins-bot: Enable readinglist for phase0 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) (owner: 10Bernard Wang) [20:24:21] (03PS1) 10Krinkle: LR: Only unset deriv state for inner rendering [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1329370 (https://phabricator.wikimedia.org/T434456) [20:24:36] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1328239|Enable readinglist for phase0 wikis (T435258)]] [20:24:40] T435258: Enable ReadingLists for all logged-in users [Aug 25] on phase 0 wikis (bnwiki, cswiki, viwiki and zhwiki) - https://phabricator.wikimedia.org/T435258 [20:25:26] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 25 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1329370 (https://phabricator.wikimedia.org/T434456) (owner: 10Krinkle) [20:26:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.48% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:26:35] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12254127 (10cmooney) They seemed to have pulled their finger out now. ` 2026-08-25 20:22:05 GMT - The maintenance window starts at 12am CST. We will follow up again with the TNOC once th... [20:29:35] !log cjming@deploy1003 cjming, bwang: Backport for [[gerrit:1328239|Enable readinglist for phase0 wikis (T435258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:29:41] T435258: Enable ReadingLists for all logged-in users [Aug 25] on phase 0 wikis (bnwiki, cswiki, viwiki and zhwiki) - https://phabricator.wikimedia.org/T435258 [20:31:00] (03PS2) 10Kamila Součková: Add title-case mapping to support migration to PHP 8.5 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) [20:31:06] bwang: gtg on 2nd patch? [20:31:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.48% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:31:19] What’s the second one sorry [20:31:37] The enable reading list one right [20:32:08] (03CR) 10Cathal Mooney: [C:03+2] RPM Probes: add for all other transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1329319 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [20:32:43] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [20:32:43] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [20:32:43] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [20:32:44] (03CR) 10Cathal Mooney: [C:03+2] "merged so I can compare the current stats at other sites to magru before making a call on un-draining. we can remove if it seems a good i" [homer/public] - 10https://gerrit.wikimedia.org/r/1329319 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [20:33:33] Ok we are good to go [20:33:34] Thanks clare [20:33:37] (03Merged) 10jenkins-bot: RPM Probes: add for all other transport circuits [homer/public] - 10https://gerrit.wikimedia.org/r/1329319 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [20:33:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:34:11] !log cjming@deploy1003 cjming, bwang: Continuing with deployment [20:34:25] bwang: np :) [20:37:21] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [20:38:39] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328239|Enable readinglist for phase0 wikis (T435258)]] (duration: 14m 03s) [20:38:44] T435258: Enable ReadingLists for all logged-in users [Aug 25] on phase 0 wikis (bnwiki, cswiki, viwiki and zhwiki) - https://phabricator.wikimedia.org/T435258 [20:40:12] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329339 (https://phabricator.wikimedia.org/T435887) (owner: 10Hamish) [20:40:30] (03CR) 10Kamila Součková: "File checked and commented with https://gitlab.wikimedia.org/repos/sre/php-upgrade-tools/-/blob/main/src/php_upgrade_tools/__main__.py?ref" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [20:41:23] (03Merged) 10jenkins-bot: zhwiki: allow TAV remove themselves and grant abusefilter-revert to sysop and abusefilter [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329339 (https://phabricator.wikimedia.org/T435887) (owner: 10Hamish) [20:41:42] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1329339|zhwiki: allow TAV remove themselves and grant abusefilter-revert to sysop and abusefilter (T435887)]] [20:41:47] T435887: User group rights changes for abuse filter and temporary-account-viewer on zhwiki - https://phabricator.wikimedia.org/T435887 [20:41:59] (03CR) 10Eevans: [C:03+2] Record that a principal was created for user mkrolik-wmf [puppet] - 10https://gerrit.wikimedia.org/r/1328736 (https://phabricator.wikimedia.org/T434877) (owner: 10Eevans) [20:43:00] (03CR) 10ArielGlenn: [C:03+1] "Here's my +1." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [20:43:30] 06SRE, 10SRE-Access-Requests, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12254190 (10Eevans) 05In progress→03Resolved a:03Eevans Closing; Please feel free to reopen... [20:44:14] !log cjming@deploy1003 cjming, hamishz: Backport for [[gerrit:1329339|zhwiki: allow TAV remove themselves and grant abusefilter-revert to sysop and abusefilter (T435887)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:45:04] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [20:45:06] (03PS1) 10Cwhite: beta-logs: enable security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1329373 (https://phabricator.wikimedia.org/T350516) [20:45:16] cjming: LGTM [20:45:21] !log cjming@deploy1003 cjming, hamishz: Continuing with deployment [20:46:02] ty :) [20:46:22] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12254196 (10Eevans) [20:48:16] yw! [20:48:53] (03PS1) 10Eevans: admin: replace ssh key for user bgwiki [puppet] - 10https://gerrit.wikimedia.org/r/1329374 (https://phabricator.wikimedia.org/T433313) [20:49:38] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329339|zhwiki: allow TAV remove themselves and grant abusefilter-revert to sysop and abusefilter (T435887)]] (duration: 07m 55s) [20:49:42] T435887: User group rights changes for abuse filter and temporary-account-viewer on zhwiki - https://phabricator.wikimedia.org/T435887 [20:50:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [20:50:21] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [20:50:44] (03PS1) 10Bking: presto: enable spill-to-disk for coordinators [puppet] - 10https://gerrit.wikimedia.org/r/1329375 (https://phabricator.wikimedia.org/T435862) [20:50:46] cjming: I'm up? [20:50:54] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329375 (https://phabricator.wikimedia.org/T435862) (owner: 10Bking) [20:51:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [20:51:21] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [20:51:28] arlolra: yes! are you good to self-service? [20:51:32] yes [20:51:34] thanks [20:51:37] all yours [20:51:50] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 (https://phabricator.wikimedia.org/T435830) (owner: 10Jgiannelos) [20:51:51] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329369 (https://phabricator.wikimedia.org/T432760) (owner: 10Arlolra) [20:52:54] (03Merged) 10jenkins-bot: prv: Enable parsoid rendering for 9 wikisources [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 (https://phabricator.wikimedia.org/T435830) (owner: 10Jgiannelos) [20:52:57] (03Merged) 10jenkins-bot: Set a default for wgUseParsoidMessages in prod [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329369 (https://phabricator.wikimedia.org/T432760) (owner: 10Arlolra) [20:53:17] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1328528|prv: Enable parsoid rendering for 9 wikisources (T435830)]], [[gerrit:1329369|Set a default for wgUseParsoidMessages in prod (T432760)]] [20:53:27] T435830: Rollout parsoid read views on wikisource - week of 21 Aug - https://phabricator.wikimedia.org/T435830 [20:53:27] T432760: Add configuration options to core for Parsoid by default - https://phabricator.wikimedia.org/T432760 [20:53:28] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12254225 (10BCornwall) Bless you, Rob. - {F100099161, layout=link} [20:53:33] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [20:53:33] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [20:53:33] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.130 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [20:53:38] (03PS1) 10Cwhite: opensearch: update location of ca cert to more accessible location [puppet] - 10https://gerrit.wikimedia.org/r/1329376 (https://phabricator.wikimedia.org/T350516) [20:53:41] (03PS1) 10Cathal Mooney: Fix error in template for rpm probes [homer/public] - 10https://gerrit.wikimedia.org/r/1329377 (https://phabricator.wikimedia.org/T435855) [20:55:01] (03PS2) 10Cwhite: opensearch: update location of ca cert to more accessible location [puppet] - 10https://gerrit.wikimedia.org/r/1329376 (https://phabricator.wikimedia.org/T350516) [20:55:40] (03CR) 10Cathal Mooney: [C:03+2] Fix error in template for rpm probes [homer/public] - 10https://gerrit.wikimedia.org/r/1329377 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [20:55:50] !log arlolra@deploy1003 jgiannelos, arlolra: Backport for [[gerrit:1328528|prv: Enable parsoid rendering for 9 wikisources (T435830)]], [[gerrit:1329369|Set a default for wgUseParsoidMessages in prod (T432760)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:56:30] !log arlolra@deploy1003 jgiannelos, arlolra: Continuing with deployment [20:57:32] (03CR) 10Scott French: [C:03+1] "Thanks, Jasmine!" [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) (owner: 10Jasmine) [20:58:41] (03CR) 10Cwhite: [C:03+2] opensearch: update location of ca cert to more accessible location [puppet] - 10https://gerrit.wikimedia.org/r/1329376 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:58:53] (03Merged) 10jenkins-bot: Fix error in template for rpm probes [homer/public] - 10https://gerrit.wikimedia.org/r/1329377 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [20:58:57] Hey all - would like to get out a couple of security patches if the Readers window isn’t being used. [20:59:08] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12254240 (10andrea.denisse) >>! In T435271#12251965, @Jhancock.wm wrote: > @Marostegui the cpu info isn't in netbox. i had to go through our procurement docs to figure out what each order of db server... [20:59:23] sbassett: waiting for arlolra and myself backporting, should be done in 20min [21:00:04] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T2100) [21:00:25] Krinkle: o [21:00:27] k [21:00:51] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328528|prv: Enable parsoid rendering for 9 wikisources (T435830)]], [[gerrit:1329369|Set a default for wgUseParsoidMessages in prod (T432760)]] (duration: 07m 33s) [21:00:57] T435830: Rollout parsoid read views on wikisource - week of 21 Aug - https://phabricator.wikimedia.org/T435830 [21:00:57] T432760: Add configuration options to core for Parsoid by default - https://phabricator.wikimedia.org/T432760 [21:01:01] (03PS1) 10Cwhite: opensearch_dashboards: clean up config documentation [puppet] - 10https://gerrit.wikimedia.org/r/1329379 (https://phabricator.wikimedia.org/T350516) [21:01:04] Krinkle: done [21:01:06] thx [21:01:15] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1329370 (https://phabricator.wikimedia.org/T434456) (owner: 10Krinkle) [21:01:39] (03CR) 10Cwhite: [C:03+2] beta-logs: enable security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1329373 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [21:02:48] (03Merged) 10jenkins-bot: LR: Only unset deriv state for inner rendering [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1329370 (https://phabricator.wikimedia.org/T434456) (owner: 10Krinkle) [21:03:18] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1329370|LR: Only unset deriv state for inner rendering (T434456)]] [21:03:22] T434456: Missing prime symbol after \right) in MathML and Client Side SVG - https://phabricator.wikimedia.org/T434456 [21:03:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [21:05:55] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1329370|LR: Only unset deriv state for inner rendering (T434456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:06:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:06:35] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [21:07:00] !help [21:07:01] You're not allowed to perform this action. [21:07:01] want docs? ask for "!wm-bot". all keywords? try "@regsearch .*" [21:08:43] !log krinkle@deploy1003 krinkle: Continuing with deployment [21:11:10] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:11:26] (03CR) 10Jdlrobson: [C:03+1] Add reader experiments exclusion to CNB [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329358 (https://phabricator.wikimedia.org/T435658) (owner: 10Kimberly Sarabia) [21:13:03] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329370|LR: Only unset deriv state for inner rendering (T434456)]] (duration: 09m 45s) [21:13:05] jouncebot: now [21:13:05] For the next 0 hour(s) and 46 minute(s): Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260825T2100) [21:13:08] T434456: Missing prime symbol after \right) in MathML and Client Side SVG - https://phabricator.wikimedia.org/T434456 [21:14:48] Results (Found 80): puppet, morebots, bang, nagios, bot, labs-home-wm, labs-nagios-wm, labs-morebots, gerrit-wm, wiki, bastion, extension, wm-bot, putty, gerrit, wikitech, revision, monitor, alert, password, unicorn, help, bz, os-change, leslie's-reset, damianz's-reset, credentials, queue, info, logging, ask, $realm, keys, $site, bug, pageant, blueprint-dns, pxe, ghsh, rt, erb, regsubst, bots, wt, gerrit-search, change, dn, opshelp, testwiki, sal, task, thx, depool, cluster, hiera, infobot, zuul, jouncebot, jenkins, mira, selfie, hug, instance, git, amend, security, labs, projects, instancelist, instance-del, pathconflict, terminology, access, socks-proxy, sudo, ping, tzag, emoji, rain_dance, kuai, [21:14:48] @regsearch .* [21:15:22] Krinkle: do you have a few more? [21:15:58] Les4353: do you need something? :) [21:16:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:16:28] rzl: no thanks, just curious :) [21:19:04] (03PS1) 10Cwhite: prometheus: bugfix elasticsearch_exporter ca config [puppet] - 10https://gerrit.wikimedia.org/r/1329382 (https://phabricator.wikimedia.org/T350516) [21:21:23] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:21:25] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:21:45] (03CR) 10Cwhite: [C:03+2] prometheus: bugfix elasticsearch_exporter ca config [puppet] - 10https://gerrit.wikimedia.org/r/1329382 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [21:22:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:22:21] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:24:16] (03CR) 10Aleksandar Mastilovic: presto: enable spill-to-disk for coordinators (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329375 (https://phabricator.wikimedia.org/T435862) (owner: 10Bking) [21:26:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:26:21] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:27:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [21:28:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:28:21] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:30:58] (03PS2) 10Bking: presto: enable spill-to-disk for coordinators [puppet] - 10https://gerrit.wikimedia.org/r/1329375 (https://phabricator.wikimedia.org/T435862) [21:31:08] (03CR) 10Bking: presto: enable spill-to-disk for coordinators (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329375 (https://phabricator.wikimedia.org/T435862) (owner: 10Bking) [21:35:08] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:35:08] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:35:08] (03CR) 10Aleksandar Mastilovic: [V:03+1 C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1329375 (https://phabricator.wikimedia.org/T435862) (owner: 10Bking) [21:35:09] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [21:35:09] (03PS1) 10Kimberly Sarabia: [MinMin] Do not display language count badge for 0 languages [skins/MinervaNeue] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329386 (https://phabricator.wikimedia.org/T435249) [21:35:09] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 26 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [skins/MinervaNeue] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329386 (https://phabricator.wikimedia.org/T435249) (owner: 10Kimberly Sarabia) [21:36:57] (03CR) 10Bking: [C:03+2] presto: enable spill-to-disk for coordinators [puppet] - 10https://gerrit.wikimedia.org/r/1329375 (https://phabricator.wikimedia.org/T435862) (owner: 10Bking) [21:38:21] (03PS1) 10Cwhite: beta-logs: update opensearch output to use fqdn [puppet] - 10https://gerrit.wikimedia.org/r/1329387 (https://phabricator.wikimedia.org/T350516) [21:40:26] !log Removed security mitigation for T435455 [21:40:30] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:43:04] (03CR) 10Cwhite: [C:03+2] beta-logs: update opensearch output to use fqdn [puppet] - 10https://gerrit.wikimedia.org/r/1329387 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [21:48:13] (03PS1) 10Lerickson: Increase WDQS backend pod storage to 900GiB. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329388 (https://phabricator.wikimedia.org/T436023) [21:48:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:51:25] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:52:09] (03PS1) 10Bking: Presto: Set required spill config properties [puppet] - 10https://gerrit.wikimedia.org/r/1329389 (https://phabricator.wikimedia.org/T435862) [21:52:18] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329389 (https://phabricator.wikimedia.org/T435862) (owner: 10Bking) [21:54:39] (03CR) 10Aleksandar Mastilovic: [V:03+1 C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1329389 (https://phabricator.wikimedia.org/T435862) (owner: 10Bking) [21:55:17] (03CR) 10Bking: [C:03+2] Presto: Set required spill config properties [puppet] - 10https://gerrit.wikimedia.org/r/1329389 (https://phabricator.wikimedia.org/T435862) (owner: 10Bking) [21:56:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:57:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [21:57:55] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:02:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:04:23] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [22:04:25] PROBLEM - Ensure trafficserver_exporter is running for instance backend on cp5017 is CRITICAL: PROCS CRITICAL: 3 processes with args /usr/bin/python3 /usr/bin/prometheus-trafficserver-exporter --no-procstats --no-ssl-verification --endpoint http://127.0.0.1:3128/_stats --port 9122 https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [22:04:40] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:05:25] RECOVERY - Ensure trafficserver_exporter is running for instance backend on cp5017 is OK: PROCS OK: 1 process with args /usr/bin/python3 /usr/bin/prometheus-trafficserver-exporter --no-procstats --no-ssl-verification --endpoint http://127.0.0.1:3128/_stats --port 9122 https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [22:06:09] (03PS1) 10Cwhite: opensearch: update locationmatch search endpoint to use https when required [puppet] - 10https://gerrit.wikimedia.org/r/1329391 (https://phabricator.wikimedia.org/T350516) [22:09:25] FIRING: BFDdown: BFD session down between cr2-eqiad and 208.80.154.217 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:10:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:11:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [22:11:33] (03CR) 10Cwhite: [C:03+2] opensearch: update locationmatch search endpoint to use https when required [puppet] - 10https://gerrit.wikimedia.org/r/1329391 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:12:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [22:13:12] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 26 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329358 (https://phabricator.wikimedia.org/T435658) (owner: 10Kimberly Sarabia) [22:14:25] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:15:40] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:17:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:19:07] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12254502 (10andrea.denisse) >>! In T435271#12254240, @andrea.denisse wrote: >>>! In T435271#12251965, @Jhancock.wm wrote: >> @Marostegui the cpu info isn't in netbox. i had to go through our procureme... [22:19:25] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:22:21] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [22:27:21] RESOLVED: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [22:29:25] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:32:40] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [22:37:23] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [22:38:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [22:44:16] (03PS1) 10Cwhite: opensearch: pass curator username/password through server profile [puppet] - 10https://gerrit.wikimedia.org/r/1329393 (https://phabricator.wikimedia.org/T350516) [22:44:59] (03PS1) 10Cwhite: Revert "beta-logs: enable security plugin" [puppet] - 10https://gerrit.wikimedia.org/r/1329394 [22:45:32] (03PS1) 10RLazarus: external_clouds_vendors: chown dump_cloud_ip_ranges job to Traffic [puppet] - 10https://gerrit.wikimedia.org/r/1329395 [22:46:06] (03CR) 10Cwhite: [C:03+2] Revert "beta-logs: enable security plugin" [puppet] - 10https://gerrit.wikimedia.org/r/1329394 (owner: 10Cwhite) [22:47:13] (03CR) 10RLazarus: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9313/co" [puppet] - 10https://gerrit.wikimedia.org/r/1329395 (owner: 10RLazarus) [22:49:32] (03CR) 10RLazarus: external_clouds_vendors: chown dump_cloud_ip_ranges job to Traffic [puppet] - 10https://gerrit.wikimedia.org/r/1329395 (owner: 10RLazarus) [22:55:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:00:10] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:05:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:06:25] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:06:25] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:07:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:07:21] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:10:25] FIRING: [5x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:15:25] RESOLVED: [5x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:20:25] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:25:25] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:26:56] 06SRE, 10LDAP-Access-Requests: Grant Access to wmf, for WMF staff/contractors nda group for gsduser - https://phabricator.wikimedia.org/T435852#12254673 (10Eevans) 05Open→03Resolved a:03Eevans Hi @GSduser, if I understand the request correctly, and you just need access to group !!wmf!!, then you shou... [23:27:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:30:25] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:32:10] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:38:40] FIRING: [5x] BFDdown: BFD session down between cr1-magru and 195.200.68.136 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:40:25] RESOLVED: [5x] BFDdown: BFD session down between cr1-magru and 195.200.68.136 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:41:13] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1329401 [23:41:14] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1329401 (owner: 10TrainBranchBot) [23:43:40] FIRING: [6x] BFDdown: BFD session down between cr1-magru and 195.200.68.136 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:45:25] RESOLVED: [6x] BFDdown: BFD session down between cr1-magru and 195.200.68.136 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:46:14] (03PS1) 10RLazarus: docker: Add team labels to docker-reporter jobs [puppet] - 10https://gerrit.wikimedia.org/r/1329402 [23:46:42] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1329401 (owner: 10TrainBranchBot) [23:47:24] (03CR) 10RLazarus: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9314/co" [puppet] - 10https://gerrit.wikimedia.org/r/1329402 (owner: 10RLazarus) [23:55:25] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:55:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [23:58:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown