[00:03:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:05:25] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:05:44] FIRING: [4x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [00:08:40] FIRING: [3x] BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:10:25] RESOLVED: [3x] BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:15:25] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:15:44] RESOLVED: [4x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [00:18:40] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:25:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:30:25] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:45:05] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [00:48:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:53:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:55:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:55:21] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:56:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:56:21] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:58:25] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:00:23] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:00:23] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 7/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:02:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:02:21] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:03:25] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:06:35] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [01:11:14] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1329405 [01:11:14] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1329405 (owner: 10TrainBranchBot) [01:14:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:14:55] RESOLVED: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:20:15] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1329405 (owner: 10TrainBranchBot) [01:21:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:21:21] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:22:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:25:21] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:43:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:44:39] FIRING: CoreBGPDown: Core BGP session down between cr1-codfw and cr1-magru (195.200.68.139) - group Confed_magru - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_magru&var-bgp_neighbor=cr1-magru - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [01:48:40] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [01:49:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-codfw and cr1-magru (195.200.68.139) - group Confed_magru - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_magru&var-bgp_neighbor=cr1-magru - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [01:57:57] (03CR) 10Scott French: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [02:00:37] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:04:38] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [02:05:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:08:13] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 35s) [02:10:40] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:11:55] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:16:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:21:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:21:55] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:26:40] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:42:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:47:40] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:04:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:04:23] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:04:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:05:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:05:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:08:12] (03PS9) 10Scott French: P:kubernetes::deployment_server::global_config: Update services_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1328256 (https://phabricator.wikimedia.org/T427666) [03:08:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:09:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:09:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:14:40] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:19:08] !log Upgrading Grafana on codfw - T435816 [03:19:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [03:19:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:24:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:29:23] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:29:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:30:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:30:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:30:55] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:31:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:34:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:36:55] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:39:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:03:55] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:04:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:08:40] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:15:47] PROBLEM - mailman list info on lists1004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [04:15:47] PROBLEM - mailman archives on lists1004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [04:20:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:22:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:22:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:27:10] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:35:41] RECOVERY - mailman archives on lists1004 is OK: HTTP OK: HTTP/1.1 200 OK - 55878 bytes in 5.388 second response time https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [04:35:41] RECOVERY - mailman list info on lists1004 is OK: HTTP OK: HTTP/1.1 200 OK - 9235 bytes in 5.542 second response time https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [04:45:05] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [04:45:40] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:50:40] RESOLVED: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:59:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [04:59:29] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:00:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:00:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:03:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:03:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [05:08:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:13:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:19:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:19:55] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 208.80.154.217 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:24:40] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:24:55] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:29:40] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:29:51] FIRING: ATSBackendErrorsHigh: ATS: elevated 5xx errors from lists.discovery.wmnet in eqiad #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://grafana.wikimedia.org/d/1T_4O08Wk/ats-backends-origin-servers-overview?orgId=1&viewPanel=12&var-site=eqiad&var-cluster=text&var-origin=lists.discovery.wmnet - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [05:30:17] <_joe_> we should really exclude lists.wikimedia.org from this alert [05:30:36] <_joe_> !ack [05:30:37] All incidents are already acked. [05:30:41] Do we have a task for that? [05:30:57] <_joe_> not sure, I know g.odog was talking about it [05:31:02] * Emperor here [05:31:11] ...and waking up from the p.age. Again. [05:31:13] <_joe_> this is probably some scraper [05:31:39] <_joe_> but tbh, I don't think we need to wake people up for lists.w.org sending some 500s [05:40:47] PROBLEM - mailman archives on lists1004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [05:40:47] PROBLEM - mailman list info on lists1004 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [05:42:03] (03PS1) 10Kevin Bazira: ml-services: deploy qwen36 image with tool-call history normalization [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329417 (https://phabricator.wikimedia.org/T434274) [05:47:40] 06SRE, 06SRE Observability, 10Wikimedia-Mailing-lists: Consider disabling paging for lists.wikimedia.org - https://phabricator.wikimedia.org/T436049 (10Marostegui) 03NEW [05:56:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [05:56:27] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T0600) [06:00:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:00:45] RECOVERY - mailman archives on lists1004 is OK: HTTP OK: HTTP/1.1 200 OK - 55878 bytes in 9.318 second response time https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [06:00:45] RECOVERY - mailman list info on lists1004 is OK: HTTP OK: HTTP/1.1 200 OK - 9235 bytes in 9.467 second response time https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [06:01:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:03:20] (03CR) 10Tim Starling: [C:03+1] Add title-case mapping to support migration to PHP 8.5 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [06:04:27] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:04:38] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [06:04:51] RESOLVED: ATSBackendErrorsHigh: ATS: elevated 5xx errors from lists.discovery.wmnet in eqiad #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://grafana.wikimedia.org/d/1T_4O08Wk/ats-backends-origin-servers-overview?orgId=1&viewPanel=12&var-site=eqiad&var-cluster=text&var-origin=lists.discovery.wmnet - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [06:05:23] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:06:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:06:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:08:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:11:31] (03PS1) 10Muehlenhoff: lists: Double the amount of UWSGI processes [puppet] - 10https://gerrit.wikimedia.org/r/1329418 [06:11:54] (03PS1) 10Andrea Denisse: grafana: Disable automatic updates for plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329416 (https://phabricator.wikimedia.org/T436045) [06:12:25] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:13:10] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:14:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:18:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:19:00] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12255185 (10Marostegui) This is very useful @andrea.denisse thanks a lot. I'll leave this for DCOps but I think we can definitely benefit from such info in netbox cc @wiki_willy [06:20:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:20:23] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:20:50] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12255188 (10Marostegui) For the shake of this task, does it need to be open or can this be closed - we've sort of hijacked it with different comments already. From my side what I know is: * We are coo... [06:21:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:21:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:22:04] (03PS1) 10Slyngshede: CAS update, version 7.3.8.2 [software/cas-overlay-template] - 10https://gerrit.wikimedia.org/r/1329419 [06:22:07] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1329374 (https://phabricator.wikimedia.org/T433313) (owner: 10Eevans) [06:23:25] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:24:27] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:25:36] (03PS2) 10Slyngshede: CAS update, version 7.3.8.2 [software/cas-overlay-template] - 10https://gerrit.wikimedia.org/r/1329419 [06:26:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:26:46] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [software/cas-overlay-template] - 10https://gerrit.wikimedia.org/r/1329419 (owner: 10Slyngshede) [06:27:18] (03CR) 10Slyngshede: [V:03+2 C:03+2] CAS update, version 7.3.8.2 [software/cas-overlay-template] - 10https://gerrit.wikimedia.org/r/1329419 (owner: 10Slyngshede) [06:27:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:28:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:28:25] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:30:23] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:31:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:32:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:32:25] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:33:25] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:35:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:35:23] PROBLEM - OSPF status on cr1-eqiad is CRITICAL: OSPFv2: 6/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:35:23] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:35:33] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 5/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:36:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:36:23] RECOVERY - OSPF status on cr1-eqiad is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:36:33] RECOVERY - OSPF status on cr1-magru is OK: OSPFv2: 5/5 UP : OSPFv3: 5/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:38:25] FIRING: [7x] BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:39:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:41:25] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:42:14] (03CR) 10Bartosz Wójtowicz: [C:03+1] ml-services: deploy qwen36 image with tool-call history normalization [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329417 (https://phabricator.wikimedia.org/T434274) (owner: 10Kevin Bazira) [06:42:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:43:25] RESOLVED: [6x] BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:43:48] jouncebot: next [06:43:49] In 0 hour(s) and 16 minute(s): UTC morning backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T0700) [06:47:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:49:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:49:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [06:53:23] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:53:25] FIRING: [4x] BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:54:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:54:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [06:57:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:58:25] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:59:07] (03CR) 10Gmodena: [C:03+1] Increase WDQS backend pod storage to 900GiB. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329388 (https://phabricator.wikimedia.org/T436023) (owner: 10Lerickson) [06:59:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:00:03] (03PS1) 10Slyngshede: IDP: CAS version 7.8.3.2 [dns] - 10https://gerrit.wikimedia.org/r/1329424 [07:00:04] Amir1, urbanecm, and awight: #bothumor I � Unicode. All rise for UTC morning backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T0700). [07:00:04] matthiasmullie: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:38] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy qwen36 image with tool-call history normalization [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329417 (https://phabricator.wikimedia.org/T434274) (owner: 10Kevin Bazira) [07:00:42] o/ [07:02:15] matthiasmullie: o/ you'll self deploy wouldn't you? [07:02:16] FIRING: [2x] ProbeDown: Service idp1005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:02:30] yeah [07:03:05] (03Merged) 10jenkins-bot: ml-services: deploy qwen36 image with tool-call history normalization [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329417 (https://phabricator.wikimedia.org/T434274) (owner: 10Kevin Bazira) [07:03:43] great :) [07:04:29] !log slyngshede@cumin1003 START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: eqsin switch migration, T435406] [07:04:33] T435406: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406 [07:04:37] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mlitn@deploy1003 using scap backport" [skins/MinervaNeue] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329386 (https://phabricator.wikimedia.org/T435249) (owner: 10Kimberly Sarabia) [07:04:38] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: eqsin switch migration, T435406] [07:04:50] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet [07:05:09] (03CR) 10Marostegui: [C:03+1] site.pp, db-test1003.yaml: Decommission db-test1003 [puppet] - 10https://gerrit.wikimedia.org/r/1329273 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [07:05:43] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255247 (10SLyngshede-WMF) eqsin depooled: ` slyngshede@cumin1003:~$ sudo cookbook sre.dns.admin depool eqsin -t T435406 -r "eqsin switch mi... [07:05:54] (03Merged) 10jenkins-bot: [MinMin] Do not display language count badge for 0 languages [skins/MinervaNeue] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329386 (https://phabricator.wikimedia.org/T435249) (owner: 10Kimberly Sarabia) [07:06:29] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:06:56] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [07:07:09] !log mlitn@deploy1003 Started scap sync-world: Backport for [[gerrit:1329386|[MinMin] Do not display language count badge for 0 languages (T435249)]] [07:07:13] T435249: simplified Minerva toolbelt - the Language counter displays count for artciles with no interlanguage links - https://phabricator.wikimedia.org/T435249 [07:07:15] RESOLVED: [2x] ProbeDown: Service idp1005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:07:23] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:08:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:08:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:09:47] !log mlitn@deploy1003 mlitn, ksarabia: Backport for [[gerrit:1329386|[MinMin] Do not display language count badge for 0 languages (T435249)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:11:29] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet [07:11:34] !log mlitn@deploy1003 mlitn, ksarabia: Continuing with deployment [07:14:25] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:15:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:15:58] !log mlitn@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329386|[MinMin] Do not display language count badge for 0 languages (T435249)]] (duration: 08m 49s) [07:16:03] T435249: simplified Minerva toolbelt - the Language counter displays count for artciles with no interlanguage links - https://phabricator.wikimedia.org/T435249 [07:16:22] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mlitn@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329358 (https://phabricator.wikimedia.org/T435658) (owner: 10Kimberly Sarabia) [07:16:45] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255289 (10cmooney) [07:17:17] (03Merged) 10jenkins-bot: Add reader experiments exclusion to CNB [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329358 (https://phabricator.wikimedia.org/T435658) (owner: 10Kimberly Sarabia) [07:17:37] !log mlitn@deploy1003 Started scap sync-world: Backport for [[gerrit:1329358|Add reader experiments exclusion to CNB (T435658)]] [07:17:43] T435658: [MinMin] No banners in experiment - https://phabricator.wikimedia.org/T435658 [07:18:19] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12255293 (10cmooney) ` 2026-08-26 05:08:20 GMT - The Lumen NOC has confirmed the Urgent Maintenance Network Event to address a card issue in Kansas City, MO is commencing at this time.... [07:19:08] !log installing Linux 6.12.105 on trixie hosts [07:19:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:19:12] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:19:22] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet [07:20:14] !log mlitn@deploy1003 mlitn, ksarabia: Backport for [[gerrit:1329358|Add reader experiments exclusion to CNB (T435658)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:22:34] !log mlitn@deploy1003 mlitn, ksarabia: Continuing with deployment [07:23:25] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:23:27] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:24:10] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:25:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:26:02] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet [07:26:53] !log mlitn@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329358|Add reader experiments exclusion to CNB (T435658)]] (duration: 09m 16s) [07:26:59] T435658: [MinMin] No banners in experiment - https://phabricator.wikimedia.org/T435658 [07:27:15] All done; rest of the window up for grabs [07:27:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:29:25] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:30:10] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:30:21] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:30:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:32:27] (03CR) 10Filippo Giunchedi: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9315/co" [puppet] - 10https://gerrit.wikimedia.org/r/1327853 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [07:33:03] (03CR) 10Filippo Giunchedi: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9316/co" [puppet] - 10https://gerrit.wikimedia.org/r/1327854 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [07:33:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:33:32] (03CR) 10Filippo Giunchedi: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9317/co" [puppet] - 10https://gerrit.wikimedia.org/r/1327855 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [07:34:25] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:35:09] (03CR) 10Filippo Giunchedi: "Untested but LGTM, I'll let Daniel vote" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [07:36:25] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:36:25] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:37:23] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:39:25] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:39:38] (03PS2) 10MVernon: lists: Double the amount of UWSGI processes [puppet] - 10https://gerrit.wikimedia.org/r/1329418 (owner: 10Muehlenhoff) [07:40:15] (03CR) 10CI reject: [V:04-1] lists: Double the amount of UWSGI processes [puppet] - 10https://gerrit.wikimedia.org/r/1329418 (owner: 10Muehlenhoff) [07:40:21] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:40:40] FIRING: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:42:13] (03PS3) 10Jelto: lists: Double the amount of UWSGI processes [puppet] - 10https://gerrit.wikimedia.org/r/1329418 (https://phabricator.wikimedia.org/T436047) (owner: 10Muehlenhoff) [07:42:44] (03PS4) 10MVernon: lists: Double the amount of UWSGI processes [puppet] - 10https://gerrit.wikimedia.org/r/1329418 (https://phabricator.wikimedia.org/T436047) (owner: 10Muehlenhoff) [07:44:31] RESOLVED: [8x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:44:57] 06SRE, 06SRE Observability, 10Wikimedia-Mailing-lists: Consider disabling paging for lists.wikimedia.org - https://phabricator.wikimedia.org/T436049#12255335 (10MatthewVernon) There is perhaps a broader question: "Now that we have 24/7 oncall, are there other services that we shouldn't be waking people up to... [07:45:43] 06SRE, 06Collaboration-Services, 06SRE Observability, 10Wikimedia-Mailing-lists: Consider disabling paging for lists.wikimedia.org - https://phabricator.wikimedia.org/T436049#12255336 (10Jelto) [07:45:44] (03CR) 10Brouberol: [C:03+2] Scale down mw-content-history-reconcile-enrich [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329361 (owner: 10Xcollazo) [07:48:02] (03Merged) 10jenkins-bot: Scale down mw-content-history-reconcile-enrich [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329361 (owner: 10Xcollazo) [07:48:11] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply [07:48:15] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply [07:48:22] (03CR) 10Jelto: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9318/co" [puppet] - 10https://gerrit.wikimedia.org/r/1329418 (https://phabricator.wikimedia.org/T436047) (owner: 10Muehlenhoff) [07:49:02] (03PS2) 10Andrea Denisse: Add grafana-prometheus-datasource [debs/grafana-plugins] - 10https://gerrit.wikimedia.org/r/1329426 (https://phabricator.wikimedia.org/T436045) [07:49:43] (03CR) 10Jelto: [V:03+1 C:03+1] "lgtm, thank you. Let me know when I should take care of deploying this" [puppet] - 10https://gerrit.wikimedia.org/r/1329418 (https://phabricator.wikimedia.org/T436047) (owner: 10Muehlenhoff) [07:54:17] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet [07:55:15] (03PS1) 10Phedenskog: prometheus: scrape the CI feedback time exporter [puppet] - 10https://gerrit.wikimedia.org/r/1329430 (https://phabricator.wikimedia.org/T435814) [08:00:04] hashar and andre: Deploy window MediaWiki train - Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T0800) [08:00:44] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet [08:04:44] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 26 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [08:04:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:05:02] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 26 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [08:05:04] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329453 (https://phabricator.wikimedia.org/T430836) [08:05:07] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by hashar@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329453 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [08:05:57] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329453 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [08:11:58] 06SRE: Improve reuse-parts story for standard recipes - https://phabricator.wikimedia.org/T436058 (10fgiunchedi) 03NEW [08:12:15] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet [08:12:57] !log hashar@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.17 refs T430836 [08:13:02] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [08:14:09] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on 12 hosts with reason: site migration to new Nokia switches [08:14:13] PROBLEM - Host hcaptcha-proxy5003 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:13] PROBLEM - Host hcaptcha-proxy5004 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:13] PROBLEM - Host install5004 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:17] 06SRE, 10SRE-swift-storage, 05Goal, 06Machine-Learning-Team (Q1 FY2026-27), 07OKR-Work: Provision and wire object storage for TTS v1 audio artifacts: S3 sink, Swift bucket, and credentials - https://phabricator.wikimedia.org/T432944#12255460 (10isarantopoulos) 05Open→03Declined For the purpose of... [08:14:17] PROBLEM - Host prometheus5003 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:17] PROBLEM - Host netflow5003 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:17] PROBLEM - Host ncredir5004 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:17] PROBLEM - Host tcp-proxy5003 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:17] PROBLEM - Host ncredir5003 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:18] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255465 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b884740c-078e-4c8d-a90b-e1f10dcdef27) set by cmooney@cumin1003 fo... [08:14:22] bah Wikibase issue Wikimedia\Rdbms\Platform\SQLPlatform::makeList: empty input for field page_id [08:14:23] PROBLEM - Host asw1-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [08:14:25] PROBLEM - Host doh5004 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:25] PROBLEM - Host doh5003 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:27] PROBLEM - Host durum5004 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:27] PROBLEM - Host durum5003 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:27] PROBLEM - Host bast5005 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:27] PROBLEM - Host tcp-proxy5004 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:30] and we lost eqsin [08:14:39] PROBLEM - Host mr1-eqsin IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [08:14:47] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on 13 hosts with reason: site migration to new Nokia switches [08:14:56] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255466 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=a5b20745-166d-4473-bbfe-1ba217405d16) set by cmooney@cumin1003 fo... [08:14:59] 06SRE: Improve reuse-parts story for standard recipes - https://phabricator.wikimedia.org/T436058#12255467 (10fgiunchedi) [08:15:19] PROBLEM - Host ps1-604-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [08:15:23] PROBLEM - Host ps1-603-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [08:15:35] PROBLEM - VRRP status on cr3-eqsin is CRITICAL: VRRP CRITICAL - 4 inconsistent interfaces, 0 misconfigured interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23VRRP_status [08:16:36] !log cmooney@cumin1003 DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 5:00:00 on 12 hosts with reason: site migration to new Nokia switches [08:17:38] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on 12 hosts with reason: site migration to new Nokia switches [08:17:49] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255490 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=af2e3dff-6866-4ea4-a3b1-818b0b29d804) set by cmooney@cumin1003 fo... [08:17:58] 06SRE: Improve reuse-parts story for standard recipes - https://phabricator.wikimedia.org/T436058#12255491 (10MoritzMuehlenhoff) Let's maybe go with partman/reuse-srv.cfg? All the standard recipes with reuse only retain /srv (and also the custom ones not in scope), and reuse-parts.cfg is quite vague. [08:18:04] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12255492 (10Fabfur) Thanks for the updates @Cyberpower678! Could you investigate if this is still an issue? If someone from archive.org wants to directly reach us (also mail ch... [08:19:38] I am going to rollback Special:EntityUsage is broken, I have let WMDE know on the shared Slack channel [08:19:47] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet [08:20:08] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329503 (https://phabricator.wikimedia.org/T430836) [08:20:11] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by hashar@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329503 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [08:20:59] !log Rolling back group 1 from 1.47.0-wmf.17 to 1.47.0-wmf.16 due to Special:EntityUsage being broken on Wikibase client wikis # T436060 T430836 [08:21:06] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:21:07] T436060: InvalidArgumentException: Wikimedia\Rdbms\Platform\SQLPlatform::makeList: empty input for field page_id - https://phabricator.wikimedia.org/T436060 [08:21:07] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [08:21:53] 06SRE: Improve reuse /srv story for standard recipes - https://phabricator.wikimedia.org/T436058#12255511 (10fgiunchedi) [08:23:11] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329503 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [08:24:29] (03CR) 10Ladsgroup: [C:03+1] lists: Double the amount of UWSGI processes [puppet] - 10https://gerrit.wikimedia.org/r/1329418 (https://phabricator.wikimedia.org/T436047) (owner: 10Muehlenhoff) [08:26:51] (03PS1) 10Dpogorzelski: kserve: add LLMInferenceService support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329506 (https://phabricator.wikimedia.org/T433973) [08:27:04] (03CR) 10CI reject: [V:04-1] kserve: add LLMInferenceService support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329506 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [08:27:04] (03CR) 10Federico Ceratto: [C:03+2] site.pp, db-test1003.yaml: Decommission db-test1003 [puppet] - 10https://gerrit.wikimedia.org/r/1329273 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [08:27:36] (03PS2) 10Dpogorzelski: kserve: add LLMInferenceService support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329506 (https://phabricator.wikimedia.org/T433973) [08:27:48] (03CR) 10CI reject: [V:04-1] kserve: add LLMInferenceService support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329506 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [08:28:29] !log fceratto@cumin1003 Removing db-test1003 from zarcillo T435912 [08:28:34] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) [08:28:35] T435912: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912 [08:28:40] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12255526 (10ops-monitoring-bot) db-test1003 has been deleted from zarcillo [08:28:42] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12255528 (10ops-monitoring-bot) db-test1003 has been decommissioned by Data Persistence [08:28:46] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12255529 (10ops-monitoring-bot) a:05FCeratto-WMF→03None This host is ready for DC-Ops to decommission [08:28:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:29:44] FIRING: RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:29:46] (03CR) 10Slyngshede: [C:03+2] IDP: CAS version 7.8.3.2 [dns] - 10https://gerrit.wikimedia.org/r/1329424 (owner: 10Slyngshede) [08:30:08] !log slyngshede@dns1004 START - running authdns-update [08:30:13] !log hashar@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.17 refs T430836 [08:30:18] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [08:30:48] !log Updating IDP/CAS-SSO to version 3.7.8.2 [08:30:50] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:31:14] (03CR) 10Jelto: [V:03+1 C:03+2] lists: Double the amount of UWSGI processes [puppet] - 10https://gerrit.wikimedia.org/r/1329418 (https://phabricator.wikimedia.org/T436047) (owner: 10Muehlenhoff) [08:31:52] (03CR) 10Tiziano Fogli: [C:03+2] prometheus: scrape the CI feedback time exporter [puppet] - 10https://gerrit.wikimedia.org/r/1329430 (https://phabricator.wikimedia.org/T435814) (owner: 10Phedenskog) [08:32:21] PROBLEM - OSPF status on cr1-codfw is CRITICAL: OSPFv2: 7/8 UP : OSPFv3: 7/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [08:32:23] PROBLEM - OSPF status on cr1-eqiad is CRITICAL: OSPFv2: 6/7 UP : OSPFv3: 6/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [08:32:28] !log slyngshede@dns1004 FAIL - running authdns-update [08:33:39] FIRING: CoreBGPDown: Core BGP session down between cr1-codfw and cr2-eqsin (103.102.166.149) - group Confed_eqsin - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqsin&var-bgp_neighbor=cr2-eqsin - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [08:34:10] FIRING: [2x] BFDdown: BFD session down between cr1-codfw and 103.102.166.149 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:34:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:37:05] PROBLEM - mailman3-web on lists1004 is CRITICAL: PROCS CRITICAL: 25 processes with UID = 33 (www-data), regex args /usr/bin/uwsgi https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [08:39:10] FIRING: [3x] BFDdown: BFD session down between cr1-codfw and 103.102.166.149 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [08:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:42:08] !log cmooney@cumin1003 START - Cookbook sre.metamonitoring.downtime Downtime for 4:00:00 of prometheus/deadmanswitchnotified, prometheus/deadmanswitchonamdb, prometheus/extmon on 2 host(s) with reason: eqsin site rebuild [08:42:13] !log cmooney@cumin1003 END (PASS) - Cookbook sre.metamonitoring.downtime (exit_code=0) Downtime for 4:00:00 of prometheus/deadmanswitchnotified, prometheus/deadmanswitchonamdb, prometheus/extmon on 2 host(s) with reason: eqsin site rebuild [08:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:45:05] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [08:49:25] !log installing openssl security updates on trixie [08:49:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:50:29] (03CR) 10Kevin Bazira: [C:03+1] ml-services: Deploy latest logo-detection model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329292 (https://phabricator.wikimedia.org/T435946) (owner: 10Gkyziridis) [08:57:23] PROBLEM - Host mr1-eqsin.oob IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [08:57:42] (03PS1) 10Awight: Multiple reverts for entity usage [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329510 (https://phabricator.wikimedia.org/T436060) [08:59:51] 10ops-eqiad, 06DC-Ops: Alert for device ps1-f4-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T436064 (10phaultfinder) 03NEW [09:00:38] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest changes in liftwing-openapi-server [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327001 (owner: 10Gkyziridis) [09:00:53] (03PS1) 10Awight: Revert "Removes unused join and refactors helper method." [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329511 (https://phabricator.wikimedia.org/T436060) [09:01:18] (03Abandoned) 10Awight: Multiple reverts for entity usage [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329510 (https://phabricator.wikimedia.org/T436060) (owner: 10Awight) [09:01:42] (03CR) 10Awight: [C:03+1] Revert "Removes unused join and refactors helper method." [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329511 (https://phabricator.wikimedia.org/T436060) (owner: 10Awight) [09:01:51] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm [09:02:25] RECOVERY - Host mr1-eqsin.oob IPv6 is UP: PING OK - Packet loss = 0%, RTA = 224.55 ms [09:02:48] (03Merged) 10jenkins-bot: ml-services: Deploy latest changes in liftwing-openapi-server [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327001 (owner: 10Gkyziridis) [09:03:20] (03PS1) 10Awight: Revert "This splits the joined `wbc_entity_usage` queries into individual queries because of X1 migration." [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329512 (https://phabricator.wikimedia.org/T436060) [09:04:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [09:06:38] (03PS3) 10Reedy: Periodic Jobs: Add purge_expired_recovery_codes for SUL and non SUL wikis [puppet] - 10https://gerrit.wikimedia.org/r/1325573 (https://phabricator.wikimedia.org/T422922) [09:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [09:07:07] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . [09:07:31] !log gkyziridis@deploy1003 helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . [09:09:39] !log gkyziridis@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . [09:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:16:59] (03CR) 10TrainBranchBot: [C:03+2] "Approved by hashar@deploy1003 using scap backport" [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329511 (https://phabricator.wikimedia.org/T436060) (owner: 10Awight) [09:17:00] (03CR) 10TrainBranchBot: [C:03+2] "Approved by hashar@deploy1003 using scap backport" [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329512 (https://phabricator.wikimedia.org/T436060) (owner: 10Awight) [09:18:33] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest logo-detection model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329292 (https://phabricator.wikimedia.org/T435946) (owner: 10Gkyziridis) [09:19:07] I am backporting a couple patches for Wikibase and will then roll the train again [09:19:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [09:20:50] (03Merged) 10jenkins-bot: ml-services: Deploy latest logo-detection model version on staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329292 (https://phabricator.wikimedia.org/T435946) (owner: 10Gkyziridis) [09:22:32] (03Merged) 10jenkins-bot: Revert "Removes unused join and refactors helper method." [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329511 (https://phabricator.wikimedia.org/T436060) (owner: 10Awight) [09:22:43] !log gkyziridis@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . [09:23:56] (03Merged) 10jenkins-bot: Revert "This splits the joined `wbc_entity_usage` queries into individual queries because of X1 migration." [extensions/Wikibase] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329512 (https://phabricator.wikimedia.org/T436060) (owner: 10Awight) [09:24:23] !log hashar@deploy1003 Started scap sync-world: Backport for [[gerrit:1329511|Revert "Removes unused join and refactors helper method." (T436060)]], [[gerrit:1329512|Revert "This splits the joined `wbc_entity_usage` queries into individual queries because of X1 migration." (T436060)]] [09:24:28] T436060: InvalidArgumentException: Wikimedia\Rdbms\Platform\SQLPlatform::makeList: empty input for field page_id - https://phabricator.wikimedia.org/T436060 [09:26:44] !log brouberol@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1010.eqiad.wmnet with OS bookworm [09:26:59] !log hashar@deploy1003 awight, hashar: Backport for [[gerrit:1329511|Revert "Removes unused join and refactors helper method." (T436060)]], [[gerrit:1329512|Revert "This splits the joined `wbc_entity_usage` queries into individual queries because of X1 migration." (T436060)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [09:27:08] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm [09:28:16] !log hashar@deploy1003 awight, hashar: Continuing with deployment [09:30:40] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [09:30:44] (03PS1) 10Gkyziridis: ml-services: Deploy latest logo-detection model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329515 (https://phabricator.wikimedia.org/T435946) [09:30:44] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) [09:31:20] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [09:31:25] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db-test2002.codfw.wmnet [09:32:36] !log hashar@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329511|Revert "Removes unused join and refactors helper method." (T436060)]], [[gerrit:1329512|Revert "This splits the joined `wbc_entity_usage` queries into individual queries because of X1 migration." (T436060)]] (duration: 08m 13s) [09:32:41] T436060: InvalidArgumentException: Wikimedia\Rdbms\Platform\SQLPlatform::makeList: empty input for field page_id - https://phabricator.wikimedia.org/T436060 [09:34:49] fceratto@cumin1003 decommission (PID 334318) is awaiting input [09:39:01] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329518 (https://phabricator.wikimedia.org/T430836) [09:39:04] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by hashar@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329518 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [09:39:32] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.hosts.decommission (exit_code=99) for hosts db-test2002.codfw.wmnet [09:39:34] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) [09:39:37] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [09:39:41] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db-test2002.codfw.wmnet [09:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:40:22] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329518 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [09:40:58] (03PS1) 10Andrea Denisse: grafana: Disable loading of not installed core plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) [09:40:58] (03CR) 10Andrea Denisse: "None of these plugins are isntalled but Grafana considers them to be "core plugins" and tries to load them everytime, when it can't find t" [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [09:41:46] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.hosts.decommission (exit_code=99) for hosts db-test2002.codfw.wmnet [09:41:47] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) [09:42:16] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on pc2022.codfw.wmnet,pc1022.eqiad.wmnet with reason: Working on pc2 truncation [09:42:17] (03CR) 10MSantos: [C:03+1] switch testwiki to use parsoid [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [09:42:27] (03CR) 10CI reject: [V:04-1] switch testwiki to use parsoid [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [09:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:43:59] (03CR) 10Andrea Denisse: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9319/co" [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [09:47:29] !log hashar@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.17 refs T430836 [09:47:35] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [09:54:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:55:05] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool pc1022: Pool back [09:55:05] !log marostegui@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1022: Pool back [09:55:12] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool pc1022: Pool back pc2 [09:55:12] !log marostegui@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool pc1022: Pool back pc2 [09:55:29] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool pc1022: Pool back pc2 [09:55:29] !log marostegui@cumin1003 START - Cookbook sre.mysql.parsercache [09:55:35] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0) [09:55:36] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Pool back pc2 [09:58:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:59:20] (03PS3) 10Abijeet Patro: ArticleGuidance: Configure feedback links to local talk pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329314 (https://phabricator.wikimedia.org/T433483) [09:59:33] (03CR) 10Abijeet Patro: ArticleGuidance: Configure feedback links to local talk pages (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329314 (https://phabricator.wikimedia.org/T433483) (owner: 10Abijeet Patro) [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1000) [10:02:22] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [10:02:26] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db-test2002.codfw.wmnet [10:02:46] I have rolled 1.47.0-wmf.17 to group 1 wikis and it seems quite. I am going to have a lunch break and will do some log triage when I am done [10:03:07] (03CR) 10Peterxy12: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [10:03:30] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [10:04:38] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [10:05:02] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 26 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324031 (https://phabricator.wikimedia.org/T429122) (owner: 10Abijeet Patro) [10:05:43] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) [10:06:41] (03PS4) 10Andrea Denisse: grafana: Disable loading of not installed core plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) [10:06:41] (03CR) 10Andrea Denisse: "PCC results: https://puppet-compiler.wmflabs.org/output/1329517/9320/" [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [10:07:38] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [10:09:45] (03PS4) 10Peterxy12: switch testwiki to use parsoid [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) [10:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:13:09] fceratto@cumin1003 decommission (PID 358625) is awaiting input [10:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:13:56] (03PS2) 10Mvolz: zotero: update version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329240 (https://phabricator.wikimedia.org/T435179) [10:14:16] (03CR) 10Blake: [C:03+2] mw-*: switch to envoy 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329325 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:15:52] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1010.eqiad.wmnet with OS bookworm [10:16:47] (03Merged) 10jenkins-bot: mw-*: switch to envoy 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329325 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:17:36] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-wikifunctions: apply [10:18:26] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-wikifunctions: apply [10:18:26] what's up with dns5004.wikimedia.org ? netbox updates are failing: [10:18:29] ssh: connect to host dns5003.wikimedia.org port 22: Connection timed out [10:18:32] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-wikifunctions: apply [10:18:48] RECOVERY - Host ps1-603-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.86 ms [10:18:48] RECOVERY - Host ps1-604-eqsin is UP: PING OK - Packet loss = 0%, RTA = 210.64 ms [10:18:50] RECOVERY - Host asw1-eqsin is UP: PING OK - Packet loss = 0%, RTA = 210.40 ms [10:19:04] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-wikifunctions: apply [10:19:20] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-misc: apply [10:19:22] RECOVERY - OSPF status on cr1-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [10:19:22] RECOVERY - OSPF status on cr1-eqiad is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [10:19:29] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 26 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [10:19:43] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-misc: apply [10:19:48] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-misc: apply [10:20:09] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-misc: apply [10:20:25] (03CR) 10Andrea Denisse: "Looking at the docs another option for this *could be* to disable all preinstall plugins [1] and so we only install the plugins we need th" [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [10:20:57] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply [10:21:58] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply [10:22:03] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-jobrunner: apply [10:22:59] (03CR) 10Ozge: [C:03+1] ml-services: Deploy latest logo-detection model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329515 (https://phabricator.wikimedia.org/T435946) (owner: 10Gkyziridis) [10:23:04] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: apply [10:23:14] fceratto@cumin1003 decommission (PID 358625) is awaiting input [10:23:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-codfw and cr2-eqsin (103.102.166.149) - group Confed_eqsin - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqsin&var-bgp_neighbor=cr2-eqsin - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [10:24:10] RESOLVED: [2x] BFDdown: BFD session down between cr1-codfw and 103.102.166.149 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:24:53] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [10:24:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:25:00] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-api-ext: apply [10:26:21] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply [10:27:23] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:27:46] (03PS1) 10Muehlenhoff: Extend Cumin aliases for URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1329521 (https://phabricator.wikimedia.org/T429175) [10:28:27] fceratto@cumin1003 decommission (PID 358625) is awaiting input [10:28:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:28:58] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply [10:29:12] (03PS1) 10AOkoth: aptrepo: bump gitlab-ce version [puppet] - 10https://gerrit.wikimedia.org/r/1329522 (https://phabricator.wikimedia.org/T436069) [10:30:04] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply [10:30:33] (03PS1) 10Muehlenhoff: Add cookbook to roll-restart/reboot URL downloaders [cookbooks] - 10https://gerrit.wikimedia.org/r/1329523 (https://phabricator.wikimedia.org/T429175) [10:31:59] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-api-int: apply [10:32:33] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest logo-detection model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329515 (https://phabricator.wikimedia.org/T435946) (owner: 10Gkyziridis) [10:32:58] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-api-int: apply [10:34:21] (03CR) 10Jelto: aptrepo: bump gitlab-ce version (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329522 (https://phabricator.wikimedia.org/T436069) (owner: 10AOkoth) [10:34:39] (03PS1) 10Marostegui: installserver: Do not format db1269 [puppet] - 10https://gerrit.wikimedia.org/r/1329525 [10:34:51] (03Merged) 10jenkins-bot: ml-services: Deploy latest logo-detection model version on prod. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329515 (https://phabricator.wikimedia.org/T435946) (owner: 10Gkyziridis) [10:35:22] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-api-int: apply [10:35:31] (03PS2) 10AOkoth: aptrepo: bump gitlab-ce version [puppet] - 10https://gerrit.wikimedia.org/r/1329522 (https://phabricator.wikimedia.org/T436069) [10:35:54] (03CR) 10AOkoth: aptrepo: bump gitlab-ce version (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329522 (https://phabricator.wikimedia.org/T436069) (owner: 10AOkoth) [10:36:23] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-api-int: apply [10:37:07] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' . [10:37:12] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-web: apply [10:37:18] !log gkyziridis@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' . [10:37:29] (03CR) 10Marostegui: [C:03+2] installserver: Do not format db1269 [puppet] - 10https://gerrit.wikimedia.org/r/1329525 (owner: 10Marostegui) [10:38:26] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-web: apply [10:38:47] !log slyngshede@puppetserver1001 conftool action : set/pooled=no; selector: name=dns5003.* [10:38:58] !log slyngshede@puppetserver1001 conftool action : set/pooled=no; selector: name=dns5004.* [10:39:26] !log slyngshede@dns1004 START - running authdns-update [10:39:53] (03CR) 10Jelto: [C:03+1] "lgtm thank you" [puppet] - 10https://gerrit.wikimedia.org/r/1329522 (https://phabricator.wikimedia.org/T436069) (owner: 10AOkoth) [10:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:40:33] 06SRE, 10Wikimedia-Mailing-lists: Postorius: "Message could not be found" error preventing mailing list moderation - https://phabricator.wikimedia.org/T435893#12255972 (10FastLizard4) Discarding messages is indeed now working again for me. I did retry a few times over the course of a couple hours after filing... [10:40:56] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255975 (10SLyngshede-WMF) Forgot the DNS hosts: ` slyngshede@puppetserver1001:~$ sudo -i confctl select name=dns5003.* set/pooled=no The s... [10:41:35] (03PS1) 10Majavah: P:wmcs::novaproxy: Fix ordering of wmflabs redirect [puppet] - 10https://gerrit.wikimedia.org/r/1329527 (https://phabricator.wikimedia.org/T429930) [10:41:38] !log uploaded cortobot-1.2.0 to apt1002 [10:41:41] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:41:43] !log slyngshede@dns1004 END - running authdns-update [10:42:22] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-web: apply [10:43:02] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [10:43:18] 06SRE, 06Traffic, 07Wikimedia-Incident: Wikipedia and Wiktionary pages returning "upstream connect error or disconnect/reset before headers" - https://phabricator.wikimedia.org/T436004#12255983 (10toddbradley) Hi, apparently this issue might likely still persists somehow? Apparently I cannot access Wikipedia... [10:43:34] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-web: apply [10:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:44:24] (03CR) 10FNegri: [C:03+1] P:wmcs::novaproxy: Fix ordering of wmflabs redirect [puppet] - 10https://gerrit.wikimedia.org/r/1329527 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [10:44:50] (03CR) 10AOkoth: [C:03+2] aptrepo: bump gitlab-ce version [puppet] - 10https://gerrit.wikimedia.org/r/1329522 (https://phabricator.wikimedia.org/T436069) (owner: 10AOkoth) [10:45:09] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Fix ordering of wmflabs redirect [puppet] - 10https://gerrit.wikimedia.org/r/1329527 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [10:45:13] !log cmooney@cumin1003 END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) [10:45:57] (03CR) 10Muehlenhoff: [C:03+2] Move the build of the base images from build2001 to build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1328554 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [10:49:14] PROBLEM - Host 2001:df2:e500:1:103:102:166:10 is DOWN: PING CRITICAL - Packet loss = 100% [10:50:47] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12256016 (10AlexisJazz) >>! In T435743#12255492, @Fabfur wrote: > Thanks for the updates @Cyberpower678! > Could you investigate if this is still an issue? I'm still seeing th... [10:50:59] 06SRE, 06Traffic, 07Wikimedia-Incident: Wikipedia and Wiktionary pages returning "upstream connect error or disconnect/reset before headers" - https://phabricator.wikimedia.org/T436004#12256018 (10Fabfur) Hi @toddbradley , thanks for the report. We've depooled the Singapore DC some hours ago due to a schedul... [10:54:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:56:43] (03PS1) 10Marostegui: mariadb: Update notes [puppet] - 10https://gerrit.wikimedia.org/r/1329529 [10:57:31] (03CR) 10Marostegui: [C:03+2] mariadb: Update notes [puppet] - 10https://gerrit.wikimedia.org/r/1329529 (owner: 10Marostegui) [10:57:42] (03CR) 10Marostegui: [C:03+2] "This is a noop" [puppet] - 10https://gerrit.wikimedia.org/r/1329529 (owner: 10Marostegui) [10:58:07] 06SRE, 10LDAP-Access-Requests: Grant Access to wmf, for WMF staff/contractors nda group for gsduser - https://phabricator.wikimedia.org/T435852#12256072 (10Aklapper) 05Resolved→03Declined a:05Eevans→03None Setting status to declined as no work was done in this ticket [10:58:48] (03PS1) 10Mszwarc: UserGroupsSpecialPage: Check in_array in strict mode [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1329530 (https://phabricator.wikimedia.org/T435907) [10:58:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:59:08] (03PS1) 10Mszwarc: UserGroupsSpecialPage: Check in_array in strict mode [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329531 (https://phabricator.wikimedia.org/T435907) [11:00:04] mvolz: Time to snap out of that daydream and deploy Services – Citoid / Zotero. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1100). [11:00:32] (03CR) 10Mvolz: [C:03+2] zotero: update version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329240 (https://phabricator.wikimedia.org/T435179) (owner: 10Mvolz) [11:02:33] (03Merged) 10jenkins-bot: zotero: update version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329240 (https://phabricator.wikimedia.org/T435179) (owner: 10Mvolz) [11:04:20] !log Upgrading grafana in eqiad - T435816 [11:04:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:04:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:05:27] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/zotero: apply [11:05:47] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/zotero: apply [11:08:47] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/zotero: apply [11:08:54] PROBLEM - Check correctness of the icinga configuration on alert1002 is CRITICAL: Icinga configuration contains errors https://wikitech.wikimedia.org/wiki/Icinga [11:09:14] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/zotero: apply [11:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:10:33] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/zotero: apply [11:11:01] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/zotero: apply [11:11:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:12:26] PROBLEM - Host 2001:df2:e500:2:103:102:166:36 is DOWN: PING CRITICAL - Packet loss = 100% [11:13:44] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Migrate remaining container build/report steps from build2001 to build2004 - https://phabricator.wikimedia.org/T417389#12256153 (10MoritzMuehlenhoff) [11:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:15:06] RECOVERY - Postfix SMTP on crm2001 is OK: OK - Certificate crm2001.codfw.wmnet will expire on Wed 23 Sep 2026 10:40:00 AM GMT +0000. https://wikitech.wikimedia.org/wiki/Mail%23Troubleshooting [11:16:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:23:08] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12256217 (10MoritzMuehlenhoff) [11:24:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:26:21] !log jmm@cumin2003 START - Cookbook sre.hosts.reimage for host ganeti3005.esams.wmnet with OS bookworm [11:26:32] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12256218 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin2003 for host ganeti3005.esams.wmnet with OS bookworm [11:28:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:29:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:34:55] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12256244 (10MoritzMuehlenhoff) >>! In T435873#12252130, @bking wrote: > @MoritzMuehlenhoff It seems like EQIAD only has 10G... [11:36:43] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T435989#12256250 (10MoritzMuehlenhoff) There's no need for 16G of RAM, 8G is perfectly fine and with ample headroom. Rest loo... [11:37:10] jouncebot: nowandnext [11:37:10] For the next 0 hour(s) and 22 minute(s): Services – Citoid / Zotero (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1100) [11:37:10] In 1 hour(s) and 22 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1300) [11:37:32] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1329337 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [11:37:58] Mvolz: I have a patch that I'd like to backport. Could you please let me know when/if you're done with deploying Citoid / Zotero? [11:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:42:45] (03CR) 10Hnowlan: grafana: Disable loading of not installed core plugins (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [11:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:49:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329531 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [11:49:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1329530 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [11:49:05] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test2002.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [11:50:46] !log jmm@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti3005.esams.wmnet with reason: host reimage [11:52:08] fceratto@cumin1003 decommission (PID 358625) is awaiting input [11:53:58] (03Merged) 10jenkins-bot: UserGroupsSpecialPage: Check in_array in strict mode [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329531 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [11:54:28] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti3005.esams.wmnet with reason: host reimage [11:55:03] (03Merged) 10jenkins-bot: UserGroupsSpecialPage: Check in_array in strict mode [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1329530 (https://phabricator.wikimedia.org/T435907) (owner: 10Mszwarc) [11:55:25] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1329531|UserGroupsSpecialPage: Check in_array in strict mode (T435907)]], [[gerrit:1329530|UserGroupsSpecialPage: Check in_array in strict mode (T435907)]] [11:58:01] !log mszwarc@deploy1003 mszwarc: Backport for [[gerrit:1329531|UserGroupsSpecialPage: Check in_array in strict mode (T435907)]], [[gerrit:1329530|UserGroupsSpecialPage: Check in_array in strict mode (T435907)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:58:33] !log mszwarc@deploy1003 mszwarc: Continuing with deployment [12:00:03] (03CR) 10JMeybohm: [C:03+1] docker: Add team labels to docker-reporter jobs [puppet] - 10https://gerrit.wikimedia.org/r/1329402 (owner: 10RLazarus) [12:00:24] (03PS1) 10Majavah: Remove various references to the puppetmaster class [puppet] - 10https://gerrit.wikimedia.org/r/1329540 [12:01:00] RECOVERY - Host 2001:df2:e500:1:103:102:166:10 is UP: PING OK - Packet loss = 0%, RTA = 210.09 ms [12:01:16] RECOVERY - Host 2001:df2:e500:2:103:102:166:36 is UP: PING OK - Packet loss = 0%, RTA = 210.27 ms [12:02:23] (03PS2) 10Majavah: Remove various references to the puppetmaster class [puppet] - 10https://gerrit.wikimedia.org/r/1329540 [12:02:53] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329531|UserGroupsSpecialPage: Check in_array in strict mode (T435907)]], [[gerrit:1329530|UserGroupsSpecialPage: Check in_array in strict mode (T435907)]] (duration: 07m 27s) [12:06:38] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (NOOP 17): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9322/consol" [puppet] - 10https://gerrit.wikimedia.org/r/1329540 (owner: 10Majavah) [12:08:09] (03PS1) 10Clément Goubert: P:chartmuseum: Update every minute [puppet] - 10https://gerrit.wikimedia.org/r/1329544 (https://phabricator.wikimedia.org/T435987) [12:08:50] (03PS2) 10Lerickson: Increase WDQS backend pod storage to 900GiB. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329388 (https://phabricator.wikimedia.org/T436023) [12:09:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:10:42] (03PS2) 10Lerickson: Update memory settings for wdqs-next. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329313 (https://phabricator.wikimedia.org/T434854) [12:12:20] (03CR) 10JMeybohm: "Fine by me. Keep in mind though that this runs on both chartmuseum hosts which makes it effectively "more then once per minute"" [puppet] - 10https://gerrit.wikimedia.org/r/1329544 (https://phabricator.wikimedia.org/T435987) (owner: 10Clément Goubert) [12:12:52] (03PS3) 10Abijeet Patro: ULS: Remove wgULSLanguageSelectorV2Enabled config [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324031 (https://phabricator.wikimedia.org/T429122) [12:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:15:43] (03PS6) 10Clément Goubert: mediawiki: enable forward of fatal metrics to statsd exporter [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) (owner: 10Cwhite) [12:15:46] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) (owner: 10Cwhite) [12:16:23] (03PS2) 10Clément Goubert: P:chartmuseum: Update every minute [puppet] - 10https://gerrit.wikimedia.org/r/1329544 (https://phabricator.wikimedia.org/T435987) [12:17:00] (03CR) 10Clément Goubert: "Yeah that's fine, we just want it to not cause undue waiting for operators after merging." [puppet] - 10https://gerrit.wikimedia.org/r/1329544 (https://phabricator.wikimedia.org/T435987) (owner: 10Clément Goubert) [12:19:35] (03CR) 10Sbisson: [C:03+1] ArticleGuidance: Configure feedback links to local talk pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329314 (https://phabricator.wikimedia.org/T433483) (owner: 10Abijeet Patro) [12:19:37] (03PS4) 10Majavah: P:wmcs::novaproxy: Stop writing data to Redis [puppet] - 10https://gerrit.wikimedia.org/r/1329230 (https://phabricator.wikimedia.org/T429930) [12:19:37] (03PS4) 10Majavah: P:wmcs::novaproxy: Undeploy Redis instance [puppet] - 10https://gerrit.wikimedia.org/r/1329231 (https://phabricator.wikimedia.org/T429930) [12:19:37] (03PS1) 10Majavah: dynamicproxy: Remove custom log rotation timer [puppet] - 10https://gerrit.wikimedia.org/r/1329547 (https://phabricator.wikimedia.org/T429930) [12:19:39] (03PS1) 10Majavah: P:wmcs::novaproxy: Remove active proxy concept [puppet] - 10https://gerrit.wikimedia.org/r/1329548 (https://phabricator.wikimedia.org/T429930) [12:19:42] (03PS1) 10Majavah: P:wmcs::novaproxy: Migrate to firewall::service [puppet] - 10https://gerrit.wikimedia.org/r/1329549 (https://phabricator.wikimedia.org/T429930) [12:24:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:26:40] (03CR) 10Daimona Eaytoy: "Note this is a config patch, it shouldn't be +2ed ahead of deployment (which needs to wait for next week's train for I2e8ea1bd064e351 to b" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326400 (https://phabricator.wikimedia.org/T429510) (owner: 10Daimona Eaytoy) [12:31:39] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti3005.esams.wmnet with OS bookworm [12:31:47] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12256482 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin2003 for host ganeti3005.esams.wmnet with OS bookworm completed: - ganeti3005 (**WARN**) - Downtim... [12:35:34] (03PS2) 10Brouberol: dse-k8s::worker: enable p2p pulling of large images [puppet] - 10https://gerrit.wikimedia.org/r/1329551 (https://phabricator.wikimedia.org/T436089) [12:38:41] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T436069 [12:39:25] Hi folks, I need to run a write query in production to restore a bunch of accidentally deleted event participants. The full query is in T436088#12256503. May I go ahead? [12:39:25] T436088: Restore deleted participants to https://meta.wikimedia.org/wiki/Special:EventDetails/4242 - https://phabricator.wikimedia.org/T436088 [12:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:44:26] topranks: moritzm: as oncall, see Daimona's question above [12:45:05] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:46:46] Daimona, taavi: I have absolutely idea how impactful that would be, but if we did it before it seems fine [12:47:36] It's a trivial UPDATE query that affects 316 rows, it shouldn't have any impact. I did it before on fewer (61) rows, but the effects should be the same [12:47:57] Daimona: ok, go ahead please [12:48:02] but if there is some documented procedure to notify oncallers then I fail to see it's usefulness since 90% of us will also not be able to tell. this rather sounds like something that should be clarified specifically with DBA [12:48:39] Thank you, I will run it :) And I'm not sure about procedures unfortunately :/ [12:49:14] all fine, just go ahead. you absolutely did the right thing to doublechecking for sure [12:49:30] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T436069 [12:49:45] !log Running query from T436088#12256503 in x1.wikishared [12:49:49] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:49:49] T436088: Restore deleted participants to https://meta.wikimedia.org/wiki/Special:EventDetails/4242 - https://phabricator.wikimedia.org/T436088 [12:50:41] (03PS1) 10Muehlenhoff: Move weekly build of production images to build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1329554 (https://phabricator.wikimedia.org/T417389) [12:53:16] (03CR) 10Filippo Giunchedi: [C:03+1] "Very nice" [puppet] - 10https://gerrit.wikimedia.org/r/1329230 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:53:31] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Undeploy Redis instance [puppet] - 10https://gerrit.wikimedia.org/r/1329231 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:53:45] (03CR) 10Filippo Giunchedi: [C:03+1] dynamicproxy: Remove custom log rotation timer [puppet] - 10https://gerrit.wikimedia.org/r/1329547 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:54:26] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Stop writing data to Redis [puppet] - 10https://gerrit.wikimedia.org/r/1329230 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [12:55:27] (03CR) 10Majavah: [C:03+1] labstore: clean up traffic_shaping [puppet] - 10https://gerrit.wikimedia.org/r/1327853 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [12:56:33] (03CR) 10Majavah: [C:03+1] wmcs: remove support for clouddumps client symlinks [puppet] - 10https://gerrit.wikimedia.org/r/1327854 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [12:57:21] (03CR) 10Majavah: [C:03+1] dumps: remove production support for non-lb mounts [puppet] - 10https://gerrit.wikimedia.org/r/1327855 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [12:58:33] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329554 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [12:59:37] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12256588 (10Fabfur) Can't replicate, I've just archive [[ https://web.archive.org/web/20260826125613/https://en.wikipedia.org/wiki/John_of_Sterngassen | this page ]] (new page,... [12:59:53] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Undeploy Redis instance [puppet] - 10https://gerrit.wikimedia.org/r/1329231 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:00:05] urbanecm and TheresNoTime: #bothumor I � Unicode. All rise for UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1300). [13:00:05] Peterxy and abijeet: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:29] (03CR) 10Majavah: [C:03+2] dynamicproxy: Remove custom log rotation timer [puppet] - 10https://gerrit.wikimedia.org/r/1329547 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:01:05] !log brouberol@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1010.eqiad.wmnet with OS bookworm [13:01:31] !incidents [13:01:32] 8307 (UNACKED) Alertmanager: Prometheus Watchdog Hook Processing check is still DOWN [13:01:32] 8305 (RESOLVED) Alertmanager-active: Prometheus Watchdog Alerts Presence check is still DOWN [13:01:32] 8304 (RESOLVED) Alertmanager-passive: Prometheus Watchdog Alerts Presence check is still DOWN [13:01:33] 8306 (RESOLVED) Alertmanager: Prometheus Watchdog Hook Processing check is still DOWN [13:01:33] 8303 (RESOLVED) ATSBackendErrorsHigh cache_text sre (lists.discovery.wmnet eqiad) [13:01:33] 8302 (RESOLVED) Large wiki outage [13:01:33] 8295 (RESOLVED) pc1021 (paged)/MariaDB read only pc1 (paged) [13:01:34] 8298 (RESOLVED) pc1024 (paged)/MariaDB read only pc4 (paged) [13:01:34] 8301 (RESOLVED) [8x] ATSBackendErrorsHigh cache_text sre () [13:01:35] 8297 (RESOLVED) PHPFPMTooBusy sre (mw-web main eqiad) [13:01:35] 8294 (RESOLVED) [7x] ProbeDown sre (probes/service) [13:01:36] 8300 (RESOLVED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [13:01:36] 8299 (RESOLVED) VarnishUnavailable global sre (varnish-text thanos-rule@main) [13:01:36] !ack [13:01:37] 8296 (RESOLVED) pc1015 (paged)/MariaDB read only pc5 (paged) [13:01:37] 8307 (ACKED) Alertmanager: Prometheus Watchdog Hook Processing check is still DOWN [13:02:37] (03CR) 10Majavah: [C:03+2] openstack: novastats: Update proxyleaks for singular backend object [puppet] - 10https://gerrit.wikimedia.org/r/1308063 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [13:02:49] I believe the downtime needs to be extended for the duration of the planned work.. moritzm topranks [13:03:47] tappof: yeah I was wholly too optimistic sorry [13:04:12] !log cmooney@cumin1003 START - Cookbook sre.metamonitoring.downtime Downtime for 3:00:00 of prometheus/deadmanswitchnotified, prometheus/deadmanswitchonamdb, prometheus/extmon on 2 host(s) with reason: eqsin site rebuild [13:04:17] !log cmooney@cumin1003 END (PASS) - Cookbook sre.metamonitoring.downtime (exit_code=0) Downtime for 3:00:00 of prometheus/deadmanswitchnotified, prometheus/deadmanswitchonamdb, prometheus/extmon on 2 host(s) with reason: eqsin site rebuild [13:05:01] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 12 hosts with reason: site migration to new Nokia switches [13:05:12] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12256611 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=4da7b8be-5c0d-47c2-8661-fbe2c21e768c) set by cmooney@cumin1003 fo... [13:05:32] tappof: do I need to resolve the pages - I've run that downtime cookbook again now so I think that might auto-resolve them? [13:06:09] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: site migration to new Nokia switches [13:06:19] topranks: I could also manually resolve them in victorops? [13:06:20] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12256613 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=f7d13814-60b9-44ed-b58e-b45be17b64e4) set by cmooney@cumin1003 fo... [13:06:31] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 12 hosts with reason: site migration to new Nokia switches [13:06:39] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12256614 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=a30bcd41-744e-49bf-94f4-505e01bfb698) set by cmooney@cumin1003 fo... [13:06:40] topranks: If you've run the cookbook, it should resolve itself shortly. [13:06:49] it just did [13:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [13:07:02] (03CR) 10Majavah: [C:03+2] openstack: wmf_sink: Update for singular proxy 'backend' object [puppet] - 10https://gerrit.wikimedia.org/r/1308064 (https://phabricator.wikimedia.org/T429960) (owner: 10Majavah) [13:07:41] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T436069 [13:07:59] nice - thanks guys ! [13:08:35] urbanecm: deploying? [13:09:23] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:10:33] Peterxy seems not around. [13:10:46] I'll check with abijeet. [13:11:39] (03PS1) 10Cparle: New swift_per_container_stats_objects metrics [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) [13:12:14] (03CR) 10CI reject: [V:04-1] New swift_per_container_stats_objects metrics [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [13:12:17] (03PS2) 10Cparle: New swift_per_container_stats_objects metrics [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) [13:12:25] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test2002.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [13:12:25] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:12:26] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db-test2002.codfw.wmnet [13:13:02] (03CR) 10CI reject: [V:04-1] New swift_per_container_stats_objects metrics [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [13:13:30] (03CR) 10JMeybohm: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1329551 (https://phabricator.wikimedia.org/T436089) (owner: 10Brouberol) [13:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:14:15] hi, I'm here now. I have a patch for backport. [13:14:29] abijeet: let's deploy.. [13:14:37] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet [13:15:05] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kartik@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324031 (https://phabricator.wikimedia.org/T429122) (owner: 10Abijeet Patro) [13:15:26] fceratto@cumin1003 decommission (PID 358625) is awaiting input [13:15:40] (03PS1) 10Federico Ceratto: site.pp,db-test2002.yaml: Decommission db-test2002 [puppet] - 10https://gerrit.wikimedia.org/r/1329563 (https://phabricator.wikimedia.org/T435912) [13:15:55] kart_, thanks! [13:16:02] (03Merged) 10jenkins-bot: ULS: Remove wgULSLanguageSelectorV2Enabled config [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324031 (https://phabricator.wikimedia.org/T429122) (owner: 10Abijeet Patro) [13:16:09] !log brouberol@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1010.eqiad.wmnet with reason: host reimage [13:16:15] (03CR) 10Marostegui: [C:03+1] site.pp,db-test2002.yaml: Decommission db-test2002 [puppet] - 10https://gerrit.wikimedia.org/r/1329563 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [13:16:22] !log kartik@deploy1003 Started scap sync-world: Backport for [[gerrit:1324031|ULS: Remove wgULSLanguageSelectorV2Enabled config (T429122)]] [13:16:29] T429122: Remove new ULS beta feature - https://phabricator.wikimedia.org/T429122 [13:17:31] (03PS3) 10Cparle: New swift_per_container_stats_objects metrics [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) [13:18:24] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T436069 [13:19:00] !log kartik@deploy1003 kartik, abi: Backport for [[gerrit:1324031|ULS: Remove wgULSLanguageSelectorV2Enabled config (T429122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:19:23] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:19:37] (03CR) 10CI reject: [V:04-1] New swift_per_container_stats_objects metrics [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [13:19:46] abijeet: Please test. [13:19:49] kart_, ok [13:21:12] kart_, works well [13:21:17] !log brouberol@cumin1003 END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-presto1010.eqiad.wmnet with reason: host reimage [13:21:48] (03PS2) 10Majavah: P:wmcs::novaproxy: Remove active proxy concept [puppet] - 10https://gerrit.wikimedia.org/r/1329548 (https://phabricator.wikimedia.org/T429930) [13:21:48] (03PS2) 10Majavah: P:wmcs::novaproxy: Migrate to firewall::service [puppet] - 10https://gerrit.wikimedia.org/r/1329549 (https://phabricator.wikimedia.org/T429930) [13:22:39] cool. Deploying. [13:22:47] !log kartik@deploy1003 kartik, abi: Continuing with deployment [13:22:52] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9324/console" [puppet] - 10https://gerrit.wikimedia.org/r/1329548 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:23:12] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1329549 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:24:11] (03PS4) 10Cparle: New swift_per_container_stats_objects metrics [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) [13:24:23] FIRING: [7x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:25:35] (03PS3) 10Majavah: P:wmcs::novaproxy: Migrate to nftables::service [puppet] - 10https://gerrit.wikimedia.org/r/1329549 (https://phabricator.wikimedia.org/T429930) [13:25:56] (03PS1) 10Ladsgroup: Switch eswiki to thumb.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329564 (https://phabricator.wikimedia.org/T427465) [13:26:15] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9326/console" [puppet] - 10https://gerrit.wikimedia.org/r/1329549 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:27:20] !log kartik@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324031|ULS: Remove wgULSLanguageSelectorV2Enabled config (T429122)]] (duration: 10m 57s) [13:27:25] T429122: Remove new ULS beta feature - https://phabricator.wikimedia.org/T429122 [13:27:56] abijeet: Deployed! [13:28:05] kart_, thanks! [13:29:23] (03CR) 10Ssingh: Extend Cumin aliases for URL downloaders (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329521 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:30:48] (03CR) 10FNegri: [C:03+1] P:wmcs::novaproxy: Remove active proxy concept [puppet] - 10https://gerrit.wikimedia.org/r/1329548 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:31:19] (03CR) 10Majavah: [V:03+1 C:03+2] P:wmcs::novaproxy: Remove active proxy concept [puppet] - 10https://gerrit.wikimedia.org/r/1329548 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:31:23] (03PS2) 10Muehlenhoff: Extend Cumin aliases for URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1329521 (https://phabricator.wikimedia.org/T429175) [13:31:41] (03CR) 10Muehlenhoff: Extend Cumin aliases for URL downloaders (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329521 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:32:09] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host clouddumps1002.wikimedia.org with OS trixie [13:33:37] (03CR) 10Ssingh: [C:03+1] Extend Cumin aliases for URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1329521 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:34:23] RESOLVED: [6x] GnmiInterfaceCountersDrop: cloudsw1-b1-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:34:49] (03PS1) 10Blake: mw-videoscaler: bump envoy to 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329565 (https://phabricator.wikimedia.org/T421418) [13:35:56] (03CR) 10Brouberol: [C:03+2] dse-k8s::worker: enable p2p pulling of large images [puppet] - 10https://gerrit.wikimedia.org/r/1329551 (https://phabricator.wikimedia.org/T436089) (owner: 10Brouberol) [13:36:21] FIRING: SLOBudgetBurn: Search update lag is below 95% target in eqiad - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [13:37:51] (03PS3) 10Milazg: Remove mode from RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328215 (https://phabricator.wikimedia.org/T434267) [13:39:43] !log jmm@cumin2003 END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host ganeti3005.esams.wmnet [13:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:41:11] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12256801 (10cmooney) ` Hi, Please note the issue has not resolved for us. The interface has flapped down over 2,000 times since 05:56 UTC this morning and remains very unstable, though... [13:41:21] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [13:42:47] (03CR) 10Muehlenhoff: [C:03+2] Extend Cumin aliases for URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1329521 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:46:43] (03CR) 10Clément Goubert: [C:03+1] mediawiki: enable forward of fatal metrics to statsd exporter [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) (owner: 10Cwhite) [13:46:51] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1002.wikimedia.org with reason: host reimage [13:49:30] PROBLEM - Host wikikube-worker1138 is DOWN: PING CRITICAL - Packet loss = 77%, RTA = 7794.20 ms [13:50:22] RECOVERY - Host wikikube-worker1138 is UP: PING OK - Packet loss = 0%, RTA = 0.24 ms [13:51:10] (03CR) 10Lerickson: [C:03+2] Increase WDQS backend pod storage to 900GiB. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329388 (https://phabricator.wikimedia.org/T436023) (owner: 10Lerickson) [13:53:11] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1002.wikimedia.org with reason: host reimage [13:53:42] (03Merged) 10jenkins-bot: Increase WDQS backend pod storage to 900GiB. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329388 (https://phabricator.wikimedia.org/T436023) (owner: 10Lerickson) [13:54:47] (03PS1) 10Jforrester: wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-21-213729 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329570 (https://phabricator.wikimedia.org/T433631) [13:54:50] (03PS1) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-08-19-122639 to 2026-08-19-165650 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329571 (https://phabricator.wikimedia.org/T435361) [13:56:03] jouncebot: nowandnext [13:56:03] For the next 0 hour(s) and 3 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1300) [13:56:03] In 0 hour(s) and 3 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1400) [13:56:21] RESOLVED: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [13:58:02] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [13:58:52] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T435989#12256911 (10bking) Sounds good, I will deploy as an 8GB vRAM VM shortly. [13:59:47] (03CR) 10Bking: [C:03+2] Kerberos: Prepare a VM for deployment as krb replica [puppet] - 10https://gerrit.wikimedia.org/r/1329337 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [14:00:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1400) [14:00:59] (03PS1) 10Cwhite: cfssl: limit key sizes to values defined in error string [puppet] - 10https://gerrit.wikimedia.org/r/1329572 [14:03:02] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:03:47] !log brouberol@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1010.eqiad.wmnet with OS bookworm [14:04:03] (03CR) 10Ecarg: [C:03+2] wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-21-213729 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329570 (https://phabricator.wikimedia.org/T433631) (owner: 10Jforrester) [14:04:30] Amir1: You can deploy MW-land stuff now if you want; we’re services-only. [14:04:41] ah, thank you <3 [14:04:59] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [14:06:13] (03Merged) 10jenkins-bot: wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-21-213729 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329570 (https://phabricator.wikimedia.org/T433631) (owner: 10Jforrester) [14:06:14] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329564 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [14:06:46] (03CR) 10Federico Ceratto: [C:03+2] site.pp,db-test2002.yaml: Decommission db-test2002 [puppet] - 10https://gerrit.wikimedia.org/r/1329563 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [14:07:20] Amir1: Can you help clear the git status of /srv/deployment-charts ? [14:07:23] (03Merged) 10jenkins-bot: Switch eswiki to thumb.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329564 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [14:07:30] (We lowly non-roots can’t fix that.) [14:07:35] James_F: what's about it? [14:07:44] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T435989#12256949 (10bking) 05Open→03Resolved [14:07:44] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1329564|Switch eswiki to thumb.wikimedia.org (T427465)]] [14:07:49] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [14:07:59] Amir1: It’s been dirtied (missing return?) which means the cron git update can’t do anything. [14:08:03] !log fceratto@cumin1003 Removing db-test2002 from zarcillo T435912 [14:08:08] T435912: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912 [14:08:10] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) [14:08:13] 10ops-codfw, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12256952 (10ops-monitoring-bot) db-test2002 has been deleted from zarcillo [14:08:17] 10ops-codfw, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12256954 (10ops-monitoring-bot) db-test2002 has been decommissioned by Data Persistence [14:08:19] 10ops-codfw, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12256955 (10ops-monitoring-bot) This host is ready for DC-Ops to decommission [14:08:20] let me take a look [14:08:53] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [14:08:57] Huh. Never mind, it’s not blocking it somehow. [14:09:29] !log ecarg@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:09:50] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:10:16] !log ecarg@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:10:18] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1329564|Switch eswiki to thumb.wikimedia.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:10:31] lol [14:10:44] https://www.irccloud.com/pastebin/G7TNY8jU/ [14:10:53] it's not a real change at all [14:11:04] Yeah. [14:11:34] should be fixed now [14:11:37] <3 [14:11:40] !log ecarg@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:11:53] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [14:13:26] !log bking@cumin2003 START - Cookbook sre.ganeti.makevm for new host krb1004.eqiad.wmnet [14:13:28] !log bking@cumin2003 START - Cookbook sre.dns.netbox [14:13:37] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [14:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:15:41] !log installing node-flatted security updates [14:15:44] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:16:58] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12257023 (10MoritzMuehlenhoff) [14:17:55] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329564|Switch eswiki to thumb.wikimedia.org (T427465)]] (duration: 10m 11s) [14:18:02] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [14:18:55] bking@cumin2003 makevm (PID 843572) is awaiting input [14:20:56] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12257056 (10Jhancock.wm) I'm fine with closing this task. also fwiw, there were no crash tickets for servers with the 5217 CPU, just 5317. i only mention that cause i saw that one pop up in andrea's... [14:21:31] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM krb1004.eqiad.wmnet - bking@cumin2003" [14:22:01] !log ecarg@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:22:48] (03CR) 10Ssingh: [C:03+1] "Perfectly fine and thank you for the patch." [puppet] - 10https://gerrit.wikimedia.org/r/1329395 (owner: 10RLazarus) [14:23:04] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM krb1004.eqiad.wmnet - bking@cumin2003" [14:23:04] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:23:04] !log bking@cumin2003 START - Cookbook sre.dns.wipe-cache krb1004.eqiad.wmnet on all recursors [14:23:05] !log ecarg@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:23:08] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) krb1004.eqiad.wmnet on all recursors [14:23:41] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM krb1004.eqiad.wmnet - bking@cumin2003" [14:23:45] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM krb1004.eqiad.wmnet - bking@cumin2003" [14:25:26] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [14:25:31] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db-test2001.codfw.wmnet [14:25:58] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host krb1004.eqiad.wmnet with OS bookworm [14:26:11] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T435989#12257110 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by bking@cumin2003 for host krb1... [14:27:05] !log jmm@cumin2003 START - Cookbook sre.ganeti.addnode for new host ganeti3005.esams.wmnet to cluster esams03 and group B [14:30:02] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti3005.esams.wmnet to cluster esams03 and group B [14:30:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1400) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1430) [14:30:23] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [14:33:02] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:33:25] !log ecarg@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:35:02] (03PS1) 10Ecarg: Revert "wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-21-213729" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329574 [14:35:14] (03CR) 10Ecarg: [C:03+2] Revert "wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-21-213729" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329574 (owner: 10Ecarg) [14:35:50] fceratto@cumin1003 decommission (PID 533465) is awaiting input [14:36:29] !log bking@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host clouddumps1002.wikimedia.org with OS trixie [14:36:39] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12257224 (10RobH) [14:36:50] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host clouddumps1002.wikimedia.org with OS bookworm [14:38:04] (03Merged) 10jenkins-bot: Revert "wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-21-213729" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329574 (owner: 10Ecarg) [14:38:08] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12257275 (10MoritzMuehlenhoff) [14:38:17] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on krb1004.eqiad.wmnet with reason: host reimage [14:38:32] (03PS1) 10Brouberol: kafka-mirrormaker: stop replicating the mediawiki.job topics to the jumbo cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329576 (https://phabricator.wikimedia.org/T434693) [14:39:21] !log ecarg@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:40:12] !log ecarg@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:41:24] (03PS2) 10Brouberol: kafka-mirrormaker: stop replicating the mediawiki.job topics to the jumbo cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329576 (https://phabricator.wikimedia.org/T434693) [14:42:04] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [14:42:27] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db-test2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [14:42:27] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:42:28] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db-test2001.codfw.wmnet [14:42:51] (03PS3) 10Brouberol: kafka-mirrormaker: stop replicating the mediawiki.job topics to the jumbo cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329576 (https://phabricator.wikimedia.org/T434693) [14:43:02] (03CR) 10Mforns: [C:03+1] "LGTM! thanks :-)" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329576 (https://phabricator.wikimedia.org/T434693) (owner: 10Brouberol) [14:43:02] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:45:00] !log bking@cumin2003 END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on krb1004.eqiad.wmnet with reason: host reimage [14:45:28] fceratto@cumin1003 decommission (PID 533465) is awaiting input [14:46:23] !log jelto@cumin1003 START - Cookbook sre.hosts.reboot-multiple Sequential unattended reboot of 1 host(s) [['A:owner-collaboration-services', 'P{etherpad2*}'], os=bookworm] [14:46:59] !log jelto@cumin1003 END (FAIL) - Cookbook sre.hosts.reboot-multiple (exit_code=99) Sequential unattended reboot of 1 host(s) [['A:owner-collaboration-services', 'P{etherpad2*}'], os=bookworm] [14:47:07] (03CR) 10Brouberol: [C:03+2] kafka-mirrormaker: stop replicating the mediawiki.job topics to the jumbo cluster [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329576 (https://phabricator.wikimedia.org/T434693) (owner: 10Brouberol) [14:48:21] !log brouberol@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/kafka-mirrormaker: apply [14:48:30] !log brouberol@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/kafka-mirrormaker: apply [14:51:05] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12257400 (10MoritzMuehlenhoff) 05Open→03Resolved The server is working again and has been reimaged and re-added to the esams Ganeti cluster. [14:51:12] (03CR) 10Kamila Součková: [C:03+1] P:chartmuseum: Update every minute [puppet] - 10https://gerrit.wikimedia.org/r/1329544 (https://phabricator.wikimedia.org/T435987) (owner: 10Clément Goubert) [14:53:04] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1002.wikimedia.org with reason: host reimage [14:53:20] (03PS1) 10Federico Ceratto: site.pp,db-test2001.yaml: Decommission db-test2001 [puppet] - 10https://gerrit.wikimedia.org/r/1329580 (https://phabricator.wikimedia.org/T435912) [14:53:29] (03PS1) 10Kamila Součková: deployment_server: switch mw-debug/next to PHP 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1329581 (https://phabricator.wikimedia.org/T432989) [14:58:10] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1002.wikimedia.org with reason: host reimage [14:58:23] jouncebot: nowandnext [14:58:23] For the next 0 hour(s) and 1 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1400) [14:58:23] For the next 0 hour(s) and 1 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1430) [14:58:23] In 2 hour(s) and 1 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1700) [15:00:31] (03CR) 10Scott French: [C:03+1] deployment_server: switch mw-debug/next to PHP 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1329581 (https://phabricator.wikimedia.org/T432989) (owner: 10Kamila Součková) [15:01:41] (03CR) 10Clément Goubert: [C:03+2] P:chartmuseum: Update every minute [puppet] - 10https://gerrit.wikimedia.org/r/1329544 (https://phabricator.wikimedia.org/T435987) (owner: 10Clément Goubert) [15:02:06] (03CR) 10Gmodena: [C:03+1] Update memory settings for wdqs-next. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329313 (https://phabricator.wikimedia.org/T434854) (owner: 10Lerickson) [15:02:26] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host krb1004.eqiad.wmnet with OS bookworm [15:02:26] !log bking@cumin2003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host krb1004.eqiad.wmnet [15:02:33] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T435989#12257528 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by bking@cumin2003 for host krb1004.... [15:02:41] (03PS2) 10Cwhite: cfssl: limit key sizes to values defined in error string [puppet] - 10https://gerrit.wikimedia.org/r/1329572 [15:02:47] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Migrate to nftables::service [puppet] - 10https://gerrit.wikimedia.org/r/1329549 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [15:03:13] (03PS6) 10Jelto: sre.hosts.reboot-multiple: add new cookbook for unattended reboots [cookbooks] - 10https://gerrit.wikimedia.org/r/1305707 [15:03:13] (03CR) 10Jelto: "As discussed in IRC I changed the cookbook name and the parameters. I added some test-cookbook tests below combining the parameters. What " [cookbooks] - 10https://gerrit.wikimedia.org/r/1305707 (owner: 10Jelto) [15:03:49] (03CR) 10Majavah: [V:03+1 C:03+2] P:wmcs::novaproxy: Migrate to nftables::service [puppet] - 10https://gerrit.wikimedia.org/r/1329549 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [15:05:01] (03CR) 10Marostegui: [C:03+1] site.pp,db-test2001.yaml: Decommission db-test2001 [puppet] - 10https://gerrit.wikimedia.org/r/1329580 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [15:08:52] (03CR) 10Federico Ceratto: [C:03+2] site.pp,db-test2001.yaml: Decommission db-test2001 [puppet] - 10https://gerrit.wikimedia.org/r/1329580 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [15:09:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [15:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:10:45] !log fceratto@cumin1003 Removing db-test2001 from zarcillo T435912 [15:10:50] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) [15:10:51] T435912: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912 [15:10:55] 10ops-codfw, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12257571 (10ops-monitoring-bot) db-test2001 has been deleted from zarcillo [15:10:58] 10ops-codfw, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12257573 (10ops-monitoring-bot) db-test2001 has been decommissioned by Data Persistence [15:11:00] 10ops-codfw, 06DBA, 06DC-Ops, 10decommission-hardware: Decommission db-test* hosts - https://phabricator.wikimedia.org/T435912#12257574 (10ops-monitoring-bot) This host is ready for DC-Ops to decommission [15:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:17:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:18:16] T1 [15:18:44] (03CR) 10Ssingh: [C:03+1] hieradata: use cfg for pdns v5 on dns1004 [puppet] - 10https://gerrit.wikimedia.org/r/1327151 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [15:18:58] (03CR) 10CDobbins: [V:03+1 C:03+2] hieradata: use cfg for pdns v5 on dns1004 [puppet] - 10https://gerrit.wikimedia.org/r/1327151 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [15:20:53] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=dns1004.* [reason: trixie upgrade] [15:21:14] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host dns1004.wikimedia.org with OS trixie [15:22:10] RESOLVED: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:24:48] !log Deployed patch for T432848 [15:24:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:24:54] T432848: EventLogging: Record confidence score results - https://phabricator.wikimedia.org/T432848 [15:24:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:25:08] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [15:27:25] FIRING: [5x] BFDdown: BFD session down between cr1-eqiad and 208.80.154.6 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:27:34] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:28:10] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [15:28:16] PROBLEM - Host 2620:0:861:1:208:80:154:6 is DOWN: CRITICAL - Host Unreachable (2620:0:861:1:208:80:154:6) [15:28:42] Deployed the private-side change to production. Inert until ModelToRun gains the fourth constructor parameter. SAL: https://sal.toolforge.org/log/Kb2sPqABDAZZyZXnhPJB [15:28:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:29:46] PROBLEM - Recursive DNS on 208.80.154.6 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [15:31:13] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [15:33:05] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: fix dns entries for cr2-eqiad transport interface IPs - cmooney@cumin1003" [15:33:09] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: fix dns entries for cr2-eqiad transport interface IPs - cmooney@cumin1003" [15:33:10] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:33:37] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [15:34:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [15:35:14] RECOVERY - Host 2620:0:861:1:208:80:154:6 is UP: PING OK - Packet loss = 0%, RTA = 0.32 ms [15:35:18] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [15:36:46] PROBLEM - Recursive DNS on 2620:0:861:1:208:80:154:6 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [15:37:59] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on dns1004.wikimedia.org with reason: host reimage [15:38:52] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12257699 (10MoritzMuehlenhoff) [15:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:41:22] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: fix dns entries for cr2-eqiad transport interface IPs - cmooney@cumin1003" [15:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:44:03] (03PS1) 10Federico Ceratto: preseed.yaml: Remove db-test entry [puppet] - 10https://gerrit.wikimedia.org/r/1329594 (https://phabricator.wikimedia.org/T435912) [15:44:16] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1004.wikimedia.org with reason: host reimage [15:44:25] cmooney@cumin1003 netbox (PID 583121) is awaiting input [15:44:43] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: fix dns entries for cr2-eqiad transport interface IPs - cmooney@cumin1003" [15:44:43] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:54:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:57:44] RECOVERY - Recursive DNS on 208.80.154.6 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [15:57:44] RECOVERY - Recursive DNS on 2620:0:861:1:208:80:154:6 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [15:58:31] (03CR) 10RLazarus: [C:03+2] "Thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1329395 (owner: 10RLazarus) [15:58:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:59:07] (03CR) 10RLazarus: [V:03+1 C:03+2] "Thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1329402 (owner: 10RLazarus) [15:59:55] (03CR) 10Ssingh: "Looks good but I have a question. It tries urldownloader1003 as part of this, which is expected. test-cookbook fails with" [cookbooks] - 10https://gerrit.wikimedia.org/r/1329523 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [16:02:25] FIRING: [6x] BFDdown: BFD session down between cr1-codfw and fe80::669:8f07:ece:f4e7 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:02:40] FIRING: [6x] BFDdown: BFD session down between cr1-codfw and fe80::669:8f07:ece:f4e7 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:02:50] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12257848 (10Cyberpower678) I'm told everything is working fine on the archive side of things as well. [16:03:47] PROBLEM - Host 2001:df2:e500:1:103:102:166:10 is DOWN: PING CRITICAL - Packet loss = 100% [16:03:47] PROBLEM - Host 2001:df2:e500:2:103:102:166:36 is DOWN: PING CRITICAL - Packet loss = 100% [16:04:01] RECOVERY - Host 2001:df2:e500:2:103:102:166:36 is UP: PING WARNING - Packet loss = 77%, RTA = 210.23 ms [16:04:01] RECOVERY - Host 2001:df2:e500:1:103:102:166:10 is UP: PING OK - Packet loss = 0%, RTA = 209.91 ms [16:06:41] PROBLEM - Host cp5022 is DOWN: PING CRITICAL - Packet loss = 100% [16:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:11:05] (03CR) 10Elukey: "To keep archives happy - I synced with Cole and it seems a very spot-on change, but since it is a delicate part of our codebase I'll have " [puppet] - 10https://gerrit.wikimedia.org/r/1329572 (owner: 10Cwhite) [16:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:14:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [16:17:40] FIRING: [6x] BFDdown: BFD session down between cr1-codfw and fe80::669:8f07:ece:f4e7 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:18:02] RESOLVED: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [16:18:44] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1004.wikimedia.org with OS trixie [16:20:24] !log cdobbins@cumin1003 conftool action : set/pooled=yes; selector: name=dns1004.* [reason: trixie upgrade] [16:22:25] RESOLVED: [4x] BFDdown: BFD session down between cr1-eqiad and 208.80.154.6 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:23:48] !log cmooney@cumin1003 START - Cookbook sre.metamonitoring.downtime Downtime for 4:00:00 of prometheus/deadmanswitchnotified, prometheus/deadmanswitchonamdb, prometheus/extmon on 2 host(s) with reason: eqsin site rebuild [16:23:54] !log cmooney@cumin1003 END (PASS) - Cookbook sre.metamonitoring.downtime (exit_code=0) Downtime for 4:00:00 of prometheus/deadmanswitchnotified, prometheus/deadmanswitchonamdb, prometheus/extmon on 2 host(s) with reason: eqsin site rebuild [16:24:11] (03Restored) 10CDobbins: hieradata: add cfg flag for pdns v5 for dns1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327152 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [16:24:39] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 12 hosts with reason: site migration to new Nokia switches [16:24:39] (03CR) 10JHathaway: [C:03+1] cfssl: limit key sizes to values defined in error string [puppet] - 10https://gerrit.wikimedia.org/r/1329572 (owner: 10Cwhite) [16:24:54] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12257936 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b58886a4-bace-4b3a-a42f-d960bc8c4bc4) set by cmooney@cumin1003 fo... [16:25:11] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 13 hosts with reason: site migration to new Nokia switches [16:25:16] (03CR) 10Jasmine: [C:03+2] site.pp, preseed.yaml: add conf101[0-2] [puppet] - 10https://gerrit.wikimedia.org/r/1329342 (https://phabricator.wikimedia.org/T435426) (owner: 10Jasmine) [16:25:18] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12257938 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b0850968-674c-42f7-b6a0-d37d70f8c521) set by cmooney@cumin1003 fo... [16:25:45] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 12 hosts with reason: site migration to new Nokia switches [16:25:54] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12257941 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=1866892e-3c8b-4e8a-b581-189ea17a8c5b) set by cmooney@cumin1003 fo... [16:29:24] 10ops-eqiad, 06SRE, 06DC-Ops, 06ServiceOps, and 2 others: Q1:rack/setup/install conf101[0-2] - https://phabricator.wikimedia.org/T435426#12257988 (10jasmine_) >>! In T435426#12236303, @Clement_Goubert wrote: > @jasmine_ Can you take care of adding the servers to `site.pp` and `preseed.yml` please? Sure, d... [16:29:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [16:38:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:42:25] RESOLVED: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:47:23] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [16:48:22] (03CR) 10Ssingh: [C:03+1] hieradata: add cfg flag for pdns v5 for dns1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327152 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [16:48:50] RECOVERY - Host mr1-eqsin IPv6 is UP: PING WARNING - Packet loss = 33%, RTA = 230.84 ms [16:51:14] (03CR) 10CDobbins: [C:03+2] hieradata: add cfg flag for pdns v5 for dns1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327152 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [16:54:23] cmooney@cumin1003 netbox (PID 635927) is awaiting input [16:54:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:55:19] (03PS1) 10Cathal Mooney: Add missing include for eqsin private loopback range [dns] - 10https://gerrit.wikimedia.org/r/1329608 (https://phabricator.wikimedia.org/T418439) [16:55:40] (03CR) 10Aaron Schulz: Add configurable RestModuleOverrides (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [16:57:26] (03CR) 10Cathal Mooney: [C:03+2] Add missing include for eqsin private loopback range [dns] - 10https://gerrit.wikimedia.org/r/1329608 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [16:57:53] !log cmooney@dns3003 START - running authdns-update [16:58:08] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=dns1005.* [reason: trixie upgrade] [16:58:25] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host dns1005.wikimedia.org with OS trixie [16:58:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T1700) [17:00:20] !log cmooney@dns3003 FAIL - running authdns-update [17:02:06] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: fix dns entries for cr2-eqiad transport interface IPs - cmooney@cumin1003" [17:02:28] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: fix dns entries for cr2-eqiad transport interface IPs - cmooney@cumin1003" [17:02:28] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:03:10] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 208.80.154.153 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:04:42] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12258203 (10ssingh) I have tried to reproduce now with some obscure paths as well but "Save Page" is working fine. @AlexisJazz: can you confirm your workflow here please? Just... [17:05:16] PROBLEM - Host 2620:0:861:2:208:80:154:153 is DOWN: CRITICAL - Host Unreachable (2620:0:861:2:208:80:154:153) [17:05:16] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kamila@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [17:06:46] PROBLEM - Recursive DNS on 208.80.154.153 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [17:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [17:08:10] FIRING: [4x] BFDdown: BFD session down between cr1-eqiad and 208.80.154.153 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:08:17] (03Merged) 10jenkins-bot: Add title-case mapping to support migration to PHP 8.5 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [17:08:42] !log kamila@deploy1003 Started scap sync-world: Backport for [[gerrit:1328667|Add title-case mapping to support migration to PHP 8.5 (T432985)]] [17:09:46] (03PS1) 10Aaron Schulz: Rest: Rename "mode" key to "availability" in RestModuleOverrides [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329616 (https://phabricator.wikimedia.org/T433522) [17:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:11:16] !log kamila@deploy1003 kamila: Backport for [[gerrit:1328667|Add title-case mapping to support migration to PHP 8.5 (T432985)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:11:17] (03CR) 10Aaron Schulz: [C:03+1] "This can go out once https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1327198 lands in the next wmf branch cut and that branch is out on " [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328215 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [17:11:23] T432985: [PHP 8.5] Generate and deploy title-case mapping in MediaWiki configuration (wmf-config/UcfirstOverrides.php) - https://phabricator.wikimedia.org/T432985 [17:11:48] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 26 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329616 (https://phabricator.wikimedia.org/T433522) (owner: 10Aaron Schulz) [17:12:16] RECOVERY - Host 2620:0:861:2:208:80:154:153 is UP: PING OK - Packet loss = 0%, RTA = 0.46 ms [17:13:46] PROBLEM - Recursive DNS on 2620:0:861:2:208:80:154:153 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [17:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:14:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [17:14:53] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on dns1005.wikimedia.org with reason: host reimage [17:15:42] PROBLEM - SSH on an-druid1007 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:15:59] !log kamila@deploy1003 kamila: Continuing with deployment [17:18:11] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns1005.wikimedia.org with reason: host reimage [17:20:20] !log kamila@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328667|Add title-case mapping to support migration to PHP 8.5 (T432985)]] (duration: 11m 38s) [17:20:26] T432985: [PHP 8.5] Generate and deploy title-case mapping in MediaWiki configuration (wmf-config/UcfirstOverrides.php) - https://phabricator.wikimedia.org/T432985 [17:20:49] (03PS1) 10CDobbins: hieradata: apply pdns v5 flag to all dns hosts [puppet] - 10https://gerrit.wikimedia.org/r/1329618 (https://phabricator.wikimedia.org/T401832) [17:22:11] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1002.wikimedia.org with OS bookworm [17:22:12] !log taavi@cumin1003 conftool action : set/pooled=yes; selector: name=clouddumps1002.wikimedia.org,service=dumps-(rsync|https) [17:22:40] (03PS2) 10CDobbins: hieradata: apply pdns v5 flag to all dns hosts [puppet] - 10https://gerrit.wikimedia.org/r/1329618 (https://phabricator.wikimedia.org/T401832) [17:24:42] RECOVERY - SSH on an-druid1007 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u7 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:24:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:27:42] PROBLEM - SSH on an-druid1007 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:28:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:29:44] FIRING: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [17:30:01] (03PS1) 10Cathal Mooney: Rename old site-wide vlans in eqsin to rack 603 specific [puppet] - 10https://gerrit.wikimedia.org/r/1329621 (https://phabricator.wikimedia.org/T418439) [17:32:05] (03CR) 10Kamila Součková: [C:03+2] deployment_server: switch mw-debug/next to PHP 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1329581 (https://phabricator.wikimedia.org/T432989) (owner: 10Kamila Součková) [17:32:40] RECOVERY - SSH on an-druid1007 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u7 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:33:42] PROBLEM - SSH on an-druid1006 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:36:30] (03CR) 10Krinkle: [C:03+1] Profiler: Fix excimer component regex skipping the last frame in each stack [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328191 (owner: 10SomeRandomDeveloper) [17:37:28] (03CR) 10Jasmine: [C:03+2] Add Kubernetes POD IP reverse range delegations for wikikube-ctrl2006 [dns] - 10https://gerrit.wikimedia.org/r/1285465 (https://phabricator.wikimedia.org/T406596) (owner: 10Jasmine) [17:37:44] RECOVERY - Recursive DNS on 2620:0:861:2:208:80:154:153 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [17:37:44] RECOVERY - Recursive DNS on 208.80.154.153 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [17:39:20] !log jasmine@dns2004 START - running authdns-update [17:39:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:41:43] Is anyone doing deployments right now, or is it safe for me to deploy a small MW patch (for debugging an urgent bug)? [17:41:48] (03CR) 10Kamila Součková: [C:04-1] "DNM until https://phabricator.wikimedia.org/T432985#12258393 is understood." [puppet] - 10https://gerrit.wikimedia.org/r/1329581 (https://phabricator.wikimedia.org/T432989) (owner: 10Kamila Součková) [17:41:54] !log jasmine@dns2004 END - running authdns-update [17:42:07] RoanKattouw: Raine is scapping I believe [17:42:12] (03PS1) 10Catrope: Log failures to generate recovery codes in the account creation email [extensions/OATHAuth] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329624 (https://phabricator.wikimedia.org/T436066) [17:42:34] RoanKattouw: Raine really wanted to be scapping, but php has decided otherwise [17:42:36] (03CR) 10Jasmine: [C:03+2] wmnet: add wikikube-ctrl2006 to etcd-server SRV record [dns] - 10https://gerrit.wikimedia.org/r/1249423 (https://phabricator.wikimedia.org/T406596) (owner: 10Jasmine) [17:42:49] !log jasmine@dns2004 START - running authdns-update [17:42:52] the floor is yours, unfortunately [17:43:01] womp womp :( [17:43:17] My condolences and thank you, this should be quick (<10 mins) [17:43:29] rzl: oh hi you're here I may have something for you :D [17:43:41] (03CR) 10TrainBranchBot: [C:03+2] "Approved by catrope@deploy1003 using scap backport" [extensions/OATHAuth] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329624 (https://phabricator.wikimedia.org/T436066) (owner: 10Catrope) [17:43:45] rzl? who is rzl? I am Guy Incognito [17:43:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:44:00] (03PS1) 10Cwhite: opensearch: change pki algo to rsa [puppet] - 10https://gerrit.wikimedia.org/r/1329625 (https://phabricator.wikimedia.org/T350516) [17:44:20] (03PS1) 10Cathal Mooney: wmf-netbox.py: remove exception to vlan naming convention for eqsin [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1329626 (https://phabricator.wikimedia.org/T418439) [17:45:14] !log jasmine@dns2004 END - running authdns-update [17:46:11] (03Merged) 10jenkins-bot: Log failures to generate recovery codes in the account creation email [extensions/OATHAuth] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329624 (https://phabricator.wikimedia.org/T436066) (owner: 10Catrope) [17:46:35] !log catrope@deploy1003 Started scap sync-world: Backport for [[gerrit:1329624|Log failures to generate recovery codes in the account creation email (T436066)]] [17:46:40] T436066: Initial 2FA not created on otrs_wikiwiki - https://phabricator.wikimedia.org/T436066 [17:47:59] (03CR) 10Ssingh: [C:03+1] Rename old site-wide vlans in eqsin to rack 603 specific [puppet] - 10https://gerrit.wikimedia.org/r/1329621 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [17:49:10] !log catrope@deploy1003 catrope: Backport for [[gerrit:1329624|Log failures to generate recovery codes in the account creation email (T436066)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:49:30] (03CR) 10Jasmine: [C:03+2] wikikube: add wikikube-ctrl2006 [puppet] - 10https://gerrit.wikimedia.org/r/1249321 (https://phabricator.wikimedia.org/T406596) (owner: 10Jasmine) [17:49:52] !log catrope@deploy1003 catrope: Continuing with deployment [17:50:18] (03CR) 10Cathal Mooney: [C:03+2] Rename old site-wide vlans in eqsin to rack 603 specific [puppet] - 10https://gerrit.wikimedia.org/r/1329621 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [17:51:34] RECOVERY - SSH on an-druid1006 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u7 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [17:51:36] PROBLEM - Druid historical on an-druid1006 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [17:53:36] RECOVERY - Druid historical on an-druid1006 is OK: PROCS OK: 1 process with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [17:54:06] (03PS1) 10Cathal Mooney: BIRD profile eqsin: delete file as no longer needed [puppet] - 10https://gerrit.wikimedia.org/r/1329628 (https://phabricator.wikimedia.org/T418439) [17:54:10] !log catrope@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329624|Log failures to generate recovery codes in the account creation email (T436066)]] (duration: 07m 35s) [17:54:16] T436066: Initial 2FA not created on otrs_wikiwiki - https://phabricator.wikimedia.org/T436066 [17:56:22] (03CR) 10Kamila Součková: [C:03+2] "I goofed, everything looks good (famous last words)." [puppet] - 10https://gerrit.wikimedia.org/r/1329581 (https://phabricator.wikimedia.org/T432989) (owner: 10Kamila Součková) [17:57:05] Raine: I'm done, the floor is yours again (if you're ready for it) [17:57:13] RoanKattouw: perfect timing :D thanks! [17:57:31] (I was ready for about 2 seconds when you wrote :D) [17:58:54] (03CR) 10Ssingh: [C:03+1] BIRD profile eqsin: delete file as no longer needed [puppet] - 10https://gerrit.wikimedia.org/r/1329628 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [18:00:45] jouncebot: nowandnext [18:00:45] No deployments scheduled for the next 1 hour(s) and 59 minute(s) [18:00:45] In 1 hour(s) and 59 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T2000) [18:01:08] (03CR) 10Cathal Mooney: [C:03+2] BIRD profile eqsin: delete file as no longer needed [puppet] - 10https://gerrit.wikimedia.org/r/1329628 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [18:01:51] (03PS2) 10Cwhite: opensearch: change pki algo to rsa [puppet] - 10https://gerrit.wikimedia.org/r/1329625 (https://phabricator.wikimedia.org/T350516) [18:03:10] RESOLVED: [4x] BFDdown: BFD session down between cr1-eqiad and 208.80.154.153 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:04:44] FIRING: [10x] RipeAtlasAnchorUnreachable: ipv4 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133212 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [18:05:12] !log kamila@deploy1003 Started scap sync-world: Switch mw-debug/next to 8.5 - T432989 [18:05:17] T432989: [PHP 8.5] Switch mw-debug/next and Scap to PHP 8.5 images - https://phabricator.wikimedia.org/T432989 [18:07:59] (03PS3) 10Cwhite: cfssl: limit key sizes to values defined in error string [puppet] - 10https://gerrit.wikimedia.org/r/1329572 [18:09:57] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:10:05] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [18:10:55] !log kamila@deploy1003 Finished scap sync-world: Switch mw-debug/next to 8.5 - T432989 (duration: 06m 00s) [18:11:00] T432989: [PHP 8.5] Switch mw-debug/next and Scap to PHP 8.5 images - https://phabricator.wikimedia.org/T432989 [18:12:13] (03CR) 10Ssingh: "Looks good. Let's run PCC on a few hosts to confirm NOOP? 2-3 random DNS hosts is fine." [puppet] - 10https://gerrit.wikimedia.org/r/1329618 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [18:12:54] (03PS1) 10Cathal Mooney: Eqsin routed ganeti: remove defs for BGP peering to CRs [puppet] - 10https://gerrit.wikimedia.org/r/1329629 (https://phabricator.wikimedia.org/T418439) [18:13:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:14:29] (03CR) 10Ssingh: [C:03+1] Eqsin routed ganeti: remove defs for BGP peering to CRs [puppet] - 10https://gerrit.wikimedia.org/r/1329629 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [18:14:44] FIRING: [10x] RipeAtlasAnchorUnreachable: ipv4 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133212 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [18:15:04] (03CR) 10Cathal Mooney: [C:03+2] Eqsin routed ganeti: remove defs for BGP peering to CRs [puppet] - 10https://gerrit.wikimedia.org/r/1329629 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [18:15:30] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (NOOP 3): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9328/console" [puppet] - 10https://gerrit.wikimedia.org/r/1329618 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [18:16:36] (03CR) 10Ssingh: [C:03+1] "OK to merge tomorrow, just in case." [puppet] - 10https://gerrit.wikimedia.org/r/1329618 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [18:17:01] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns1005.wikimedia.org with OS trixie [18:18:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:19:10] (03CR) 10CDobbins: [V:03+1] "I'll merge it tomorrow morning. Thanks for the review!" [puppet] - 10https://gerrit.wikimedia.org/r/1329618 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [18:20:02] (03CR) 10Ssingh: [C:03+1] "Makes sense, in my limited understanding of this, but as it pertains to the work today and the comment!" [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1329626 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [18:20:24] (03CR) 10Cathal Mooney: [C:03+2] wmf-netbox.py: remove exception to vlan naming convention for eqsin [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1329626 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [18:25:46] (03PS7) 10Andrea Denisse: grafana: Disable loading of not installed core plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) [18:25:46] (03CR) 10Andrea Denisse: [V:03+1] "PCC results: https://puppet-compiler.wmflabs.org/output/1329517/9331/" [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [18:26:09] !log sukhe@cumin1003 START - Cookbook sre.hosts.remove-downtime for dns1005.wikimedia.org [18:26:10] !log sukhe@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns1005.wikimedia.org [18:26:37] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=dns1005.wikimedia.org [reason: trixie upgrade completed] [18:27:02] !log sukhe@dns1004 START - running authdns-update [18:28:54] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:29:19] !log sukhe@dns1004 END - running authdns-update [18:29:44] FIRING: [10x] RipeAtlasAnchorUnreachable: ipv4 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133212 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [18:30:30] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12258634 (10cmooney) Lumen have come back twice to say everything is ok still. I have in both cases responded to them showing logs and letting them know we need the service to work and... [18:34:04] elukey: My IRC client is being weird. Please ask jnuche about the repos/test-platform/catalyst/patchdemo images. [18:36:54] (03PS1) 10Cathal Mooney: Eqsin routed ganeti: one more time with feeling [puppet] - 10https://gerrit.wikimedia.org/r/1329632 (https://phabricator.wikimedia.org/T418439) [18:37:22] (03CR) 10Ssingh: [C:03+1] "<3" [puppet] - 10https://gerrit.wikimedia.org/r/1329632 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [18:37:35] (03CR) 10Cathal Mooney: [C:03+2] Eqsin routed ganeti: one more time with feeling [puppet] - 10https://gerrit.wikimedia.org/r/1329632 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [18:39:16] (03CR) 10RLazarus: [C:03+1] mw-videoscaler: bump envoy to 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329565 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [18:39:57] FIRING: [2x] SystemdUnitFailed: rsync-srv_firmwares.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:43:54] FIRING: [2x] SystemdUnitFailed: rsync-srv_firmwares.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:46:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:51:44] (03CR) 10Eevans: [C:03+2] admin: replace ssh key for user bgwiki [puppet] - 10https://gerrit.wikimedia.org/r/1329374 (https://phabricator.wikimedia.org/T433313) (owner: 10Eevans) [18:55:11] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12258737 (10Eevans) 05Open→03Resolved It may take as much as an hour from now to take effect, but you should be good to go @Bethany; Let... [18:57:41] (03PS1) 10Bking: Kerberos: Apply kerberos role to newly-provisioned host [puppet] - 10https://gerrit.wikimedia.org/r/1329636 (https://phabricator.wikimedia.org/T435873) [18:58:36] !log "homer 'lsw1-d8-codfw*' commit "enabled BGP on new wikikube-ctrl2006 host - T406596" [18:58:41] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:58:42] T406596: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596 [18:59:08] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329636 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [18:59:10] !log jasmine@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl2006.codfw.wmnet [18:59:13] !log jasmine@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl2006.codfw.wmnet [18:59:25] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12258759 (10ops-monitoring-bot) Cookbook cookbooks.sre.k8s.pool-depool-node started by jasmine@cumin1003 pool for host wikikube-ctrl2006.codfw.wmnet compl... [19:04:57] FIRING: [2x] SystemdUnitFailed: rsync-srv_firmwares.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:06:39] (03CR) 10Bking: "As in previous CRs, the PCC failure is due to `krb1004` being too new, and can be safely ignored." [puppet] - 10https://gerrit.wikimedia.org/r/1329636 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [19:07:03] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [19:07:06] RECOVERY - Host doh5003 is UP: PING OK - Packet loss = 0%, RTA = 211.88 ms [19:07:08] RECOVERY - Host durum5003 is UP: PING OK - Packet loss = 0%, RTA = 210.45 ms [19:07:08] RECOVERY - Host durum5004 is UP: PING OK - Packet loss = 0%, RTA = 210.23 ms [19:07:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:07:12] FIRING: [13x] JobUnavailable: Reduced availability for job cache_haproxy_tls in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [19:07:15] FIRING: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [19:07:28] RECOVERY - Host tcp-proxy5003 is UP: PING OK - Packet loss = 0%, RTA = 210.26 ms [19:07:30] RECOVERY - Host ncredir5003 is UP: PING OK - Packet loss = 0%, RTA = 210.23 ms [19:07:30] RECOVERY - Host ncredir5004 is UP: PING OK - Packet loss = 0%, RTA = 210.37 ms [19:07:32] FIRING: [2x] SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [19:07:35] FIRING: FNMNotReported: FastNetMon metrics not reported - https://wikitech.wikimedia.org/wiki/Fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DFNMNotReported [19:07:38] RECOVERY - Host tcp-proxy5004 is UP: PING OK - Packet loss = 0%, RTA = 210.51 ms [19:07:38] RECOVERY - Host hcaptcha-proxy5003 is UP: PING OK - Packet loss = 0%, RTA = 230.56 ms [19:07:38] RECOVERY - Host hcaptcha-proxy5004 is UP: PING OK - Packet loss = 0%, RTA = 230.57 ms [19:07:38] RECOVERY - Host install5004 is UP: PING OK - Packet loss = 0%, RTA = 230.70 ms [19:07:38] RECOVERY - Host prometheus5003 is UP: PING OK - Packet loss = 0%, RTA = 210.46 ms [19:07:39] RECOVERY - Host doh5004 is UP: PING OK - Packet loss = 0%, RTA = 210.20 ms [19:07:39] RECOVERY - Host netflow5003 is UP: PING OK - Packet loss = 0%, RTA = 230.68 ms [19:07:40] FIRING: [3x] ProbeDown: Service text:80 has failed probes (http_text_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:07:46] o/ [19:07:48] !ack [19:07:48] 8309 (ACKED) [3x] ProbeDown sre (probes/service eqsin) [19:07:59] eqsin fiber flap? [19:08:00] RECOVERY - Host bast5005 is UP: PING OK - Packet loss = 0%, RTA = 230.61 ms [19:08:16] is the recabling work done thre? [19:08:18] yeah looks that way -- no point depooling at this point probably [19:08:43] (03PS4) 10Krinkle: Profiler: Fix excimer component regex skipping the last frame in each stack [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328191 (owner: 10SomeRandomDeveloper) [19:08:43] (03PS1) 10Krinkle: Profiler: Fold excimer line parsing into one method [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329640 [19:09:09] hmm, would be a great time for the eqsin data on https://grafana.wikimedia.org/d/000000093/cdn-frontend-network?from=now-3h&to=now&timezone=utc&var-site=$__all to be working [19:09:26] rzl: let's --> #-sre [19:09:28] oh no, it's just zeroed out since 07:00, already depooled I guess [19:09:29] 👍 [19:09:42] sorry, yeah that's what I mean - is that still ongoing :) [19:09:57] RESOLVED: [2x] SystemdUnitFailed: rsync-srv_firmwares.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:11:20] (03PS2) 10Krinkle: Profiler: Fold excimer line parsing into one method [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329640 [19:12:03] RESOLVED: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [19:12:10] RESOLVED: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:12:12] RESOLVED: [31x] JobUnavailable: Reduced availability for job benthos in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [19:12:21] RESOLVED: [2x] SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [19:12:23] RESOLVED: FNMNotReported: FastNetMon metrics not reported - https://wikitech.wikimedia.org/wiki/Fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DFNMNotReported [19:12:28] FIRING: [8x] ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:14:33] (03PS3) 10Krinkle: Profiler: Fold excimer line parsing into one method [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329640 [19:14:44] RESOLVED: [8x] RipeAtlasAnchorUnreachable: ipv4 ping to eqsin RIPE Atlas anchor: failures over threshold for measurement 95145503 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [19:15:05] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12258835 (10jasmine_) 05Open→03Resolved [19:17:02] (03PS4) 10Krinkle: Profiler: Fold excimer line parsing into one method [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329640 [19:29:08] (03CR) 10Ryan Kemper: [C:03+1] Kerberos: Apply kerberos role to newly-provisioned host [puppet] - 10https://gerrit.wikimedia.org/r/1329636 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [19:32:15] RESOLVED: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [19:35:05] (03PS1) 10CDanis: geodns: Prometheus metric for per-site pooled state [puppet] - 10https://gerrit.wikimedia.org/r/1329644 [19:35:22] (03PS1) 10Ahmon Dancy: README: say when to borrow a plugin path and when not to [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1329645 [19:35:58] (03PS2) 10Ahmon Dancy: README: say when to borrow a plugin path and when not to [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1329645 [19:36:19] (03CR) 10Ahmon Dancy: [C:03+2] README: say when to borrow a plugin path and when not to [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1329645 (owner: 10Ahmon Dancy) [19:36:58] (03Merged) 10jenkins-bot: README: say when to borrow a plugin path and when not to [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1329645 (owner: 10Ahmon Dancy) [19:39:06] (03PS2) 10CDanis: geodns: Prometheus metric for per-site pooled state [puppet] - 10https://gerrit.wikimedia.org/r/1329644 [19:39:07] (03CR) 10CDanis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329644 (owner: 10CDanis) [19:39:07] (03CR) 10CDanis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329644 (owner: 10CDanis) [19:40:07] (03PS1) 10Cathal Mooney: Remove overrides for LVS servers in eqsin to let them peer with TOR [puppet] - 10https://gerrit.wikimedia.org/r/1329647 (https://phabricator.wikimedia.org/T418439) [19:40:42] (03CR) 10Ssingh: [C:03+1] Remove overrides for LVS servers in eqsin to let them peer with TOR [puppet] - 10https://gerrit.wikimedia.org/r/1329647 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [19:41:47] (03PS1) 10Ahmon Dancy: wm-wheres-my-code-running: Show branch when hovering over a place chip [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1329648 (https://phabricator.wikimedia.org/T434726) [19:42:46] (03CR) 10Ahmon Dancy: [C:03+2] wm-wheres-my-code-running: Show branch when hovering over a place chip [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1329648 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [19:43:35] (03Merged) 10jenkins-bot: wm-wheres-my-code-running: Show branch when hovering over a place chip [software/gerrit] (deploy/wmf/stable-3.10) - 10https://gerrit.wikimedia.org/r/1329648 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [19:44:54] !log dancy@deploy1003 Started deploy [gerrit/gerrit@637dbd8]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1329648 [19:45:08] !log dancy@deploy1003 Finished deploy [gerrit/gerrit@637dbd8]: Deploying https://gerrit.wikimedia.org/r/c/operations/software/gerrit/+/1329648 (duration: 00m 14s) [19:45:23] !log brett@cumin2003 START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, no task ID specified] [19:45:47] !log brett@cumin2003 END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: pool eqsin [reason: no reason specified, no task ID specified] [19:46:17] (03CR) 10Ssingh: [C:03+2] Remove overrides for LVS servers in eqsin to let them peer with TOR [puppet] - 10https://gerrit.wikimedia.org/r/1329647 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [19:46:40] !log brett@puppetserver1001 conftool action : set/pooled=yes; selector: name=dns5.* [19:47:23] (03PS3) 10CDanis: geodns: Prometheus metric for per-site pooled state [puppet] - 10https://gerrit.wikimedia.org/r/1329644 [19:49:54] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin and A:liberica [19:49:58] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.admin depooling P{lvs5004.eqsin.wmnet} and A:liberica [19:50:13] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs5004.eqsin.wmnet} and A:liberica [19:50:36] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.admin pooling P{lvs5004.eqsin.wmnet} and A:liberica [19:51:02] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs5004.eqsin.wmnet} and A:liberica [19:51:08] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.admin depooling P{lvs5005.eqsin.wmnet} and A:liberica [19:52:20] FIRING: [4x] ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:52:53] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs5005.eqsin.wmnet} and A:liberica [19:53:16] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.admin pooling P{lvs5005.eqsin.wmnet} and A:liberica [19:53:41] !log sukhe@dns1004 START - running authdns-update [19:53:42] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs5005.eqsin.wmnet} and A:liberica [19:53:49] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.admin depooling P{lvs5006.eqsin.wmnet} and A:liberica [19:54:04] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs5006.eqsin.wmnet} and A:liberica [19:54:27] !log sukhe@cumin1003 START - Cookbook sre.loadbalancer.admin pooling P{lvs5006.eqsin.wmnet} and A:liberica [19:54:37] (03CR) 10RLazarus: [C:03+1] "Looks good, and thanks for the admin_state.tpl.erb pointer, which makes verifying the key structure a lot easier." [puppet] - 10https://gerrit.wikimedia.org/r/1329644 (owner: 10CDanis) [19:54:56] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs5006.eqsin.wmnet} and A:liberica [19:54:58] (03CR) 10CDanis: [C:03+2] geodns: Prometheus metric for per-site pooled state [puppet] - 10https://gerrit.wikimedia.org/r/1329644 (owner: 10CDanis) [19:54:58] !log sukhe@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin and A:liberica [19:55:57] !log sukhe@dns1004 END - running authdns-update [19:57:20] RESOLVED: [8x] ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:58:37] !log sukhe@puppetserver1001 conftool action : set/pooled=no; selector: name=cp5022.eqsin.wmnet [19:58:56] (03PS1) 10C. Scott Ananian: Simply ParserCache configuration [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329649 [19:59:00] !incidents [19:59:00] 8309 (RESOLVED) [3x] ProbeDown sre (probes/service eqsin) [19:59:00] 8308 (RESOLVED) Alertmanager: Prometheus Watchdog Hook Processing check is still DOWN [19:59:00] 8307 (RESOLVED) Alertmanager: Prometheus Watchdog Hook Processing check is still DOWN [19:59:01] 8305 (RESOLVED) Alertmanager-active: Prometheus Watchdog Alerts Presence check is still DOWN [19:59:01] 8304 (RESOLVED) Alertmanager-passive: Prometheus Watchdog Alerts Presence check is still DOWN [19:59:01] 8306 (RESOLVED) Alertmanager: Prometheus Watchdog Hook Processing check is still DOWN [19:59:01] 8303 (RESOLVED) ATSBackendErrorsHigh cache_text sre (lists.discovery.wmnet eqiad) [20:00:04] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor I � Unicode. All rise for UTC late backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T2000). [20:00:04] AaronSchulz: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:01:27] RESOLVED: [4x] ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) #page - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:02:00] (from SRE's side, nothing ongoing that would preclude deployments in the backport window) [20:06:23] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough-eqsin [20:07:56] I'd be happy to provide deployment service for AaronSchulz but he should also feel free to deploy himself if he wants [20:08:18] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough-eqsin [20:09:50] (03PS5) 10CDanis: New swift_per_container_stats_objects metrics [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [20:09:52] (03CR) 10CDanis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [20:10:29] !log brett@cumin2003 START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, T435406] [20:10:34] T435406: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406 [20:10:35] !log brett@cumin2003 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: no reason specified, T435406] [20:10:45] (03PS1) 10Cwhite: monitoring groups: add asw1-60[34]-eqsin to monitoring groups [puppet] - 10https://gerrit.wikimedia.org/r/1329653 (https://phabricator.wikimedia.org/T418439) [20:14:21] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12259070 (10Andrew) >>! In T435537#12248577, @BLiviero-WMF wrote: > can you say more about what "optimizin... [20:20:01] (03PS1) 10Cathal Mooney: Drain HE circuits to codfw to take Arelion path instead. [homer/public] - 10https://gerrit.wikimedia.org/r/1329654 [20:20:01] (03PS1) 10Cathal Mooney: Eqsin changes following migration to Nokia switches [homer/public] - 10https://gerrit.wikimedia.org/r/1329655 (https://phabricator.wikimedia.org/T418439) [20:21:19] (03CR) 10CI reject: [V:04-1] Eqsin changes following migration to Nokia switches [homer/public] - 10https://gerrit.wikimedia.org/r/1329655 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [20:22:23] (03PS2) 10C. Scott Ananian: Simplify ParserCache configuration [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329649 [20:23:09] (03PS2) 10Cathal Mooney: Eqsin changes following migration to Nokia switches [homer/public] - 10https://gerrit.wikimedia.org/r/1329655 (https://phabricator.wikimedia.org/T418439) [20:24:20] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm [20:27:08] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cassandra-dev2002.codfw.wmnet with OS bookworm [20:28:11] !log cmooney@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet,cumin1003.eqiad.wmnet with reason: Homer - remove eqsin workaround from wmf-plugin for eqsin L3 switches - cmooney@cumin1003 [20:28:29] (03CR) 10Cathal Mooney: [C:03+2] Drain HE circuits to codfw to take Arelion path instead. [homer/public] - 10https://gerrit.wikimedia.org/r/1329654 (owner: 10Cathal Mooney) [20:29:49] (03Merged) 10jenkins-bot: Drain HE circuits to codfw to take Arelion path instead. [homer/public] - 10https://gerrit.wikimedia.org/r/1329654 (owner: 10Cathal Mooney) [20:29:50] !log cmooney@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet,cumin1003.eqiad.wmnet with reason: Homer - remove eqsin workaround from wmf-plugin for eqsin L3 switches - cmooney@cumin1003 [20:31:13] (03CR) 10Cathal Mooney: [C:03+2] Eqsin changes following migration to Nokia switches [homer/public] - 10https://gerrit.wikimedia.org/r/1329655 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [20:32:31] (03Merged) 10jenkins-bot: Eqsin changes following migration to Nokia switches [homer/public] - 10https://gerrit.wikimedia.org/r/1329655 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [20:32:44] (03PS1) 10Eevans: cassandra-dev: configure installer for partition reuse [puppet] - 10https://gerrit.wikimedia.org/r/1329659 (https://phabricator.wikimedia.org/T435395) [20:36:25] (03CR) 10Eevans: [C:03+2] cassandra-dev: configure installer for partition reuse [puppet] - 10https://gerrit.wikimedia.org/r/1329659 (https://phabricator.wikimedia.org/T435395) (owner: 10Eevans) [20:41:52] (03PS2) 10JHathaway: rake_modules: support Debian 13 (Trixie) facts [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [20:43:38] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thanks!!" [puppet] - 10https://gerrit.wikimedia.org/r/1329653 (https://phabricator.wikimedia.org/T418439) (owner: 10Cwhite) [20:44:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 18.39% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:44:29] (03PS1) 10CDanis: WIP: SLOP: gate ProbeDown on if a PoP is pooled [alerts] - 10https://gerrit.wikimedia.org/r/1329662 [20:44:47] (03CR) 10CI reject: [V:04-1] rake_modules: support Debian 13 (Trixie) facts [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [20:45:24] (03PS1) 10CDanis: geodns: also export pooledness within each PoP [puppet] - 10https://gerrit.wikimedia.org/r/1329663 [20:45:30] (03CR) 10CI reject: [V:04-1] WIP: SLOP: gate ProbeDown on if a PoP is pooled [alerts] - 10https://gerrit.wikimedia.org/r/1329662 (owner: 10CDanis) [20:48:33] !log cdobbins@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp5022.eqsin.wmnet with reason: needs repair [20:49:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 18.49% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:49:49] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in add_ip6_mapped [puppet] - 10https://gerrit.wikimedia.org/r/1329664 (https://phabricator.wikimedia.org/T435225) [20:49:56] RoanKattouw: I got distracted for a bit [20:50:33] (03PS2) 10JHathaway: Puppet 8: Replace legacy facts in add_ip6_mapped [puppet] - 10https://gerrit.wikimedia.org/r/1329664 (https://phabricator.wikimedia.org/T435225) [20:50:35] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329664 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [20:53:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aaron@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329616 (https://phabricator.wikimedia.org/T433522) (owner: 10Aaron Schulz) [20:54:02] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in mariadb::config [puppet] - 10https://gerrit.wikimedia.org/r/1329665 (https://phabricator.wikimedia.org/T435225) [20:54:30] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329665 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [20:57:23] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm [21:00:05] Deploy window Wikifunctions Services UTC Late (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T2100) [21:02:05] (03Merged) 10jenkins-bot: Rest: Rename "mode" key to "availability" in RestModuleOverrides [core] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329616 (https://phabricator.wikimedia.org/T433522) (owner: 10Aaron Schulz) [21:02:13] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, and 2 others: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12259282 (10cmooney) 05Open→03Resolved Everything is now done here. I'm not sure if I totally followed the plan or not as it was a hectic enoug... [21:02:29] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in profile::tlsproxy::envoy spec [puppet] - 10https://gerrit.wikimedia.org/r/1329667 (https://phabricator.wikimedia.org/T435225) [21:02:32] !log aaron@deploy1003 Started scap sync-world: Backport for [[gerrit:1329616|Rest: Rename "mode" key to "availability" in RestModuleOverrides (T433522)]] [21:02:37] T433522: REST vNext: consider renaming module "mode" to "availability". - https://phabricator.wikimedia.org/T433522 [21:02:44] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329667 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [21:04:48] (03PS2) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-08-19-122639 to 2026-08-19-165650 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329571 (https://phabricator.wikimedia.org/T435361) [21:04:48] (03PS1) 10Jforrester: wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-26-163757 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329668 (https://phabricator.wikimedia.org/T432291) [21:05:04] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-26-163757 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329668 (https://phabricator.wikimedia.org/T432291) (owner: 10Jforrester) [21:05:23] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in community-civi.my.cnf.erb [puppet] - 10https://gerrit.wikimedia.org/r/1329669 (https://phabricator.wikimedia.org/T435225) [21:06:19] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329669 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [21:07:18] !log aaron@deploy1003 aaron: Backport for [[gerrit:1329616|Rest: Rename "mode" key to "availability" in RestModuleOverrides (T433522)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:08:15] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [21:08:31] (03Merged) 10jenkins-bot: wikifunctions: Upgrade evaluators from 2026-08-19-122930 to 2026-08-26-163757 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329668 (https://phabricator.wikimedia.org/T432291) (owner: 10Jforrester) [21:09:19] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:10:15] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:10:26] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:10:55] !log aaron@deploy1003 aaron: Continuing with deployment [21:12:18] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12259361 (10cmooney) Lumen did not respond to my updates on the ticket and closed it as "resolved". I opened another ticket with them ref #35211090. [21:12:23] (03CR) 10SomeRandomDeveloper: Profiler: Fold excimer line parsing into one method (032 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329640 (owner: 10Krinkle) [21:12:27] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:12:45] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [21:13:16] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-08-19-122639 to 2026-08-19-165650 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329571 (https://phabricator.wikimedia.org/T435361) (owner: 10Jforrester) [21:13:23] (03PS5) 10Krinkle: Profiler: Fold excimer line parsing into one method [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329640 [21:13:28] (03CR) 10Krinkle: Profiler: Fold excimer line parsing into one method (032 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329640 (owner: 10Krinkle) [21:14:16] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:14:30] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:14:37] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:15:15] !log aaron@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329616|Rest: Rename "mode" key to "availability" in RestModuleOverrides (T433522)]] (duration: 12m 43s) [21:15:20] T433522: REST vNext: consider renaming module "mode" to "availability". - https://phabricator.wikimedia.org/T433522 [21:16:25] OK, done [21:17:06] (03Merged) 10jenkins-bot: wikifunctions: Upgrade orchestrator from 2026-08-19-122639 to 2026-08-19-165650 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329571 (https://phabricator.wikimedia.org/T435361) (owner: 10Jforrester) [21:18:59] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:19:17] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:19:28] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:20:07] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:20:14] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [21:20:47] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:20:54] (03CR) 10SomeRandomDeveloper: [C:03+1] "Looks good, didn't test in any way though" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329640 (owner: 10Krinkle) [21:21:28] (03CR) 10JHathaway: "@hashar@free.fr I pushed an alternative patch which instead bumps puppet to 7.34. This is not the puppet version we run, but there are no " [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [21:23:10] (03PS1) 10LWatson: Minimal Minerva: change label from "comments" to "discussion" [skins/MinervaNeue] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329672 (https://phabricator.wikimedia.org/T436141) [21:25:26] (03PS1) 10Jforrester: wikifunctions: Enable useReentrance in staging orchestrator (but not evaluators) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329674 (https://phabricator.wikimedia.org/T415616) [21:26:04] (03CR) 10Jforrester: [C:03+2] wikifunctions: Enable useReentrance in staging orchestrator (but not evaluators) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329674 (https://phabricator.wikimedia.org/T415616) (owner: 10Jforrester) [21:29:18] (03Merged) 10jenkins-bot: wikifunctions: Enable useReentrance in staging orchestrator (but not evaluators) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329674 (https://phabricator.wikimedia.org/T415616) (owner: 10Jforrester) [21:30:59] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:31:18] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:33:46] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1329540 (owner: 10Majavah) [21:34:58] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cassandra-dev2002.codfw.wmnet with OS bookworm [21:35:16] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm [21:36:19] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install pki2003 - https://phabricator.wikimedia.org/T436179 (10RobH) 03NEW [21:36:48] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install pki2003 - https://phabricator.wikimedia.org/T436179#12259483 (10RobH) [21:39:46] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install pki2003 - https://phabricator.wikimedia.org/T436179#12259502 (10RobH) a:03MoritzMuehlenhoff Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/SRE/Dc-operatio... [21:41:40] (03PS1) 10Jforrester: wikifunctions: Set ORCHESTRATOR_CALLBACK_API_URI in staging evaluators [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329675 (https://phabricator.wikimedia.org/T415616) [21:41:43] (03PS1) 10Jforrester: wikifunctions: Enable acceptCallbacks in staging orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329676 (https://phabricator.wikimedia.org/T415616) [21:41:44] 10ops-codfw, 06SRE, 06Data-Persistence, 06DC-Ops: Q1:rack/setup/install apus-be200[7-9] - https://phabricator.wikimedia.org/T436180 (10RobH) 03NEW [21:41:58] 10ops-codfw, 06SRE, 06Data-Persistence, 06DC-Ops: Q1:rack/setup/install apus-be200[7-9] - https://phabricator.wikimedia.org/T436180#12259527 (10RobH) [21:42:22] 10ops-codfw, 06SRE, 06Data-Persistence, 06DC-Ops: Q1:rack/setup/install apus-be200[7-9] - https://phabricator.wikimedia.org/T436180#12259529 (10RobH) a:03MatthewVernon Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/SRE/Dc-operations) an... [21:43:07] (03PS3) 10Cwhite: opensearch: change pki algo to rsa [puppet] - 10https://gerrit.wikimedia.org/r/1329625 (https://phabricator.wikimedia.org/T350516) [21:43:27] 10ops-eqiad, 06SRE, 06Data-Persistence, 06DC-Ops: Q1:rack/setup/install apus-be100[7-9] - https://phabricator.wikimedia.org/T436181 (10RobH) 03NEW [21:43:39] 10ops-eqiad, 06SRE, 06Data-Persistence, 06DC-Ops: Q1:rack/setup/install apus-be100[7-9] - https://phabricator.wikimedia.org/T436181#12259551 (10RobH) [21:44:05] (03CR) 10Jforrester: [C:03+2] wikifunctions: Set ORCHESTRATOR_CALLBACK_API_URI in staging evaluators [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329675 (https://phabricator.wikimedia.org/T415616) (owner: 10Jforrester) [21:44:09] 10ops-eqiad, 06SRE, 06Data-Persistence, 06DC-Ops: Q1:rack/setup/install apus-be100[7-9] - https://phabricator.wikimedia.org/T436181#12259552 (10RobH) a:03MatthewVernon Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/SRE/Dc-operations) an... [21:46:24] (03Merged) 10jenkins-bot: wikifunctions: Set ORCHESTRATOR_CALLBACK_API_URI in staging evaluators [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329675 (https://phabricator.wikimedia.org/T415616) (owner: 10Jforrester) [21:47:19] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:47:58] 10ops-eqsin: Inbound errors on interface cr2-eqsin:et-0/0/2 (Core: asw1-604-eqsin:ethernet-1/55) - https://phabricator.wikimedia.org/T436182 (10phaultfinder) 03NEW [21:48:07] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:50:46] !log jasmine@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1260.eqiad.wmnet [21:50:50] !log jasmine@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1260.eqiad.wmnet [21:51:22] !log jasmine@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1260.eqiad.wmnet [21:51:42] (03PS2) 10Jforrester: wikifunctions: Enable acceptCallbacks in staging orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329676 (https://phabricator.wikimedia.org/T415616) [21:51:43] (03PS1) 10Jforrester: wikifunctions: Switch ORCHESTRATOR_CALLBACK_API_URI to localhost:4974 in staging orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329678 (https://phabricator.wikimedia.org/T415616) [21:51:53] !log jasmine@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1260.eqiad.wmnet with OS trixie [21:51:59] (03CR) 10Jforrester: [C:03+2] wikifunctions: Switch ORCHESTRATOR_CALLBACK_API_URI to localhost:4974 in staging orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329678 (https://phabricator.wikimedia.org/T415616) (owner: 10Jforrester) [21:52:20] !log jasmine@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1260 [21:52:52] !log jasmine@cumin1003 START - Cookbook sre.dns.netbox [21:54:19] (03Merged) 10jenkins-bot: wikifunctions: Switch ORCHESTRATOR_CALLBACK_API_URI to localhost:4974 in staging orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329678 (https://phabricator.wikimedia.org/T415616) (owner: 10Jforrester) [21:55:10] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:55:48] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:57:33] (03CR) 10Jforrester: [C:03+2] wikifunctions: Enable acceptCallbacks in staging orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329676 (https://phabricator.wikimedia.org/T415616) (owner: 10Jforrester) [21:58:09] !log jasmine@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1260 - jasmine@cumin1003" [21:58:13] !log jasmine@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1260 - jasmine@cumin1003" [21:58:13] !log jasmine@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [21:58:13] !log jasmine@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1260.eqiad.wmnet 28.32.64.10.in-addr.arpa 8.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:58:17] !log jasmine@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1260.eqiad.wmnet 28.32.64.10.in-addr.arpa 8.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:58:17] !log jasmine@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1260 [21:59:00] !log jasmine@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1260 [21:59:00] !log jasmine@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1260 [21:59:56] (03Merged) 10jenkins-bot: wikifunctions: Enable acceptCallbacks in staging orchestrator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329676 (https://phabricator.wikimedia.org/T415616) (owner: 10Jforrester) [22:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260826T2200) [22:00:17] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [22:00:27] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [22:00:36] PROBLEM - Druid historical on an-druid1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [22:00:58] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [22:01:37] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [22:01:51] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [22:02:26] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [22:02:36] RECOVERY - Druid historical on an-druid1007 is OK: PROCS OK: 1 process with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [22:10:05] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [22:11:25] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12259690 (10Bethany) thanks so much @Eevans [22:20:16] !log jasmine@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1260.eqiad.wmnet with reason: host reimage [22:23:38] !log jasmine@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1260.eqiad.wmnet with reason: host reimage [22:24:37] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [22:29:17] (03PS1) 10Cwhite: pki: remove logging env ecdsa certificates [puppet] - 10https://gerrit.wikimedia.org/r/1329679 (https://phabricator.wikimedia.org/T350516) [22:29:50] RESOLVED: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [22:31:08] (03PS1) 10Cwhite: logging: remove ecdsa pki keys [labs/private] - 10https://gerrit.wikimedia.org/r/1329682 (https://phabricator.wikimedia.org/T350516) [22:33:30] (03CR) 10Cwhite: [V:03+2 C:03+2] logging: remove ecdsa pki keys [labs/private] - 10https://gerrit.wikimedia.org/r/1329682 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:33:36] (03CR) 10Cwhite: [C:03+2] pki: remove logging env ecdsa certificates [puppet] - 10https://gerrit.wikimedia.org/r/1329679 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:36:49] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [22:43:17] (03PS1) 10Cwhite: pki: add new BetaLogs_OpenSearch RSA CA for logging env [puppet] - 10https://gerrit.wikimedia.org/r/1329684 (https://phabricator.wikimedia.org/T350516) [22:45:03] (03PS1) 10Cwhite: logging: add new betalogs pki keys [labs/private] - 10https://gerrit.wikimedia.org/r/1329685 (https://phabricator.wikimedia.org/T350516) [22:45:55] (03CR) 10Cwhite: [V:03+2 C:03+2] logging: add new betalogs pki keys [labs/private] - 10https://gerrit.wikimedia.org/r/1329685 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:46:01] (03CR) 10Cwhite: [C:03+2] pki: add new BetaLogs_OpenSearch RSA CA for logging env [puppet] - 10https://gerrit.wikimedia.org/r/1329684 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:46:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:58:14] !log jasmine@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1260.eqiad.wmnet with OS trixie [23:01:14] jasmine@cumin1003 renumber-node (PID 709262) is awaiting input [23:01:20] eevans@cumin1003 reimage (PID 708126) is awaiting input [23:03:02] (03PS1) 10Cwhite: pki: rename betalogs cert [puppet] - 10https://gerrit.wikimedia.org/r/1329688 (https://phabricator.wikimedia.org/T350516) [23:07:40] (03CR) 10Cwhite: [C:03+2] pki: rename betalogs cert [puppet] - 10https://gerrit.wikimedia.org/r/1329688 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:10:50] !log jasmine@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1260.eqiad.wmnet [23:10:52] !log jasmine@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1260.eqiad.wmnet [23:10:54] !log jasmine@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1260.eqiad.wmnet [23:16:32] (03PS1) 10BryanDavis: mediawiki::deployment::server: Automate scope=pretrain deployments [puppet] - 10https://gerrit.wikimedia.org/r/1329690 (https://phabricator.wikimedia.org/T436178) [23:17:50] (03PS6) 10Dduvall: drivers: Driver registration and factory [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1326398 (https://phabricator.wikimedia.org/T434957) [23:17:50] (03PS1) 10Dduvall: drivers: Remove label attribute in favor of arguments [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329691 (https://phabricator.wikimedia.org/T434957) [23:18:30] (03PS4) 10Dduvall: drivers: Move all docker client calls to driver [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) [23:18:31] (03PS7) 10Dduvall: drivers: Driver registration and factory [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1326398 (https://phabricator.wikimedia.org/T434957) [23:18:31] (03PS2) 10Dduvall: drivers: Remove label attribute in favor of arguments [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329691 (https://phabricator.wikimedia.org/T434957) [23:20:39] (03PS2) 10BryanDavis: mediawiki::deployment::server: Automate scope=pretrain deployments [puppet] - 10https://gerrit.wikimedia.org/r/1329690 (https://phabricator.wikimedia.org/T436178) [23:21:26] (03CR) 10Dduvall: "I left it as is for now since I'm actually trying to implement an auth-less http registry option for local functional testing." [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [23:23:05] (03CR) 10CI reject: [V:04-1] drivers: Remove label attribute in favor of arguments [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329691 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [23:23:14] (03CR) 10BryanDavis: "As noted in T436178, this should not be merged until communications steps about enabling this new automated deployment have been completed" [puppet] - 10https://gerrit.wikimedia.org/r/1329690 (https://phabricator.wikimedia.org/T436178) (owner: 10BryanDavis) [23:25:28] (03PS1) 10Catrope: Use IDBAccessObject::READ_LATEST in some cases [extensions/OATHAuth] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329692 (https://phabricator.wikimedia.org/T436066) [23:25:32] (03CR) 10BryanDavis: mediawiki::deployment::server: Automate scope=pretrain deployments (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329690 (https://phabricator.wikimedia.org/T436178) (owner: 10BryanDavis) [23:26:03] (03CR) 10Cwhite: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1329416 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [23:26:13] (03CR) 10Cwhite: [C:03+1] "Looks good!" [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [23:29:57] (03PS3) 10Dduvall: drivers: Remove label attribute in favor of arguments [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329691 (https://phabricator.wikimedia.org/T434957) [23:32:42] (03CR) 10TrainBranchBot: [C:03+2] "Approved by catrope@deploy1003 using scap backport" [extensions/OATHAuth] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329692 (https://phabricator.wikimedia.org/T436066) (owner: 10Catrope) [23:34:16] (03Merged) 10jenkins-bot: Use IDBAccessObject::READ_LATEST in some cases [extensions/OATHAuth] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329692 (https://phabricator.wikimedia.org/T436066) (owner: 10Catrope) [23:34:40] !log catrope@deploy1003 Started scap sync-world: Backport for [[gerrit:1329692|Use IDBAccessObject::READ_LATEST in some cases (T436066)]] [23:34:44] T436066: Initial 2FA not created on otrs_wikiwiki - https://phabricator.wikimedia.org/T436066 [23:35:02] (03CR) 10Dduvall: drivers: Move all docker client calls to driver (031 comment) [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [23:35:18] (03CR) 10Dduvall: drivers: Driver registration and factory (031 comment) [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1326398 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [23:37:14] !log eevans@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cassandra-dev2002.codfw.wmnet with OS bookworm [23:39:43] !log catrope@deploy1003 catrope: Backport for [[gerrit:1329692|Use IDBAccessObject::READ_LATEST in some cases (T436066)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [23:39:48] T436066: Initial 2FA not created on otrs_wikiwiki - https://phabricator.wikimedia.org/T436066 [23:40:43] jouncebot: nowandnext [23:40:43] No deployments scheduled for the next 6 hour(s) and 19 minute(s) [23:40:44] In 6 hour(s) and 19 minute(s): MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T0600) [23:40:44] In 6 hour(s) and 19 minute(s): Primary database switchover (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T0600) [23:40:46] !log catrope@deploy1003 catrope: Continuing with deployment [23:41:09] RoanKattouw: once you're done, would you mind pinging me? [23:41:22] Will do, sorry for the unscheduled deploy [23:41:22] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1329693 [23:41:22] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1329693 (owner: 10TrainBranchBot) [23:41:26] I'll be done once this patch finishes [23:45:13] !log catrope@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329692|Use IDBAccessObject::READ_LATEST in some cases (T436066)]] (duration: 10m 33s) [23:45:18] T436066: Initial 2FA not created on otrs_wikiwiki - https://phabricator.wikimedia.org/T436066 [23:49:38] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1329693 (owner: 10TrainBranchBot) [23:56:48] 06SRE, 10Beta-Cluster-Infrastructure, 06Traffic: Name new CDN servers deployment-cp-(text|upload)0x instead of deployment-cache-(text|upload)0x - https://phabricator.wikimedia.org/T280393#12259942 (10bd808) [23:57:09] (03PS1) 10Dduvall: WIP buildx: New driver for building images via Buildx/BuildKit [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958)