[00:00:51] (03PS7) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [00:04:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:17:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [00:18:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [00:23:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:28:10] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:28:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:33:25] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:34:03] 10ops-eqiad, 06DC-Ops: Inbound errors on interface cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://phabricator.wikimedia.org/T435748 (10phaultfinder) 03NEW [00:38:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 16.05% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:58:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:00:05] (03Abandoned) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1328338 (owner: 10TrainBranchBot) [01:01:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:02:04] (03PS1) 10MusikAnimal: CodeMirror: enable JSONC mode [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328386 (https://phabricator.wikimedia.org/T127682) [01:02:40] PROBLEM - statsv Varnishkafka log producer on cp4043 is CRITICAL: PROCS CRITICAL: 3 processes with args /usr/bin/varnishkafka -S /etc/varnishkafka/statsv.conf https://wikitech.wikimedia.org/wiki/Analytics/Systems/Varnishkafka [01:03:40] RECOVERY - statsv Varnishkafka log producer on cp4043 is OK: PROCS OK: 1 process with args /usr/bin/varnishkafka -S /etc/varnishkafka/statsv.conf https://wikitech.wikimedia.org/wiki/Analytics/Systems/Varnishkafka [01:06:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:12:13] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1328387 [01:12:13] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1328387 (owner: 10TrainBranchBot) [01:30:34] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:33:32] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:36:38] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:37:34] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [01:54:57] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [01:57:10] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:02:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:05:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [02:08:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:30:39] FIRING: TransitBGPDown: Transit BGP session down between cr2-esams and Init7 (77.109.134.113) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=esams&var-device=cr2-esams:9804&var-bgp_group=Transit4&var-bgp_neighbor=Init7 - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [02:33:53] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:35:39] FIRING: [2x] TransitBGPDown: Transit BGP session down between cr2-esams and Init7 (2001:1620:1000::85) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [02:39:57] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:40:39] RESOLVED: [2x] TransitBGPDown: Transit BGP session down between cr2-esams and Init7 (2001:1620:1000::85) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [02:49:14] PROBLEM - OSPF status on cr2-magru is CRITICAL: OSPFv2: 2/3 UP : OSPFv3: 2/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:49:57] RESOLVED: [3x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:50:14] RECOVERY - OSPF status on cr2-magru is OK: OSPFv2: 3/3 UP : OSPFv3: 3/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:50:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [02:51:10] FIRING: BFDdown: BFD session down between cr2-magru and fe80::ee38:73ff:fee8:9c58 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:56:10] RESOLVED: BFDdown: BFD session down between cr2-magru and fe80::ee38:73ff:fee8:9c58 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:01:14] PROBLEM - OSPF status on cr2-magru is CRITICAL: OSPFv2: 3/3 UP : OSPFv3: 2/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:02:39] FIRING: CoreBGPDown: Core BGP session down between cr2-eqdfw and cr2-magru (195.200.68.153) - group Confed_magru - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-eqdfw:9804&var-bgp_group=Confed_magru&var-bgp_neighbor=cr2-magru - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:03:14] RECOVERY - OSPF status on cr2-magru is OK: OSPFv2: 3/3 UP : OSPFv3: 3/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:04:10] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:07:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between cr2-eqdfw and cr2-magru (195.200.68.153) - group Confed_magru - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:09:10] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:09:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:09:57] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [03:11:38] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [03:12:54] FIRING: [2x] CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:14:57] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [03:18:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [03:19:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:20:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:23:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:23:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [03:24:25] FIRING: [7x] BFDdown: BFD session down between cr1-eqiad and 195.200.68.137 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:25:40] RESOLVED: [4x] BFDdown: BFD session down between cr1-eqiad and 195.200.68.137 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:44:10] FIRING: BFDdown: BFD session down between cr1-eqiad and 195.200.68.137 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:46:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:46:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [03:49:10] RESOLVED: [2x] BFDdown: BFD session down between cr1-eqiad and 195.200.68.137 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:52:10] FIRING: BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [03:56:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [03:57:10] RESOLVED: [2x] BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:04:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:20:20] !log wikifunctionsclient_usage from s2 T434541 [05:20:24] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:20:25] T434541: Drop the local wikifunctionsclient_usage tables from all wikis, replaced by wikifunctions_usage / wikifunctions_usage_wiki on x1 - https://phabricator.wikimedia.org/T434541 [05:23:32] !log wikifunctionsclient_usage from s3 T434541 [05:23:36] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:26:12] (03PS1) 10Marostegui: db1282: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1328390 (https://phabricator.wikimedia.org/T407942) [05:26:51] (03CR) 10Marostegui: [C:03+2] db1282: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1328390 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [05:28:09] (03PS1) 10Marostegui: instances.yaml: Add db1282 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1328391 (https://phabricator.wikimedia.org/T407942) [05:45:16] (03PS1) 10KartikMistry: Update cxserver to 2026-08-24-032908-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328393 (https://phabricator.wikimedia.org/T213262) [05:46:41] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1282 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1328391 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [05:48:27] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1282 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96225 and previous config saved to /var/cache/conftool/dbconfig/20260824-054826-marostegui.json [05:48:32] T407942: Productionize db12[65-90] - https://phabricator.wikimedia.org/T407942 [05:48:41] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1282: Pool back [05:50:12] !log Stop mariadb on sanitarium s1,s3,s8,s5,x3 there will be lag on wikireplicas for those sections T407942 [05:50:13] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: Cloning [05:50:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:51:47] (03PS1) 10Marostegui: mariadb: Productionize db1269 [puppet] - 10https://gerrit.wikimedia.org/r/1328394 (https://phabricator.wikimedia.org/T407942) [05:54:57] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [05:54:57] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [05:56:02] (03CR) 10Marostegui: [C:03+2] mariadb: Productionize db1269 [puppet] - 10https://gerrit.wikimedia.org/r/1328394 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [05:57:54] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [05:59:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:04:10] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:18:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [06:18:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [06:33:04] !log wikifunctionsclient_usage from s5 T434541 [06:33:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:33:10] T434541: Drop the local wikifunctionsclient_usage tables from all wikis, replaced by wikifunctions_usage / wikifunctions_usage_wiki on x1 - https://phabricator.wikimedia.org/T434541 [06:33:45] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pool back [06:46:15] !log wikifunctionsclient_usage from s7 T434541 [06:46:19] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:46:20] T434541: Drop the local wikifunctionsclient_usage tables from all wikis, replaced by wikifunctions_usage / wikifunctions_usage_wiki on x1 - https://phabricator.wikimedia.org/T434541 [06:46:23] (03CR) 10Marostegui: [C:03+2] table-catalog.yaml: Remove wikifunctionsclient_usage [puppet] - 10https://gerrit.wikimedia.org/r/1328164 (https://phabricator.wikimedia.org/T434541) (owner: 10Marostegui) [06:46:26] (03PS1) 10Muehlenhoff: Add puppetserver200[56] and ganeti205[1-6] to site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1328398 (https://phabricator.wikimedia.org/T434698) [06:52:28] (03PS1) 10Marostegui: db1182: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1328399 (https://phabricator.wikimedia.org/T434869) [06:53:07] (03CR) 10Marostegui: [C:03+2] db1182: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1328399 (https://phabricator.wikimedia.org/T434869) (owner: 10Marostegui) [06:54:01] (03PS1) 10Marostegui: instances.yaml: Remove db1182 from dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1328400 (https://phabricator.wikimedia.org/T434869) [06:54:52] (03CR) 10Marostegui: [C:03+2] instances.yaml: Remove db1182 from dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1328400 (https://phabricator.wikimedia.org/T434869) (owner: 10Marostegui) [06:55:47] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove db1182 from dbctl T434869', diff saved to https://phabricator.wikimedia.org/P96236 and previous config saved to /var/cache/conftool/dbconfig/20260824-065547-marostegui.json [06:55:53] T434869: decommission db1182.eqiad.wmnet - https://phabricator.wikimedia.org/T434869 [06:59:07] (03PS1) 10Bartosz Wójtowicz: ml-services: Update revise-tone-task-generator. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328401 (https://phabricator.wikimedia.org/T433319) [06:59:10] (03CR) 10Muehlenhoff: [C:03+2] Add puppetserver200[56] and ganeti205[1-6] to site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1328398 (https://phabricator.wikimedia.org/T434698) (owner: 10Muehlenhoff) [07:00:05] Amir1, urbanecm, and awight: UTC morning backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T0700). Please do the needful. [07:00:05] Hide_on_rosie: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:24] o/ [07:01:01] well, i wanted to deploy this patch but since it would be my first time i'd prefer having someone supervise just in case... but it seems like the requester isn't here either [07:11:34] (03CR) 10Anzx: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) (owner: 10Anzx) [07:11:38] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [07:12:59] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 13Patch-For-Review: Q1:rack/setup/install ganeti205[1-6] - https://phabricator.wikimedia.org/T434698#12245271 (10MoritzMuehlenhoff) >>! In T434698#12244711, @Jhancock.wm wrote: > preseed has these servers in it, but site.pp doesn't. Sorry! Fixed [07:13:13] (03PS4) 10Anzx: Lift IP cap for Mapudungun editathon on 2026-08-29 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) [07:13:21] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 13Patch-For-Review: Q1:rack/setup/install (2) new puppetservers - https://phabricator.wikimedia.org/T432998#12245272 (10MoritzMuehlenhoff) >>! In T432998#12244726, @Jhancock.wm wrote: > these servers aren't in preseed or site.pp. but there is a... [07:13:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:14:20] o/ robertsky [07:14:25] i assume you're here for the enwiki patch? [07:14:32] Yes [07:14:34] (03PS1) 10Kevin Bazira: ml-services: update outlink isvc to return INVALID_ARGUMENT for client errors on gRPC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328402 (https://phabricator.wikimedia.org/T435586) [07:14:47] gotcha [07:14:57] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:15:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:16:56] Amir1 urbanecm (or anyone else here): either of you willing to tag along for the deploy? i'm sure this is a low-risk patch but https://wikitech.wikimedia.org/wiki/How_to_deploy_code#Other_Prerequisites tells me i should get someone with me for my first time [07:18:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [07:20:10] RESOLVED: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:21:21] (03CR) 10Bartosz Wójtowicz: [C:03+1] ml-services: update outlink isvc to return INVALID_ARGUMENT for client errors on gRPC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328402 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [07:21:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [07:22:53] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update outlink isvc to return INVALID_ARGUMENT for client errors on gRPC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328402 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [07:23:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:25:57] (03CR) 10AikoChou: [C:03+1] "LGTM!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328401 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [07:26:57] (03CR) 10Muehlenhoff: [C:03+2] Re-apply the urldownloader role to the new Trixie hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328166 (https://phabricator.wikimedia.org/T427282) (owner: 10Muehlenhoff) [07:30:52] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Update revise-tone-task-generator. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328401 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [07:32:17] (03PS1) 10Brouberol: Deduplicate the proxy configuration value [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328403 [07:34:35] 06SRE, 10Castor, 06cloud-services-team, 06Collaboration-Services, and 2 others: Completion of castor-save-workspace-cache stucks - https://phabricator.wikimedia.org/T435699#12245304 (10hashar) I worked around the issue by having the Castor agent to be connected via the instance IPv4 address, the hostname r... [07:34:41] chlod: I guess I can tag along for the deploy and hope it doesn’t take more than, say, 15 minutes 😅 [07:34:46] (because at some point I should be going to the office) [07:34:50] \o/ [07:34:53] surely it shouldn't [07:34:58] robertsky: still around? [07:35:08] if it’s a config change, should be fine [07:35:13] alrighty [07:35:18] it's not exactly testable either [07:35:35] our favorite kind [07:35:38] would only know if it works when we see accounts getting promoted on enwiki [07:35:40] :3 [07:35:43] (03CR) 10TrainBranchBot: [C:03+2] "Approved by chlod@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307270 (https://phabricator.wikimedia.org/T431060) (owner: 10Novem Linguae) [07:36:44] yeah looks like a very harmless change to me ^^ [07:36:57] chlod: are you using spiderpig or manual scap? [07:37:03] manual scap [07:37:08] ah ok [07:37:26] i thought it'd probably be better to do the more manual method first [07:37:40] I guess that makes sense under the “learn how to do manual deploys first, then don’t do them that way” rule that I think is somewhere on wiki ^^ [07:37:57] it's good practice in general [07:38:13] I am still around [07:38:26] awesome [07:40:06] morning [07:40:20] Gd afternoon [07:40:29] o/ [07:40:35] evening squire [07:41:17] why is that change taking so long to merge? [07:41:22] 06SRE, 10Castor, 06Collaboration-Services, 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team: Completion of castor-save-workspace-cache stucks - https://phabricator.wikimedia.org/T435699#12245330 (10taavi) [07:41:28] omg zuul ._. [07:41:31] yeah i just saw [07:41:33] youch [07:42:06] some of the codehealth changes have been there for 37 hr 36 min??? [07:42:08] i'm assuming 38 hours elapsed jobs aren't ideal... [07:42:29] it was apparently broken over the weekend [07:42:34] ah... [07:42:54] zuul should prioritize gate-and-submit, especially for deployment branches and repos, but that still requires an open slot from other jobs finishing [07:44:10] guess we're stuck here for a bit [07:44:47] oh dear [07:45:19] So no deploy? ^.^ [07:45:40] good question [07:45:45] hi, we got some outage/delay/oddity in CI which happened over the week-end [07:46:36] most jobs in Jenkins trigger a child job in charge of saving the package managers caches, and that child job could not run because Jenkins is unable to reach the WMCS instance over IPv6 [07:46:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:47:15] I have made Jenkins to use IPv4 for connection and that solved it. CI is now processing the backlog, including the branch cut for MediaWiki train [07:47:39] (03Merged) 10jenkins-bot: ml-services: update outlink isvc to return INVALID_ARGUMENT for client errors on gRPC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328402 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [07:47:42] jouncebot: next [07:47:42] In 2 hour(s) and 12 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1000) [07:48:04] chlod: i'm not in a hurry at least, so if you don't mind waiting it out we can still ship the patch [07:48:14] that'd be great :) [07:48:16] i'm also not in a hurry [07:48:27] robertsky might be but this seems low-risk enough [07:48:32] ah, it seems it's going through CI now [07:48:38] I am monitoring the chat as well. [07:48:42] \o/ [07:48:44] yeah, it's running now, shoudn't be too long [07:48:56] hashar: thanks for solving it <3 [07:49:05] Will be round for about another 3 hours [07:49:27] chlod: I have to go to the office now, probably back in ca. 20 minutes and maybe I can still help with the deploy then (otherwise yay for taavi being here) [07:49:40] sure thing, thanks for coming! [07:49:42] (03Merged) 10jenkins-bot: InitialiseSettings: change enwiki extendedconfirmed autopromote settings [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307270 (https://phabricator.wikimedia.org/T431060) (owner: 10Novem Linguae) [07:49:49] there we go [07:49:58] !log chlod@deploy1003 Started scap sync-world: Backport for [[gerrit:1307270|InitialiseSettings: change enwiki extendedconfirmed autopromote settings (T431060)]] [07:50:03] T431060: Extended confirmed age calculation on en-wiki - https://phabricator.wikimedia.org/T431060 [07:50:04] and I think I found the root cause (IPv6 connection from prod to WMCS is not allowed because the v6 network is not listed =) ) [07:50:20] (03Merged) 10jenkins-bot: ml-services: Update revise-tone-task-generator. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328401 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [07:50:20] chlod: are you `scap backport`ing or spiderpig'ing? [07:50:27] scap backporting [07:50:28] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1328383 (owner: 10TrainBranchBot) [07:51:10] FIRING: BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:51:16] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [07:51:22] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [07:52:15] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1328387 (owner: 10TrainBranchBot) [07:52:37] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [07:53:32] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [07:55:50] (03PS1) 10Bartosz Wójtowicz: changeprop: enable additional languages for revise-tone-task-generator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328405 (https://phabricator.wikimedia.org/T433319) [07:56:10] RESOLVED: BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [07:58:54] (03CR) 10AikoChou: [C:03+1] changeprop: enable additional languages for revise-tone-task-generator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328405 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [08:02:35] (03PS1) 10Dpogorzelski: kserve: add llmisvc-controller [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1328406 [08:03:30] chlod: fwiw usually the automated job to build wmf/next images very early every morning means a bunch of caches needed for the image build are there, but since that failed due to the CI issues, your build is taking much longer than usual :/ [08:03:51] yeah i've been staring at it for a while [08:04:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:05:25] if you want to be super prepared you can open logstash already [08:05:57] the two main dashboards you'll want are 'mwdebug servers' and 'mediawiki-NEW-errors', both in the second row on the links panel on the home page [08:06:39] got them open now [08:06:46] and seems like the build finally finished [08:06:56] yep [08:07:42] (03CR) 10Bartosz Wójtowicz: [C:03+2] changeprop: enable additional languages for revise-tone-task-generator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328405 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [08:07:46] and the names are hopefully pretty self explanatory, the mwdebug one will show all the logs coming from the test servers, and the other one will show errors from production that aren't long-term known logspam [08:08:30] we've got a lot of long-term known logspam 😅 [08:09:01] well most of those rules are disabled to see if they're truly fixed and waiting to be cleaned up, but yes [08:09:45] so the basic idea is that after testing on mwdebug you'll want to look at the mwdebug one (don't forget to refresh as you opened it before) to see if there's anything unusual, and then once the patch has rolled out double-check the new-errors one just in case there's something that wasn't noticed on mwdebug [08:10:14] knowing whether something is 'anything unusual' comes mostly from experience [08:10:32] gotcha [08:11:09] !log chlod@deploy1003 novemlinguae, chlod: Backport for [[gerrit:1307270|InitialiseSettings: change enwiki extendedconfirmed autopromote settings (T431060)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [08:11:14] T431060: Extended confirmed age calculation on en-wiki - https://phabricator.wikimedia.org/T431060 [08:11:15] there we go [08:11:46] Whee [08:12:08] indeed i see some stuff and it's not dying [08:12:23] can't really test the patch since it's an autopromote condition [08:12:29] i assume i go ahead with the sync now? [08:12:47] chlod: i'll let you deal with organizing the testing with the patch author, but lmk if you need assistence or opinions [08:13:10] o/ back [08:13:11] even if you can't test the actual logic working, it's generally a good idea to at least load some pages on mwdebug and check the wiki still loads fine [08:13:23] yeah i was doing that just now [08:13:31] seems to be working all fine [08:13:39] yeah, and logstash too? [08:13:49] yup, nothing out of the ordinary [08:13:59] indeed :) [08:14:04] so go ahead and say yes [08:14:09] !log chlod@deploy1003 novemlinguae, chlod: Continuing with deployment [08:14:41] the other log source you can open is `logspam-watch` on mwlog1003.eqiad.wmnet btw, I still find that more useful than the logstashes to quickly see if anything is particularly wrong [08:15:11] (03Merged) 10jenkins-bot: changeprop: enable additional languages for revise-tone-task-generator [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328405 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [08:15:30] (03PS1) 10Jgiannelos: prv: Enable parsoid rendering for 6 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 [08:15:52] will keep in mind [08:16:24] 06SRE, 10Castor, 06Collaboration-Services, 10Continuous-Integration-Infrastructure, 06Release-Engineering-Team: Completion of castor-save-workspace-cache stucks - https://phabricator.wikimedia.org/T435699#12245399 (10hashar) 05Open→03Resolved a:03hashar [08:16:35] (03PS2) 10Dpogorzelski: kserve: add llmisvc-controller [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1328406 [08:16:54] (I have a script for that: https://wikitech.wikimedia.org/wiki/Backport_windows/Deployers/Script) [08:17:51] !log bwojtowicz@deploy1003 helmfile [staging] START helmfile.d/services/changeprop: apply [08:17:55] that page still references `scap sync-file`? :P [08:18:00] !log bwojtowicz@deploy1003 helmfile [staging] DONE helmfile.d/services/changeprop: apply [08:19:01] !log bwojtowicz@deploy1003 helmfile [codfw] START helmfile.d/services/changeprop: sync [08:19:02] fwiw for chlod: `scap sync-file` is more or less obsolete now, `scap sync-world` might still rarely see some use (e.g. when manually deploying security patches), and `scap backport`/spiderpig have replaced most uses of both of those [08:19:22] erm, i'm seeing a go stacktrace, should i be worried? "mw-web-migration-eqiad failed" [08:19:24] !log bwojtowicz@deploy1003 helmfile [codfw] DONE helmfile.d/services/changeprop: sync [08:19:28] the era of having to think of ordering when backporting is thankfully in the distant past [08:19:29] not seeing anything severely bad in logstash [08:19:44] can you paste the full error? [08:19:54] 08:17:44 mw-web-migration-eqiad failed (exit status 2): (cd /srv/deployment-charts/helmfile.d/services/mw-web && helmfile -e eqiad --selector name=migration apply --context 5) [08:19:54] stdout: [08:19:54] stderr: skipping missing values file matching "/etc/helmfile-defaults/private/main_services/mw-web/eqiad.yaml" [08:19:54] panic: runtime error: invalid memory address or nil pointer dereference [08:19:54] [signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0x47ac8c] [08:20:25] file a task for serviceops but don't worry about it further, I'd say [08:20:31] gotcha [08:20:42] yeah the k8s deployment has 0 failures so [08:21:04] <_joe_> it's possible that that specific release is running an old version of the code though [08:21:10] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org [08:21:15] <_joe_> and wow, a segfault in helmfile [08:22:01] !log bwojtowicz@deploy1003 helmfile [eqiad] START helmfile.d/services/changeprop: sync [08:22:07] _joe_: per https://gerrit.wikimedia.org/g/operations/deployment-charts/+/master/helmfile.d/services/mw-web/values-migration.yaml mw-web-migration is thankfully a whole 0 replicas that could be outdated, unless i'm reading that wrong [08:22:19] ah ok [08:22:26] !log bwojtowicz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/changeprop: sync [08:22:29] I was gonna say maybe do a scap sync-world to ensure everything’s up to date [08:22:31] <_joe_> I was about to ask :) [08:22:32] but then maybe not needed [08:22:48] also `kube_env mw-web eqiad ; kubectl get pod` shows no 'migration' pods [08:23:10] <_joe_> well that is expected given the zero replicas :) [08:23:36] (03CR) 10Atsuko: [C:03+2] dse-k8s-eqiad: provision new airflow namespace [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327551 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [08:24:38] yeah but i just wanted to check that the zero i found in that fiel wasnät being overridden by some other file i wasn't aware of [08:24:57] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org [08:25:34] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [08:25:40] taavi: wiki page updated :P thanks [08:26:44] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [08:27:16] chlod: I guess you could !log that the UTC morning backport+config window is done? (optional but I like to do it ^^) [08:27:28] actually, did that go segfault mean that the successful deployment wasn’t logged? I don’t see it in SAL [08:27:36] err [08:27:39] i just got asked [08:27:40] only the “continuing with deployment” and then unrelated stuff [08:27:45] 08:27:26 K8s deployment to stage production failed: K8s Deployment had the following errors: [08:27:45] Deployment of mw-web-migration-eqiad failed: The deployment of mw-web-migration-eqiad failed [08:27:45] K8s deployment to stage production failed. [08:27:45] [r] Retry deployment [08:27:45] [b] Roll back all stages and terminate [08:27:45] What do you want to do? (default: [r]): [08:27:48] ah [08:27:56] retry, I think. taavi? [08:27:56] retry seems ok [08:27:59] gotcha [08:28:13] (awkward that retry and rollback both start with r) [08:28:29] so probably scap will !log it after the retry then [08:28:34] yup [08:28:40] this time it's moving much faster [08:28:50] (03CR) 10Jelto: [C:03+1] "lgtm, thank you" [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [08:28:58] !log chlod@deploy1003 Finished scap sync-world: Backport for [[gerrit:1307270|InitialiseSettings: change enwiki extendedconfirmed autopromote settings (T431060)]] (duration: 39m 00s) [08:29:01] there we go :) [08:29:02] T431060: Extended confirmed age calculation on en-wiki - https://phabricator.wikimedia.org/T431060 [08:29:03] Lucas_WMDE: tbh those options beats the `scap train` utility which has numbered options, so testwikis is option 0 and group0 is option 1 and so on [08:29:33] lol [08:29:46] that'd be bad to brainfart on [08:30:00] in any case taavi, Lucas_WMDE thanks so much y'all for the help :D [08:30:28] trial by fire I see :) [08:30:33] :D [08:30:34] yw. [08:30:36] congrats on your first deploy [08:30:43] 🎉 [08:30:50] drinks on me next time [08:31:00] NovemLinguae: the hard part of deploying is knowing what to do when it goes wrong, so not a bad thing i'd say :P [08:31:02] Chug the first glass down [08:31:26] it kinda somewhat went wrong but it was a very recoverable wrong [08:31:33] (03CR) 10Muehlenhoff: [C:03+2] jenkins:: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328169 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [08:31:49] hopefully that segfault won't confuse someone else [08:31:59] famous last words [08:32:11] LOL [08:32:29] did you file a task for that? [08:32:50] have not yet [08:32:52] (03PS1) 10JavierMonton: subject: html-kafka-offset-lag alert [alerts] - 10https://gerrit.wikimedia.org/r/1328530 [08:33:32] (03Merged) 10jenkins-bot: dse-k8s-eqiad: provision new airflow namespace [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327551 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [08:34:11] (03CR) 10Federico Ceratto: "You mean the files in hieradata like hieradata/hosts/db2901.yaml? They were added in https://gerrit.wikimedia.org/r/c/operations/puppet/+/" [puppet] - 10https://gerrit.wikimedia.org/r/1327586 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [08:35:25] (03CR) 10Marostegui: [C:03+1] "Ah thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1327586 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [08:36:30] (03CR) 10Dpogorzelski: [C:03+2] kserve: add llmisvc-controller [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1328406 (owner: 10Dpogorzelski) [08:38:17] (03CR) 10Dpogorzelski: [V:03+2 C:03+2] kserve: add llmisvc-controller [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1328406 (owner: 10Dpogorzelski) [08:39:52] (03CR) 10Mpostoronca: [C:03+1] ModelToRun: Add getContentPolicyName() before getModelName() changes [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328280 (https://phabricator.wikimedia.org/T432848) (owner: 10Dreamy Jazz) [08:42:19] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [08:43:20] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [08:44:09] !log atsuko@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [08:45:03] !log atsuko@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [08:45:50] (03PS1) 10Marostegui: redact_sanitarium.sh: Add db1269 [puppet] - 10https://gerrit.wikimedia.org/r/1328531 (https://phabricator.wikimedia.org/T407942) [08:45:56] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [08:46:43] (03CR) 10Marostegui: [C:03+2] redact_sanitarium.sh: Add db1269 [puppet] - 10https://gerrit.wikimedia.org/r/1328531 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [08:46:54] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [08:48:37] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [08:49:30] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [08:53:20] (03CR) 10Joal: [C:03+1] "LGTM!" [alerts] - 10https://gerrit.wikimedia.org/r/1328530 (owner: 10JavierMonton) [08:55:06] (03PS1) 10Muehlenhoff: Enable cumin alias check on cumin2003 [puppet] - 10https://gerrit.wikimedia.org/r/1328533 (https://phabricator.wikimedia.org/T427897) [08:57:43] (03PS1) 10Blake: mw-*: Upgrade to envoy 1.39.0 in the MW canary releases and mw-debug. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328532 (https://phabricator.wikimedia.org/T421418) [08:59:57] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:01:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [09:02:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:03:14] (03PS1) 10Marostegui: wmen: Add backup aliases [dns] - 10https://gerrit.wikimedia.org/r/1328535 [09:04:04] (03CR) 10Marostegui: "@cwilliams@wikimedia.org what do you think of this CNAMEs?" [dns] - 10https://gerrit.wikimedia.org/r/1328535 (owner: 10Marostegui) [09:04:06] (03PS2) 10Jgiannelos: prv: Enable parsoid rendering for 6 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 [09:05:00] (03PS1) 10Kevin Bazira: ml-services: fix numpy/fasttext dependency conflict in outlink isvc [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328536 (https://phabricator.wikimedia.org/T435586) [09:07:10] RESOLVED: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:07:18] (03CR) 10Federico Ceratto: [C:03+2] site.pp: Set MariaDB role for db190[1-3] db290[1-2] [puppet] - 10https://gerrit.wikimedia.org/r/1327586 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [09:08:52] (03CR) 10Kevin Bazira: [C:03+2] ml-services: fix numpy/fasttext dependency conflict in outlink isvc [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328536 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [09:11:16] (03CR) 10Giuseppe Lavagetto: [C:03+1] mw-*: Upgrade to envoy 1.39.0 in the MW canary releases and mw-debug. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328532 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [09:11:46] (03Merged) 10jenkins-bot: ml-services: fix numpy/fasttext dependency conflict in outlink isvc [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328536 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [09:18:36] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [09:20:50] (03CR) 10Lucas Werkmeister (WMDE): [C:03+1] Lift IP cap for Mapudungun editathon on 2026-08-29 (032 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) (owner: 10Anzx) [09:24:48] (03CR) 10Muehlenhoff: [C:03+2] Enable cumin alias check on cumin2003 [puppet] - 10https://gerrit.wikimedia.org/r/1328533 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [09:25:22] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [09:26:02] !log kevinbazira@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [09:28:58] !log test [09:29:01] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:29:59] !log [09:25] kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [09:30:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:30:08] !log [09:26] kevinbazira@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [09:30:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:30:27] ^ repeated !logs during stashbot maintenance (T435223) [09:30:28] T435223: Migrate storage to new OpenSearch v2 Toolforge cluster - https://phabricator.wikimedia.org/T435223 [09:31:27] (03PS1) 10Brouberol: airflow: move the datahub connectiom block to the environment values [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328538 [09:32:20] (03CR) 10Gmodena: [C:03+2] Update the index S3 bucket for staging only. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328237 (https://phabricator.wikimedia.org/T435674) (owner: 10Lerickson) [09:32:48] (03PS1) 10Majavah: P:toolforge::elasticsearch: Disable access to old cluster [puppet] - 10https://gerrit.wikimedia.org/r/1328539 (https://phabricator.wikimedia.org/T401818) [09:36:54] (03Merged) 10jenkins-bot: Update the index S3 bucket for staging only. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328237 (https://phabricator.wikimedia.org/T435674) (owner: 10Lerickson) [09:38:53] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:40:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [09:43:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:45:09] (03PS2) 10Brouberol: airflow: move the datahub connection block to the environment values [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328538 [09:46:05] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [09:46:05] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) [09:46:06] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db2901.codfw.wmnet [09:47:09] (03PS1) 10Giuseppe Lavagetto: aptrepo: add thirdparty/gvisor to newer distros as well [puppet] - 10https://gerrit.wikimedia.org/r/1328540 (https://phabricator.wikimedia.org/T435758) [09:47:43] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [09:49:57] FIRING: [3x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:52:29] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [09:53:21] (03CR) 10Ladsgroup: "Adding Luca based on git blame 😄" [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) (owner: 10Ladsgroup) [09:54:57] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [09:58:00] fceratto@cumin1003 netbox (PID 2808180) is awaiting input [09:58:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:59:49] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [09:59:50] 10SRE-swift-storage, 10Cloud-VPS (Debian Bullseye Deprecation): Migrate swift away from Debian Bullseye to Bookworm/Trixie - https://phabricator.wikimedia.org/T435783#12246060 (10Peachey88) [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1000) [10:00:14] (03PS1) 10Ladsgroup: Expose the new thumbnail endpoint on dewiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328542 (https://phabricator.wikimedia.org/T427465) [10:00:15] i plan to start a helmfile-only deploy momentarily in order to upgrade envoy in the canaries and mw-debug [10:00:28] (03CR) 10Blake: [C:03+2] mw-*: Upgrade to envoy 1.39.0 in the MW canary releases and mw-debug. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328532 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:00:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [10:03:08] (03PS1) 10Fabfur: WIP: kapow_score support [puppet] - 10https://gerrit.wikimedia.org/r/1328543 [10:03:42] (03Merged) 10jenkins-bot: mw-*: Upgrade to envoy 1.39.0 in the MW canary releases and mw-debug. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328532 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:04:13] (03PS1) 10MVernon: ceph: update to upstream 18.2.8 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1328544 (https://phabricator.wikimedia.org/T428385) [10:04:23] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db[12]90[1-3] ipv6 addr - fceratto@cumin1003" [10:04:26] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Delete db[12]90[1-3] ipv6 addr - fceratto@cumin1003" [10:04:27] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:05:55] !log blake@deploy1003 Started scap sync-world: upgrade envoy in debug and canary for T421418 [10:06:01] T421418: Upgrade Envoy to v1.39.0 - https://phabricator.wikimedia.org/T421418 [10:06:17] (03PS2) 10Fabfur: WIP: kapow_score support [puppet] - 10https://gerrit.wikimedia.org/r/1328543 (https://phabricator.wikimedia.org/T435799) [10:06:58] 10SRE-swift-storage, 10Cloud-VPS (Debian Bullseye Deprecation): Migrate swift away from Debian Bullseye to Bookworm/Trixie - https://phabricator.wikimedia.org/T435783#12246113 (10MatthewVernon) a:03Eevans [the remaining older systems in this project are cassandra-related] [10:10:16] (03CR) 10FNegri: [C:03+1] P:toolforge::elasticsearch: Disable access to old cluster [puppet] - 10https://gerrit.wikimedia.org/r/1328539 (https://phabricator.wikimedia.org/T401818) (owner: 10Majavah) [10:11:16] !log blake@deploy1003 Finished scap sync-world: upgrade envoy in debug and canary for T421418 (duration: 06m 46s) [10:11:21] T421418: Upgrade Envoy to v1.39.0 - https://phabricator.wikimedia.org/T421418 [10:12:34] (03CR) 10MVernon: "I've done a test build locally, which worked." [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1328544 (https://phabricator.wikimedia.org/T428385) (owner: 10MVernon) [10:15:36] (03CR) 10Marostegui: [C:03+1] ceph: update to upstream 18.2.8 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1328544 (https://phabricator.wikimedia.org/T428385) (owner: 10MVernon) [10:17:53] (03CR) 10Brouberol: [C:03+2] airflow: move the datahub connection block to the environment values [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328538 (owner: 10Brouberol) [10:18:05] (03PS3) 10Jgiannelos: prv: Enable parsoid rendering for 6 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 [10:18:38] (03CR) 10MVernon: [V:03+2 C:03+2] ceph: update to upstream 18.2.8 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1328544 (https://phabricator.wikimedia.org/T428385) (owner: 10MVernon) [10:18:51] (03CR) 10Hnowlan: [C:03+1] "lgtm - could you drop some stub credentials for beta into labs-private so we can see some pcc on future changes please?" [puppet] - 10https://gerrit.wikimedia.org/r/1328244 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [10:19:44] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply [10:19:57] FIRING: [3x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:20:12] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply [10:20:41] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply [10:21:19] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply [10:22:03] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply [10:22:37] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply [10:23:06] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply [10:23:44] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply [10:24:41] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply [10:25:03] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply [10:27:01] (03PS1) 10Blake: rest-gateway: Upgrade to envoy 1.39.0-1 in production. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328549 (https://phabricator.wikimedia.org/T421418) [10:27:58] (03CR) 10Hnowlan: prometheus: configure elasticsearch exporter on disable_security_plugin state (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [10:33:25] (03PS1) 10Muehlenhoff: docker::baseimages: Switch to /srv/debuerreotype as the build directory [puppet] - 10https://gerrit.wikimedia.org/r/1328551 (https://phabricator.wikimedia.org/T417389) [10:34:10] FIRING: BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:35:28] (03PS1) 10Ladsgroup: thumbor: Limit map (memory map) in image magick [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328552 [10:35:58] (03PS2) 10Blake: rest-gateway: Upgrade to envoy 1.39.0-1 in production. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328549 (https://phabricator.wikimedia.org/T421418) [10:37:28] (03CR) 10JMeybohm: [C:03+2] coredns: Remove loop from commonplugins [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328163 (https://phabricator.wikimedia.org/T427864) (owner: 10JMeybohm) [10:38:34] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [10:39:10] RESOLVED: BFDdown: BFD session down between cr1-eqiad and fe80::b6f9:5d07:c30:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:39:14] (03CR) 10Hnowlan: [C:03+1] thumbor: Limit map (memory map) in image magick [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328552 (owner: 10Ladsgroup) [10:39:20] (03PS1) 10Muehlenhoff: Move the build of the base images from build2001 to build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1328554 (https://phabricator.wikimedia.org/T417389) [10:46:11] (03CR) 10Majavah: [C:03+2] P:toolforge::elasticsearch: Disable access to old cluster [puppet] - 10https://gerrit.wikimedia.org/r/1328539 (https://phabricator.wikimedia.org/T401818) (owner: 10Majavah) [10:46:17] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Alerting: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801 (10cmooney) 03NEW p:05Triage→03Low [10:47:00] (03Merged) 10jenkins-bot: coredns: Remove loop from commonplugins [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328163 (https://phabricator.wikimedia.org/T427864) (owner: 10JMeybohm) [10:49:30] (03CR) 10Ladsgroup: [C:03+2] thumbor: Limit map (memory map) in image magick [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328552 (owner: 10Ladsgroup) [10:49:45] (03CR) 10CWilliams: "LGTM" [dns] - 10https://gerrit.wikimedia.org/r/1328535 (owner: 10Marostegui) [10:50:02] (03CR) 10Marostegui: [C:03+2] wmen: Add backup aliases [dns] - 10https://gerrit.wikimedia.org/r/1328535 (owner: 10Marostegui) [10:50:07] !log marostegui@dns1004 START - running authdns-update [10:51:58] (03Merged) 10jenkins-bot: thumbor: Limit map (memory map) in image magick [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328552 (owner: 10Ladsgroup) [10:52:28] !log marostegui@dns1004 END - running authdns-update [10:53:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [10:53:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [10:54:09] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [10:54:18] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [10:54:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [10:56:23] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:57:10] FIRING: [2x] BFDdown: BFD session down between cr1-eqiad and 195.200.68.137 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:58:41] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [10:59:37] (03PS1) 10Cathal Mooney: Re-enable HE transport from codfw to magru [homer/public] - 10https://gerrit.wikimedia.org/r/1328558 (https://phabricator.wikimedia.org/T435543) [11:00:02] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [11:01:40] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328554 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:02:10] RESOLVED: [2x] BFDdown: BFD session down between cr1-eqiad and 195.200.68.137 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:03:21] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12246220 (10Aklapper) Has this been reported to archive.org? What does `server does not respond.` mean exactly? I have a hard time to imagine that Commons or English Wikipedia... [11:04:49] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [11:05:32] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [11:08:24] (03PS1) 10Jelto: wikikube production: Update to cert-manager 1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328560 (https://phabricator.wikimedia.org/T427402) [11:09:26] (03CR) 10Cathal Mooney: [C:03+2] Re-enable HE transport from codfw to magru [homer/public] - 10https://gerrit.wikimedia.org/r/1328558 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [11:09:43] !log installing Bind security updates (client-side libs/tools) [11:09:45] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:10:56] (03Merged) 10jenkins-bot: Re-enable HE transport from codfw to magru [homer/public] - 10https://gerrit.wikimedia.org/r/1328558 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [11:13:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr1-magru and cr1-eqiad (195.200.68.136) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [11:14:57] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:16:18] (03CR) 10JavierMonton: [C:03+2] subject: html-kafka-offset-lag alert [alerts] - 10https://gerrit.wikimedia.org/r/1328530 (owner: 10JavierMonton) [11:18:00] (03Merged) 10jenkins-bot: subject: html-kafka-offset-lag alert [alerts] - 10https://gerrit.wikimedia.org/r/1328530 (owner: 10JavierMonton) [11:21:43] !log re-enable codfw-magru traffic path over HE T435543 [11:21:48] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:21:49] T435543: Aug 2026: Packet loss on new HE transports to magru - https://phabricator.wikimedia.org/T435543 [11:21:54] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply [11:22:25] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply [11:24:17] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply [11:24:39] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply [11:25:20] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply [11:25:53] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply [11:26:13] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply [11:26:36] (03CR) 10Elukey: [C:03+1] Move the build of the base images from build2001 to build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1328554 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:26:49] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply [11:26:54] (03CR) 10JMeybohm: [C:04-1] wikikube production: Update to cert-manager 1.19.6 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328560 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [11:27:31] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply [11:28:04] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply [11:28:12] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply [11:28:40] (03CR) 10Elukey: docker::baseimages: Switch to /srv/debuerreotype as the build directory (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328551 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:28:48] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply [11:31:42] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply [11:32:16] (03CR) 10Muehlenhoff: docker::baseimages: Switch to /srv/debuerreotype as the build directory (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328551 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:32:17] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply [11:32:20] (03PS2) 10Jelto: wikikube production: Update to cert-manager 1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328560 (https://phabricator.wikimedia.org/T427402) [11:32:51] (03CR) 10Jelto: wikikube production: Update to cert-manager 1.19.6 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328560 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [11:33:06] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply [11:33:33] Deploying cxserver.. [11:33:37] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply [11:33:46] (03CR) 10KartikMistry: [C:03+2] Update cxserver to 2026-08-24-032908-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328393 (https://phabricator.wikimedia.org/T213262) (owner: 10KartikMistry) [11:33:50] (03PS1) 10Cathal Mooney: magru: change ospf metric on codfw link. [homer/public] - 10https://gerrit.wikimedia.org/r/1328564 (https://phabricator.wikimedia.org/T435543) [11:35:10] (03CR) 10Elukey: "Looks good! Keep in mind that you'll also need to do something like https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1308121 (I'd sug" [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) (owner: 10Ladsgroup) [11:35:15] (03PS1) 10JMeybohm: coredns: Enable kubepods plugin on all clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328565 (https://phabricator.wikimedia.org/T428573) [11:36:23] RESOLVED: CertAlmostExpired: gNMI TLS certificate for lsw1-c4-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [11:36:31] (03PS1) 10Marostegui: check_private_data_report: Add db1269 [puppet] - 10https://gerrit.wikimedia.org/r/1328566 (https://phabricator.wikimedia.org/T407942) [11:36:41] (03Merged) 10jenkins-bot: Update cxserver to 2026-08-24-032908-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328393 (https://phabricator.wikimedia.org/T213262) (owner: 10KartikMistry) [11:38:37] !log kartik@deploy1003 helmfile [staging] START helmfile.d/services/cxserver: apply [11:39:00] !log kartik@deploy1003 helmfile [staging] DONE helmfile.d/services/cxserver: apply [11:39:03] jouncebot: nowandnext [11:39:04] No deployments scheduled for the next 1 hour(s) and 20 minute(s) [11:39:04] In 1 hour(s) and 20 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1300) [11:40:11] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328542 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [11:41:21] (03Merged) 10jenkins-bot: Expose the new thumbnail endpoint on dewiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328542 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [11:42:16] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1328542|Expose the new thumbnail endpoint on dewiki (T427465)]] [11:42:22] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [11:45:10] !log kartik@deploy1003 helmfile [codfw] START helmfile.d/services/cxserver: apply [11:45:43] !log kartik@deploy1003 helmfile [codfw] DONE helmfile.d/services/cxserver: apply [11:45:59] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip_encapsulation for apertium in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1328348 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [11:46:03] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip_encapsulation for apertium in eqiad. [puppet] - 10https://gerrit.wikimedia.org/r/1328349 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [11:46:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:47:01] !log kartik@deploy1003 helmfile [eqiad] START helmfile.d/services/cxserver: apply [11:47:26] (03CR) 10JMeybohm: [C:03+2] coredns: Enable kubepods plugin on all clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328565 (https://phabricator.wikimedia.org/T428573) (owner: 10JMeybohm) [11:47:30] !log kartik@deploy1003 helmfile [eqiad] DONE helmfile.d/services/cxserver: apply [11:48:19] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1328542|Expose the new thumbnail endpoint on dewiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:48:22] !log installing php7.4 security updates [11:48:23] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [11:48:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:48:28] !log jayme@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [11:49:25] !log jayme@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [11:49:34] !log jayme@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [11:50:08] !log Updated cxserver to 2026-08-24-032908-production (T213262, T213277, T253501, T428612) [11:50:18] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:50:19] T213262: CX2: "citation needed" template adapted with unnecessary HTML markup and VE attributes - https://phabricator.wikimedia.org/T213262 [11:50:19] T213277: CX2: Should not substitute citation templates when publishing - https://phabricator.wikimedia.org/T213277 [11:50:20] T253501: CX2: Tripled templates + bdi tags - https://phabricator.wikimedia.org/T253501 [11:50:20] T428612: ContentTranslation cannot translate reused reference - https://phabricator.wikimedia.org/T428612 [11:50:26] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [11:50:31] !log jayme@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [11:50:42] !log jayme@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [11:51:17] !log jayme@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [11:52:28] !log jayme@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [11:52:37] !log jayme@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [11:53:43] !log jayme@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [11:53:51] (03CR) 10JMeybohm: [C:03+1] wikikube production: Update to cert-manager 1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328560 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [11:53:55] !log jayme@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [11:54:09] !log jayme@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [11:56:50] !log jayme@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [11:56:58] !log jayme@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [11:57:51] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328542|Expose the new thumbnail endpoint on dewiki (T427465)]] (duration: 15m 34s) [11:57:55] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [11:58:33] !log installing ruby2.7 security updates [11:58:36] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:00:25] (03PS1) 10Atsuko: idp: add the airflow_experiment_platform service [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) [12:04:17] (03PS1) 10Atsuko: provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1328576 (https://phabricator.wikimedia.org/T416709) [12:04:28] (03CR) 10Atsuko: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:08:31] !log jayme@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [12:08:40] !log jayme@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [12:08:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:10:51] !log jayme@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [12:10:59] !log jayme@deploy1003 helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. [12:11:53] (03CR) 10JMeybohm: [C:03+1] profile::docker_registry: create the "main" docker distribution instance [puppet] - 10https://gerrit.wikimedia.org/r/1327549 (https://phabricator.wikimedia.org/T435499) (owner: 10Elukey) [12:12:41] !log jayme@deploy1003 helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. [12:12:49] !log jayme@deploy1003 helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. [12:13:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:17:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:19:13] PROBLEM - Host ms-fe2012 is DOWN: CRITICAL - Time to live exceeded (10.192.37.24) [12:19:13] PROBLEM - Host db2236 #page is DOWN: CRITICAL - Time to live exceeded (10.192.37.6) [12:19:32] RECOVERY - Host db2236 #page is UP: PING OK - Packet loss = 0%, RTA = 32.95 ms [12:19:35] RECOVERY - Host ms-fe2012 is UP: PING OK - Packet loss = 0%, RTA = 32.87 ms [12:19:40] !log jayme@deploy1003 helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. [12:19:48] !log jayme@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. [12:20:18] (03CR) 10CI reject: [V:04-1] Localisation updates from https://translatewiki.net. [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1328579 (owner: 10L10n-bot) [12:20:27] what happed there with db2236, network loop? [12:20:40] !log jayme@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. [12:20:48] !log jayme@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:20:48] the server is up since 27 days, seems like the BFDdown error above [12:20:51] topranks: ^ [12:20:55] right [12:20:58] topranks: was that some kind of maintenance or error? [12:21:31] there is at least nothing in the vendor maint calendar [12:22:11] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:24:34] (03PS4) 10Jgiannelos: prv: Enable parsoid rendering for 6 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 [12:25:00] (03PS1) 10PipelineBot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328581 [12:25:47] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [12:27:22] !log jayme@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:27:30] !log jayme@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [12:27:40] FIRING: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:28:22] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [12:28:31] !incidents [12:28:32] 8288 (RESOLVED) Host db2236 (paged) [12:28:55] !log jayme@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [12:28:59] 10ops-eqiad, 06DC-Ops: Outbound errors on interface cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://phabricator.wikimedia.org/T435810 (10phaultfinder) 03NEW [12:29:02] !log jayme@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. [12:30:20] !log jayme@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:30:27] !log jayme@deploy1003 helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. [12:30:46] (03CR) 10Elukey: [C:03+2] setup.py: add setuptools to tests env [software/spicerack] - 10https://gerrit.wikimedia.org/r/1326858 (owner: 10Elukey) [12:30:54] (03CR) 10Elukey: [C:03+2] profile::docker_registry: create the "main" docker distribution instance [puppet] - 10https://gerrit.wikimedia.org/r/1327549 (https://phabricator.wikimedia.org/T435499) (owner: 10Elukey) [12:31:47] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:31:48] !log jayme@deploy1003 helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. [12:32:25] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:32:45] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:33:29] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 7/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:34:34] (03PS1) 10Atsuko: Add analytics-experiment system user and groups [puppet] - 10https://gerrit.wikimedia.org/r/1328582 (https://phabricator.wikimedia.org/T416709) [12:34:57] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:35:17] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [12:35:35] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [12:36:59] 10ops-eqiad, 06DC-Ops, 10Kafka-Infrastructure: Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12246522 (10brouberol) [12:37:25] RESOLVED: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:37:59] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:38:06] (03PS1) 10Tiziano Fogli: prometheus/jobunavailable: double the alert firing time [alerts] - 10https://gerrit.wikimedia.org/r/1328587 [12:38:09] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [12:38:49] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [12:39:02] (03CR) 10Jelto: [C:03+2] wikikube production: Update to cert-manager 1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328560 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [12:39:27] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:41:39] (03PS2) 10Cathal Mooney: magru: change ospf metric on codfw link. [homer/public] - 10https://gerrit.wikimedia.org/r/1328564 (https://phabricator.wikimedia.org/T435543) [12:42:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:43:53] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:44:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:46:36] (03CR) 10TheDJ: wgCategoryTreeCategoryPageMode: use the new values [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319945 (https://phabricator.wikimedia.org/T152294) (owner: 10TheDJ) [12:47:31] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 24 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319945 (https://phabricator.wikimedia.org/T152294) (owner: 10TheDJ) [12:47:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:48:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:49:01] 10ops-eqiad, 06DC-Ops, 10Kafka-Infrastructure: Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12246582 (10brouberol) I just had a conversation with @cmooney where we went over the plan and constraints. Our network has 3 different types of failure zones: - racks - r... [12:49:18] !log fceratto@cumin1003 START - Cookbook sre.mysql.clone of db-test1002.eqiad.wmnet onto db1901.eqiad.wmnet [12:49:18] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.clone (exit_code=99) of db-test1002.eqiad.wmnet onto db1901.eqiad.wmnet [12:49:19] (03Merged) 10jenkins-bot: wikikube production: Update to cert-manager 1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328560 (https://phabricator.wikimedia.org/T427402) (owner: 10Jelto) [12:49:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:50:05] (03CR) 10Cathal Mooney: [C:03+2] magru: change ospf metric on codfw link. [homer/public] - 10https://gerrit.wikimedia.org/r/1328564 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [12:51:36] (03Merged) 10jenkins-bot: magru: change ospf metric on codfw link. [homer/public] - 10https://gerrit.wikimedia.org/r/1328564 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [12:52:25] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:52:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:53:25] 10ops-eqiad, 06DC-Ops, 10Kafka-Infrastructure: Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12246601 (10cmooney) Overall the plan makes sense, thanks for taking the time to explain the current setup and why we want to change. >>! In T435775#12246324, @brouberol... [12:53:33] (03CR) 10Brouberol: provision the airflow-experiment-platform DNS records (032 comments) [dns] - 10https://gerrit.wikimedia.org/r/1328576 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:53:34] !log jelto@deploy1003 helmfile [codfw] START helmfile.d/admin 'sync'. [12:53:49] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [12:54:06] !log jelto@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'sync'. [12:54:08] !log update cert-manager to 1.19.6 on wikikube codfw - T427402 [12:54:13] (03CR) 10Lucas Werkmeister (WMDE): [C:03+1] arwikiquote: update wordmark [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328315 (https://phabricator.wikimedia.org/T435505) (owner: 10Anzx) [12:54:14] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:54:15] T427402: Update cert-manager to 1.19 - https://phabricator.wikimedia.org/T427402 [12:54:49] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:55:11] (03CR) 10Brouberol: "CI is failing because a dummy password should be added to https://gerrit.wikimedia.org/r/plugins/gitiles/labs/private/+/refs/heads/master/" [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:55:24] 10ops-eqiad, 06DC-Ops, 10Kafka-Infrastructure: Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12246608 (10cmooney) >>! In T435775#12246581, @brouberol wrote: > I just had a conversation with @cmooney where we went over the plan and constraints. Crossed wires I wro... [12:56:21] (03CR) 10Brouberol: [C:03+1] "Adding Moritz as reviewer as there's a new sudo rule being added" [puppet] - 10https://gerrit.wikimedia.org/r/1328582 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:56:26] (03PS1) 10WMDE-Fisch: Add feature flag to beta cluster and test wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328590 (https://phabricator.wikimedia.org/T431544) (owner: 10Mareike Heuer) [12:56:27] (03CR) 10WMDE-Fisch: "Just some linting issue could be fixed by composer autofix I guess :-)" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328590 (https://phabricator.wikimedia.org/T431544) (owner: 10Mareike Heuer) [12:57:20] (03CR) 10Atsuko: "Dropped sudo from this patch, adding sudo in the following patch." [puppet] - 10https://gerrit.wikimedia.org/r/1328582 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:57:25] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:57:33] (03PS2) 10Atsuko: Add analytics-experiment system user and groups [puppet] - 10https://gerrit.wikimedia.org/r/1328582 (https://phabricator.wikimedia.org/T416709) [12:58:39] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1901.eqiad.wmnet with reason: Cloning [12:59:01] 10ops-eqiad, 06DC-Ops, 10Kafka-Infrastructure: Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12246622 (10brouberol) [12:59:18] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1901.eqiad.wmnet with reason: Cloning [12:59:36] (03PS1) 10Atsuko: Grant sudo privileges for the analytics-experiment-users group [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: I seem to be stuck in Groundhog week. Sigh. Time for (yet another) UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1300). [13:00:05] MichaelG_WMF, sadiya_wmde28, and anzx: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:00:12] o/ [13:00:15] * MichaelG_WMF is here [13:00:16] o/ [13:00:19] I can deploy! [13:00:55] Lucas_WMDE: Thanks! My change should just be config cleanup of config that is not needed anymore. It should be a noop. I can test it [13:01:14] anzx: do you want to quickly change those IP “ranges” to 'IP' or should I just deploy that change as is? ^^ [13:01:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:01:29] (I just noticed that the “possible patch” in the task description also had it under 'IP' instead of 'range') [13:02:06] Lucas_WMDE: i think they provided it by mistake [13:02:40] what do you mean? [13:02:48] !log jelto@deploy1003 helmfile [eqiad] START helmfile.d/admin 'sync'. [13:03:15] !log jelto@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'sync'. [13:03:20] !log update cert-manager to 1.19.6 on wikikube eqiad - T427402 [13:03:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:03:29] T427402: Update cert-manager to 1.19 - https://phabricator.wikimedia.org/T427402 [13:03:52] (03PS2) 10Atsuko: provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1328576 (https://phabricator.wikimedia.org/T416709) [13:03:55] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:03:55] Im here [13:04:01] above they /32 range , in possible patch they provided single IP [13:04:10] (03CR) 10Atsuko: provision the airflow-experiment-platform DNS records (032 comments) [dns] - 10https://gerrit.wikimedia.org/r/1328576 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [13:04:24] (anyway, we can already start with the wordmark and the config cleanup) [13:04:30] anzx: yes, a /32 “range” is a single IP [13:04:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:04:41] * Lucas_WMDE waves at sadiya_wmde [13:04:47] anzx: at least in IPv4 ^^ [13:04:57] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:05:05] Lucas_WMDE: i will update it, you can go ahead with rest patch [13:05:10] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:05:11] alright, thanks! [13:05:20] (03CR) 10Brouberol: [C:03+1] Add analytics-experiment system user and groups [puppet] - 10https://gerrit.wikimedia.org/r/1328582 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [13:05:34] (03CR) 10Brouberol: [C:03+1] provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1328576 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [13:05:59] MichaelG_WMF: scap says that the wmf.15 branch “is a likely rollback target” and the config change depends on a commit that was only included in wmf.16 [13:06:06] how bad is it if the train gets rolled back and the config is already gone? [13:06:38] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12246644 (10VRiley-WMF) @Andrew Here it is the instructions https://www.supermicro.com/en/support/manuals... [13:06:39] would that just revert to showing the “Wikipedia is written by X users” area on the signup page? [13:06:40] (03PS5) 10Anzx: Lift IP cap for Mapudungun editathon on 2026-08-29 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) [13:07:20] Lucas_WMDE: not bad, but this is also not urgent. If we're not yet comfortable with wmf.16 being stable, then let's postpone this. We can just deploy it tomorrow without any harm done. [13:07:27] Lucas_WMDE: updated [13:07:33] eh, I think it’s pretty likely stable [13:07:40] (03CR) 10CI reject: [V:04-1] Lift IP cap for Mapudungun editathon on 2026-08-29 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) (owner: 10Anzx) [13:07:42] let’s go with it anyway [13:07:46] (03CR) 10TrainBranchBot: [C:03+2] "Approved by lucaswerkmeister-wmde@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325527 (https://phabricator.wikimedia.org/T433783) (owner: 10Michael Große) [13:07:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by lucaswerkmeister-wmde@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328315 (https://phabricator.wikimedia.org/T435505) (owner: 10Anzx) [13:07:49] 👍 [13:07:58] as long as the consequence isn’t catastrophic [13:08:26] anzx: hm, CI is failing, let me see [13:08:40] (03PS1) 10Atsuko: idp: provision the airflow_experiment_platform placeholder secret [labs/private] - 10https://gerrit.wikimedia.org/r/1328598 (https://phabricator.wikimedia.org/T416709) [13:08:43] (03Merged) 10jenkins-bot: Growth: Drop now unused config for benefits block [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325527 (https://phabricator.wikimedia.org/T433783) (owner: 10Michael Große) [13:08:49] (03Merged) 10jenkins-bot: arwikiquote: update wordmark [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328315 (https://phabricator.wikimedia.org/T435505) (owner: 10Anzx) [13:08:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:08:55] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:08:56] Yes, the only consequence I would expect is that a very small amount users would not see custom copy on Special:CreateAccount [13:08:58] oh good, the test doesn’t allow the 'IP' value to be an array [13:09:01] !log lucaswerkmeister-wmde@deploy1003 Started scap sync-world: Backport for [[gerrit:1325527|Growth: Drop now unused config for benefits block (T433783)]], [[gerrit:1328315|arwikiquote: update wordmark (T435505)]] [13:09:07] T433783: Remove the benefits block from Special:CreateAccount - https://phabricator.wikimedia.org/T433783 [13:09:08] T435505: Fix Arabic wikiquote wordmark - https://phabricator.wikimedia.org/T435505 [13:09:35] Lucas_WMDE: should i update it [13:09:44] I’m updating the test [13:09:47] unless you want to do it ^^ [13:09:57] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:09:58] (03CR) 10Atsuko: [V:03+2 C:03+2] idp: provision the airflow_experiment_platform placeholder secret [labs/private] - 10https://gerrit.wikimedia.org/r/1328598 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [13:10:21] (03PS2) 10Atsuko: idp: add the airflow_experiment_platform service [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) [13:10:27] (03CR) 10Atsuko: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [13:10:32] Lucas_WMDE: please go ahead and update test [13:11:01] !log lucaswerkmeister-wmde@deploy1003 lucaswerkmeister-wmde, anzx, migr: Backport for [[gerrit:1325527|Growth: Drop now unused config for benefits block (T433783)]], [[gerrit:1328315|arwikiquote: update wordmark (T435505)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:11:10] checking [13:12:04] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1901.eqiad.wmnet with reason: Cloning [13:12:05] (03PS6) 10Lucas Werkmeister (WMDE): Lift IP cap for Mapudungun editathon on 2026-08-29 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) (owner: 10Anzx) [13:12:06] Lucas_WMDE: arwikiquote wordmark looks good, ok to sync [13:12:09] thanks [13:12:15] MichaelG_WMF: testing okay from your side too? [13:12:28] * MichaelG_WMF is looking [13:12:32] !log drain Lumen transport circuit from eqiad<->codfw due to link flaps/errors (T435810) [13:12:37] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:12:37] T435810: Outbound errors on interface cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://phabricator.wikimedia.org/T435810 [13:12:49] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:13:05] Lucas_WMDE: testing ok on my side too 👍 [13:13:10] !log lucaswerkmeister-wmde@deploy1003 lucaswerkmeister-wmde, anzx, migr: Continuing with deployment [13:13:12] thanks! [13:13:19] ^^^ FWIW we are seeing flaps on the transport link from Lumen from eqiad to codfw - this was responsible for the DB ping failure page a short time ago [13:13:32] I am draining the link now and will follow up with the carrier [13:13:43] MichaelG_WMF: maybe you can take a quick look at the ThrottleTest change I added to https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1328310, so I’m not just deploying my own unreviewed code? :) [13:13:53] topranks: yikes, thanks [13:13:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:13:58] anything to worry about regarding deployments? [13:14:06] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [13:14:16] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) [13:14:27] Lucas_WMDE: no the errors were marginal enough, traffic won't be using that link now though so you should be ok [13:14:28] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [13:14:29] (03PS1) 10Brouberol: Add the postgresql-superset-metrics fixture password [labs/private] - 10https://gerrit.wikimedia.org/r/1328601 [13:14:33] ok, thanks [13:14:38] Lucas_WMDE: I'm looking but am not familiar with the code. Pls give me a second [13:14:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:14:39] (03CR) 10Brouberol: [C:03+2] Add the postgresql-superset-metrics fixture password [labs/private] - 10https://gerrit.wikimedia.org/r/1328601 (owner: 10Brouberol) [13:14:41] (03CR) 10Brouberol: [V:03+2 C:03+2] Add the postgresql-superset-metrics fixture password [labs/private] - 10https://gerrit.wikimedia.org/r/1328601 (owner: 10Brouberol) [13:14:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:15:33] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:17:31] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:17:35] !log lucaswerkmeister-wmde@deploy1003 Finished scap sync-world: Backport for [[gerrit:1325527|Growth: Drop now unused config for benefits block (T433783)]], [[gerrit:1328315|arwikiquote: update wordmark (T435505)]] (duration: 08m 33s) [13:17:43] T433783: Remove the benefits block from Special:CreateAccount - https://phabricator.wikimedia.org/T433783 [13:17:43] T435505: Fix Arabic wikiquote wordmark - https://phabricator.wikimedia.org/T435505 [13:17:47] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:18:03] (03PS1) 10Kevin Bazira: ml-services: update outlink isvc to resolve non-canonical wiki_id aliases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328604 (https://phabricator.wikimedia.org/T435586) [13:18:47] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:19:01] (03PS2) 10Mareike Heuer: Add feature flag to beta cluster and test wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328590 (https://phabricator.wikimedia.org/T431544) [13:19:04] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [13:19:46] (03CR) 10Michael Große: [C:03+1] "code review requested during deployment window:" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) (owner: 10Anzx) [13:20:02] Lucas_WMDE: I gave it +1, it makes sense to me [13:20:10] thanks! [13:20:26] then I think we can deploy that and sadiya_wmde’s change together, should be low risk [13:20:51] (03PS1) 10Brouberol: secrets: provision the pg-superset-metrics passowrd in the analytics HDFS home [puppet] - 10https://gerrit.wikimedia.org/r/1328607 (https://phabricator.wikimedia.org/T432104) [13:21:00] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update outlink isvc to resolve non-canonical wiki_id aliases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328604 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [13:21:04] (03CR) 10Lucas Werkmeister (WMDE): [C:03+1] Disable Wikidata Bridge on Catalan Wikipedia (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326848 (https://phabricator.wikimedia.org/T433713) (owner: 10Sadiya.mohammed13) [13:21:17] also I wrote ^ that comment like 30 minutes ago but apparently didn’t send it, oops [13:21:28] (03CR) 10TrainBranchBot: [C:03+2] "Approved by lucaswerkmeister-wmde@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) (owner: 10Anzx) [13:21:29] (03CR) 10TrainBranchBot: [C:03+2] "Approved by lucaswerkmeister-wmde@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326848 (https://phabricator.wikimedia.org/T433713) (owner: 10Sadiya.mohammed13) [13:21:39] (03PS2) 10Brouberol: secrets: provision the pg-superset-metrics password in the analytics HDFS home [puppet] - 10https://gerrit.wikimedia.org/r/1328607 (https://phabricator.wikimedia.org/T432104) [13:22:31] (03Merged) 10jenkins-bot: Lift IP cap for Mapudungun editathon on 2026-08-29 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328310 (https://phabricator.wikimedia.org/T435523) (owner: 10Anzx) [13:22:35] (03Merged) 10jenkins-bot: Disable Wikidata Bridge on Catalan Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326848 (https://phabricator.wikimedia.org/T433713) (owner: 10Sadiya.mohammed13) [13:22:48] !log lucaswerkmeister-wmde@deploy1003 Started scap sync-world: Backport for [[gerrit:1328310|Lift IP cap for Mapudungun editathon on 2026-08-29 (T435523)]], [[gerrit:1326848|Disable Wikidata Bridge on Catalan Wikipedia (T433713)]] [13:22:54] (03CR) 10Brouberol: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328607 (https://phabricator.wikimedia.org/T432104) (owner: 10Brouberol) [13:22:54] T435523: Lift IP cap for Mapudungun editathon on 2026-08-29 - https://phabricator.wikimedia.org/T435523 [13:22:55] T433713: Verify if removing the Wikidata Bridge code reduces page load - https://phabricator.wikimedia.org/T433713 [13:23:39] (03Merged) 10jenkins-bot: ml-services: update outlink isvc to resolve non-canonical wiki_id aliases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328604 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [13:23:49] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:23:54] (03CR) 10WMDE-Fisch: [C:03+1] Add feature flag to beta cluster and test wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328590 (https://phabricator.wikimedia.org/T431544) (owner: 10Mareike Heuer) [13:24:46] !log lucaswerkmeister-wmde@deploy1003 lucaswerkmeister-wmde, anzx, sadiyamohammed13: Backport for [[gerrit:1328310|Lift IP cap for Mapudungun editathon on 2026-08-29 (T435523)]], [[gerrit:1326848|Disable Wikidata Bridge on Catalan Wikipedia (T433713)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:25:22] anzx: anything to test? ^^ [13:25:25] 10ops-eqiad, 06SRE, 06DC-Ops: Outbound errors on interface cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://phabricator.wikimedia.org/T435810#12246730 (10cmooney) This link has been flapping up and down repeatedly, light coming/going on the interface {F99846500 width=600} Fir... [13:25:25] Lucas_WMDE: nothing to test on throttle [13:25:28] sadiya_wmde: please test on mwdebug :) [13:25:31] yup, fair enough [13:25:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:25:55] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12246732 (10AlexisJazz) >>! In T435743#12246220, @Aklapper wrote: > Has this been reported to archive.org? What does `server does not respond.` mean exactly? I have a hard time... [13:25:55] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and 208.80.154.216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:25:55] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [13:26:29] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:26:45] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:26:57] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [13:27:28] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport instability (Aug 2026). - https://phabricator.wikimedia.org/T435810#12246735 (10cmooney) [13:28:23] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12246738 (10MoritzMuehlenhoff) >>! In T435354#12242593, @bking wrote: > - Are you aware of any physical servers that could replace `krb1002`? I'll start... [13:30:27] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:30:41] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:31:07] (03PS3) 10Brouberol: secrets: provision the pg-superset-metrics password in the analytics HDFS home [puppet] - 10https://gerrit.wikimedia.org/r/1328607 (https://phabricator.wikimedia.org/T432104) [13:32:07] (03CR) 10Brouberol: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9301/co" [puppet] - 10https://gerrit.wikimedia.org/r/1328607 (https://phabricator.wikimedia.org/T432104) (owner: 10Brouberol) [13:33:31] Lucas_WMDE: please run `echo 'https://en.wikipedia.org/static/images/mobile/copyright/wikiquote-wordmark-ar.svg' | mwscript purgeList.php` to purge wordmark image, i think i am still seeing old wordmark [13:33:36] ah right, thanks [13:34:36] !log lucaswerkmeister-wmde@deploy1003:/srv/mediawiki-staging$ echo 'https://en.wikipedia.org/static/images/mobile/copyright/wikiquote-wordmark-ar.svg' | mwscript-k8s --comment=T435505 --attach -- purgeList enwiki [13:34:40] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:34:41] T435505: Fix Arabic wikiquote wordmark - https://phabricator.wikimedia.org/T435505 [13:34:42] anzx: done [13:35:03] (03PS3) 10Atsuko: idp: add the airflow_experiment_platform service [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) [13:35:16] !log upgrading apus codfw cluster to Reef 18.2.8 [13:35:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:35:37] Lucas_WMDE: thanks now new workmark is displaying correctly [13:35:40] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:35:42] yay [13:35:47] thanks for the reminder ^^ [13:36:03] sadiya_wmde: are you still testing? [13:36:54] yh [13:37:05] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [13:37:11] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [13:37:19] ok [13:37:20] !log kevinbazira@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [13:38:49] (03PS2) 10Ladsgroup: turnilo: Expose X-analytics thumb_generated in webrequest_sampled_live [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) [13:39:19] (03PS3) 10SomeRandomDeveloper: Profiler: Fix excimer component regex skipping the last frame in each stack [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328191 [13:39:19] (03CR) 10Brouberol: [C:03+1] idp: add the airflow_experiment_platform service [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [13:39:25] (03CR) 10Ladsgroup: "Thanks! I made https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1328608" [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) (owner: 10Ladsgroup) [13:40:17] (03CR) 10SomeRandomDeveloper: Profiler: Fix excimer component regex skipping the last frame in each stack (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328191 (owner: 10SomeRandomDeveloper) [13:40:40] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:41:49] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12246796 (10TheDJ) Maybe @cyberpower678 can help find the right contacts [13:41:52] done with test looks good [13:41:57] alright, thanks! [13:42:01] !log lucaswerkmeister-wmde@deploy1003 lucaswerkmeister-wmde, anzx, sadiyamohammed13: Continuing with deployment [13:43:16] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12246803 (10ssingh) We are looking at the internal traffic logs and checking if archive.org is matching some other rule, or one of their own rules (and if something changed the... [13:43:29] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:45:25] !log andrew@cumin2003 START - Cookbook sre.hosts.reimage for host cloudcephosd1041.eqiad.wmnet with OS bookworm [13:45:40] RESOLVED: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:46:16] !log lucaswerkmeister-wmde@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328310|Lift IP cap for Mapudungun editathon on 2026-08-29 (T435523)]], [[gerrit:1326848|Disable Wikidata Bridge on Catalan Wikipedia (T433713)]] (duration: 23m 28s) [13:46:23] T435523: Lift IP cap for Mapudungun editathon on 2026-08-29 - https://phabricator.wikimedia.org/T435523 [13:46:23] T433713: Verify if removing the Wikidata Bridge code reduces page load - https://phabricator.wikimedia.org/T433713 [13:46:25] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport instability (Aug 2026). - https://phabricator.wikimedia.org/T435810#12246823 (10cmooney) Ticket ID 35191415 opened with Lumen. [13:46:25] !log UTC afternoon backport+config window done [13:46:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:46:29] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:47:14] Lucas_WMDE: thanks for deploying [13:47:50] (03CR) 10A-pizzata: [C:03+1] secrets: provision the pg-superset-metrics password in the analytics HDFS home [puppet] - 10https://gerrit.wikimedia.org/r/1328607 (https://phabricator.wikimedia.org/T432104) (owner: 10Brouberol) [13:50:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [13:51:13] (03CR) 10Brouberol: [V:03+1 C:03+2] secrets: provision the pg-superset-metrics password in the analytics HDFS home [puppet] - 10https://gerrit.wikimedia.org/r/1328607 (https://phabricator.wikimedia.org/T432104) (owner: 10Brouberol) [13:51:25] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [13:51:41] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99) [13:51:55] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:53:35] (03PS1) 10Atsuko: Fix PCC for IDP [puppet] - 10https://gerrit.wikimedia.org/r/1328610 (https://phabricator.wikimedia.org/T435815) [13:53:50] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport instability (Aug 2026). - https://phabricator.wikimedia.org/T435810#12246845 (10cmooney) FWIW this led to a page for dropped ping to db2236, and we can also see the effects in the blackbox probes graphs: {F99849806 width=600} [13:53:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:54:33] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [13:54:39] (03PS1) 10Atsuko: Revert "idp: provision the airflow_experiment_platform placeholder secret" [labs/private] - 10https://gerrit.wikimedia.org/r/1328611 [13:54:48] (03CR) 10Atsuko: [V:03+2 C:03+2] Revert "idp: provision the airflow_experiment_platform placeholder secret" [labs/private] - 10https://gerrit.wikimedia.org/r/1328611 (owner: 10Atsuko) [13:54:54] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [13:54:57] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [13:55:19] (03PS1) 10Atsuko: Re-do "idp: provision the airflow_experiment_platform placeholder secret" [labs/private] - 10https://gerrit.wikimedia.org/r/1328612 [13:55:40] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [13:55:40] RESOLVED: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:55:50] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [13:56:08] (03PS2) 10Atsuko: Re-do "idp: provision the airflow_experiment_platform placeholder secret" [labs/private] - 10https://gerrit.wikimedia.org/r/1328612 [13:58:04] (03CR) 10Atsuko: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328610 (https://phabricator.wikimedia.org/T435815) (owner: 10Atsuko) [13:58:34] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499#12246855 (10elukey) p:05Triage→03High a:03elukey [13:59:12] (03PS1) 10Brouberol: secrets: add missing include [puppet] - 10https://gerrit.wikimedia.org/r/1328613 (https://phabricator.wikimedia.org/T432104) [13:59:35] PROBLEM - PyBal IPVS diff check on lvs1019 is CRITICAL: (CRITICAL: Mismatch between IPVS and PyBal https://wikitech.wikimedia.org/wiki/PyBal [14:03:16] (03CR) 10Brouberol: [C:03+2] secrets: add missing include [puppet] - 10https://gerrit.wikimedia.org/r/1328613 (https://phabricator.wikimedia.org/T432104) (owner: 10Brouberol) [14:03:59] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12246876 (10VRiley-WMF) a:05VRiley-WMF→03Andrew [14:04:35] RECOVERY - PyBal IPVS diff check on lvs1019 is OK: OK: no difference between hosts in IPVS/PyBal https://wikitech.wikimedia.org/wiki/PyBal [14:04:50] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818 (10bking) 03NEW [14:04:52] (03PS2) 10Atsuko: Fix PCC for IDP [puppet] - 10https://gerrit.wikimedia.org/r/1328610 (https://phabricator.wikimedia.org/T435815) [14:05:20] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12246892 (10bking) a:05bking→03VRiley-WMF [14:05:27] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12246895 (10VRiley-WMF) @bking thanks for this! Is there a specific name we'd like to use for this host? [14:05:38] (03CR) 10Atsuko: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328610 (https://phabricator.wikimedia.org/T435815) (owner: 10Atsuko) [14:07:32] !log andrew@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cloudcephosd1041.eqiad.wmnet with reason: host reimage [14:08:10] (03CR) 10Atsuko: [V:03+2 C:03+2] Re-do "idp: provision the airflow_experiment_platform placeholder secret" [labs/private] - 10https://gerrit.wikimedia.org/r/1328612 (owner: 10Atsuko) [14:11:55] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12246915 (10Fabfur) from a rapid look at our logs, we're not currently blocking archive.org bot and trying to save a couple of pages from different wiki (enwiki, itwiki) works... [14:13:58] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499#12246922 (10LSobanski) 05Open→03In progress [14:14:05] (03CR) 10CDanis: [V:03+1 C:03+1] "thanks for running PCC -- good to see the expected no-op" [puppet] - 10https://gerrit.wikimedia.org/r/1328610 (https://phabricator.wikimedia.org/T435815) (owner: 10Atsuko) [14:15:13] !log andrew@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudcephosd1041.eqiad.wmnet with reason: host reimage [14:17:22] (03CR) 10Atsuko: [C:03+2] Fix PCC for IDP [puppet] - 10https://gerrit.wikimedia.org/r/1328610 (https://phabricator.wikimedia.org/T435815) (owner: 10Atsuko) [14:19:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:20:02] (03PS3) 10Atsuko: Add analytics-experiment system user and groups [puppet] - 10https://gerrit.wikimedia.org/r/1328582 (https://phabricator.wikimedia.org/T416709) [14:20:02] (03PS2) 10Atsuko: Grant sudo privileges for the analytics-experiment-users group [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) [14:22:09] (03CR) 10Scott French: [C:03+1] rest-gateway: Upgrade to envoy 1.39.0-1 in production. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328549 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [14:26:15] (03PS7) 10Atsuko: idp: add the airflow_experiment_platform service [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) [14:26:21] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1902.eqiad.wmnet with reason: Cloning [14:26:25] (03CR) 10Atsuko: [C:03+2] idp: add the airflow_experiment_platform service [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [14:26:50] (03CR) 10Atsuko: [C:03+2] "self-CR +2 as there is no change from Patchset 3" [puppet] - 10https://gerrit.wikimedia.org/r/1328573 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [14:29:13] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:29:27] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1430) [14:30:50] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1903.eqiad.wmnet with reason: Cloning [14:31:06] 06SRE, 06serviceops-deprecated, 10TimedMediaHandler, 05MW-1.47-notes (1.47.0-wmf.16; 2026-08-18), 13Patch-For-Review: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419#12247026 (10TheDJ) [14:31:42] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1903.eqiad.wmnet with reason: Cloning [14:33:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [14:35:12] 06SRE, 06serviceops-deprecated, 10TimedMediaHandler, 05MW-1.47-notes (1.47.0-wmf.16; 2026-08-18), 13Patch-For-Review: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419#12247052 (10TheDJ) { T435744} also explains some other... [14:35:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [14:38:19] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on db1903.eqiad.wmnet with reason: Cloning [14:38:28] (03PS8) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [14:38:39] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Alerting: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801#12247067 (10tappof) The alerts come from Pint, a linter/validator for Prometheus. It just says that the expression defined in the TransitBGPDown rul... [14:38:40] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [14:39:31] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Alerting: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801#12247073 (10tappof) [14:39:37] jouncebot: nowandnext [14:39:37] For the next 0 hour(s) and 20 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1430) [14:39:37] In 0 hour(s) and 50 minute(s): Wikimedia Portals Update (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1530) [14:39:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:39:52] (03PS1) 10Ladsgroup: Fix keyframe interval regression from fpsmax refactor [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328626 (https://phabricator.wikimedia.org/T318419) [14:40:12] (03CR) 10Ladsgroup: [C:03+2] Fix keyframe interval regression from fpsmax refactor [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328626 (https://phabricator.wikimedia.org/T318419) (owner: 10Ladsgroup) [14:41:03] !log vriley@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host krb1003 [14:41:24] !log vriley@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host krb1003 [14:41:59] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [14:43:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:43:14] !log fceratto@cumin1003 START - Cookbook sre.mysql.update-replication [14:43:34] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0) [14:43:49] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247098 (10VRiley-WMF) rack C3 U29 cableID 3155 port 28 [14:44:06] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247100 (10VRiley-WMF) Using the name krb1003 [14:44:44] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247104 (10bking) Per IRC conversation, we're going with `krb1003` as the new hostname. I'll get a Puppet patch started for that. [14:45:53] !log vriley@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [krb1003] - vriley@cumin1003" [14:45:57] !log vriley@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [krb1003] - vriley@cumin1003" [14:45:57] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:48:10] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:48:23] 06SRE, 06Infrastructure-Foundations, 10Puppet-Infrastructure: Remove Puppet 5 CA cert from wmf-certificates cert bundle - https://phabricator.wikimedia.org/T415255#12247141 (10MoritzMuehlenhoff) 05Open→03Resolved a:03MoritzMuehlenhoff This is rolled out fleet-wide [14:49:18] (03PS5) 10Jgiannelos: prv: Enable parsoid rendering for 10 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328528 (https://phabricator.wikimedia.org/T435830) [14:51:15] (03PS1) 10Bking: Kerberos: Add emergency replacement host to site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1328634 (https://phabricator.wikimedia.org/T435818) [14:51:20] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack: wait_for_optimal() should ignore acked alerts - https://phabricator.wikimedia.org/T319277#12247196 (10LSobanski) 05In progress→03Resolved a:03SLyngshede-WMF [14:51:47] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack: wait_for_optimal() should ignore acked alerts - https://phabricator.wikimedia.org/T319277#12247202 (10LSobanski) Also related: https://gerrit.wikimedia.org/r/c/operations/software/spicerack/+/1140208 [14:53:10] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:54:49] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:55:49] (03CR) 10Brouberol: [C:03+1] Kerberos: Add emergency replacement host to site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1328634 (https://phabricator.wikimedia.org/T435818) (owner: 10Bking) [14:56:29] (03CR) 10Atsuko: [C:03+2] dse-k8s-eqiad: add the airflow-experiment-platform ns to the ceph tenant list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327552 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [14:56:29] (03Merged) 10jenkins-bot: Fix keyframe interval regression from fpsmax refactor [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328626 (https://phabricator.wikimedia.org/T318419) (owner: 10Ladsgroup) [14:58:10] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:58:31] (03PS1) 10Ssingh: conftool-data: switch urldownloader backends to trixie hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328635 (https://phabricator.wikimedia.org/T429175) [14:58:52] (03PS1) 10Ladsgroup: [hardening] thumbor: Flip from allow-unless-denied to vice versa [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328636 [14:58:55] (03CR) 10Aghirelli: [C:03+1] "LGTM" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [14:59:28] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply [14:59:32] !log andrew@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudcephosd1041.eqiad.wmnet with OS bookworm [15:00:20] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply [15:00:23] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1328626|Fix keyframe interval regression from fpsmax refactor (T318419 T435744)]] [15:00:30] T318419: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419 [15:00:31] T435744: Bitrate for transcoded videos has gone through the roof (40mbps 480p?) - https://phabricator.wikimedia.org/T435744 [15:00:42] (03PS2) 10Aghirelli: Remove mode from RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328215 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [15:01:34] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations, 13Patch-For-Review: Move Docker images under the /v2/releng prefix to S3 - https://phabricator.wikimedia.org/T432829#12247308 (10elukey) Also found out /v2/repos/releng: ` elukey@registry2004:~$ sudo journalctl -u docker-registry-swift.service | e... [15:02:30] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1328626|Fix keyframe interval regression from fpsmax refactor (T318419 T435744)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:02:36] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply [15:03:04] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply [15:03:10] RESOLVED: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:03:58] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [15:04:45] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:05:54] (03PS1) 10Bking: Kerberos: Acknowledge krb2002 as primary [puppet] - 10https://gerrit.wikimedia.org/r/1328637 (https://phabricator.wikimedia.org/T435354) [15:05:57] (03Merged) 10jenkins-bot: dse-k8s-eqiad: add the airflow-experiment-platform ns to the ceph tenant list [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327552 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [15:06:15] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328637 (https://phabricator.wikimedia.org/T435354) (owner: 10Bking) [15:06:24] (03CR) 10Hnowlan: [C:03+1] [hardening] thumbor: Flip from allow-unless-denied to vice versa [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328636 (owner: 10Ladsgroup) [15:06:45] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:07:13] (03CR) 10Bking: [C:03+2] Kerberos: Add emergency replacement host to site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1328634 (https://phabricator.wikimedia.org/T435818) (owner: 10Bking) [15:08:32] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328626|Fix keyframe interval regression from fpsmax refactor (T318419 T435744)]] (duration: 08m 08s) [15:08:38] T318419: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419 [15:08:38] T435744: Bitrate for transcoded videos has gone through the roof (40mbps 480p?) - https://phabricator.wikimedia.org/T435744 [15:09:27] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 24 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326410 (https://phabricator.wikimedia.org/T431372) (owner: 10Aaron Schulz) [15:09:39] (03CR) 10Muehlenhoff: Kerberos: Add emergency replacement host to site.pp (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328634 (https://phabricator.wikimedia.org/T435818) (owner: 10Bking) [15:09:56] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply [15:10:43] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply [15:10:53] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply [15:11:30] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply [15:13:45] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:14:45] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:14:57] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:16:11] 06SRE-OnFire, 06Collaboration-Services, 10Znuny: ticket.wikimedia.org should page when down - https://phabricator.wikimedia.org/T354479#12247450 (10LSobanski) Let's validate whether VRTS downtime results in an automated task creation. I consider this to be sufficient at this time. [15:18:11] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host krb1003.eqiad.wmnet with OS bookworm [15:18:24] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247476 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by bking@cumin2003 for host kr... [15:18:25] FIRING: [4x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:20:08] (03CR) 10Mmartorana: [C:03+2] [hardening] thumbor: Flip from allow-unless-denied to vice versa [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328636 (owner: 10Ladsgroup) [15:21:06] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247513 (10bking) Thanks to @VRiley-WMF for standing up the new host so quickly! I'm reimaging it now and wat... [15:21:49] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:22:35] (03CR) 10SBassett: [hardening] thumbor: Flip from allow-unless-denied to vice versa (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328636 (owner: 10Ladsgroup) [15:22:36] (03Merged) 10jenkins-bot: [hardening] thumbor: Flip from allow-unless-denied to vice versa [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328636 (owner: 10Ladsgroup) [15:23:25] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:23:42] !log bking@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host krb1003.eqiad.wmnet with OS bookworm [15:23:45] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:24:02] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247520 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by bking@cumin2003 for host krb100... [15:24:16] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host krb1003.eqiad.wmnet with OS bookworm [15:24:29] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247522 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by bking@cumin2003 for host kr... [15:25:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:26:53] (03PS2) 10Arnaudb: gerrit: allow puppetservers to reach the envoy tlsproxy [puppet] - 10https://gerrit.wikimedia.org/r/1328152 (https://phabricator.wikimedia.org/T420184) [15:28:30] (03PS9) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [15:28:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [15:29:43] (03CR) 10Dreamy Jazz: [C:03+1] ModelToRun: Add getContentPolicyName() before getModelName() changes [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328280 (https://phabricator.wikimedia.org/T432848) (owner: 10Dreamy Jazz) [15:29:45] !log bking@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host krb1003.eqiad.wmnet with OS bookworm [15:30:03] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247574 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by bking@cumin2003 for host krb100... [15:30:05] jan_drewniak: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for Wikimedia Portals Update. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1530). [15:31:51] (03PS1) 10Majavah: P:wmcs::novaproxy: Manage X-Forwarded-For header in HAProxy logic [puppet] - 10https://gerrit.wikimedia.org/r/1328642 (https://phabricator.wikimedia.org/T429930) [15:31:53] (03PS1) 10Majavah: P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) [15:31:56] (03PS1) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [15:32:53] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [15:33:07] (03PS1) 10Bking: Kerberos: use correct insetup role [puppet] - 10https://gerrit.wikimedia.org/r/1328645 (https://phabricator.wikimedia.org/T435818) [15:33:18] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [15:33:52] (03PS2) 10Majavah: P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) [15:33:55] (03PS2) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [15:33:59] (03CR) 10Bking: [C:03+2] Kerberos: use correct insetup role [puppet] - 10https://gerrit.wikimedia.org/r/1328645 (https://phabricator.wikimedia.org/T435818) (owner: 10Bking) [15:34:07] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1328645 (https://phabricator.wikimedia.org/T435818) (owner: 10Bking) [15:34:23] !log mvernon@cumin1003 START - Cookbook sre.swift.remove-ghost-objects from container wikipedia-commons-local-public.c7 in eqiad [15:35:13] !log bking@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:35:28] (03PS2) 10Majavah: P:wmcs::novaproxy: Manage X-Forwarded-For header in HAProxy logic [puppet] - 10https://gerrit.wikimedia.org/r/1328642 (https://phabricator.wikimedia.org/T429930) [15:35:28] (03PS3) 10Majavah: P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) [15:35:28] (03PS3) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [15:36:55] !log mvernon@cumin1003 END (PASS) - Cookbook sre.swift.remove-ghost-objects (exit_code=0) from container wikipedia-commons-local-public.c7 in eqiad [15:38:16] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 24 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [15:38:47] (03CR) 10Ssingh: "Thanks for the review @adenisse@wikimedia.org. Please feel free to merge and roll out as required." [software/klaxon] - 10https://gerrit.wikimedia.org/r/1322809 (owner: 10Ssingh) [15:39:16] (03CR) 10Ladsgroup: [hardening] thumbor: Flip from allow-unless-denied to vice versa (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328636 (owner: 10Ladsgroup) [15:41:37] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [15:41:47] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [15:42:25] !log bking@cumin2003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:42:28] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12247635 (10RobH) [15:42:36] !log bking@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:43:35] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12247642 (10RobH) CS5147161 filed. To avoid expedite fees, work has to be filed for in business hours and over 24 hours in advance, so the 'start time' for this is Wednesday, August 26th @... [15:44:16] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12247644 (10RobH) [15:44:47] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [15:45:22] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [15:47:26] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mpostoronca@deploy1003 using scap backport" [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328280 (https://phabricator.wikimedia.org/T432848) (owner: 10Dreamy Jazz) [15:51:12] 10ops-eqiad, 06SRE, 10Ceph, 06cloud-services-team, and 3 others: cloudcephosd1044 boot issues - https://phabricator.wikimedia.org/T429267#12247674 (10Andrew) Just checking -- is this ticket in my court or yours? [15:51:37] if you see elevated thumbor 500 p.age on codfw, that's me [15:51:39] debugging [15:52:15] !log bking@cumin2003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:52:17] (03Merged) 10jenkins-bot: ModelToRun: Add getContentPolicyName() before getModelName() changes [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328280 (https://phabricator.wikimedia.org/T432848) (owner: 10Dreamy Jazz) [15:52:32] !log mpostoronca@deploy1003 Started scap sync-world: Backport for [[gerrit:1328280|ModelToRun: Add getContentPolicyName() before getModelName() changes (T432848)]] [15:52:32] !log bking@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:52:37] T432848: EventLogging: Record confidence score results - https://phabricator.wikimedia.org/T432848 [15:52:38] sigh, IM doesnt' allow anything basically :( [15:52:43] !log bking@cumin2003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:52:51] !log bking@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:53:15] !log bking@cumin2003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:53:25] (03PS1) 10Ladsgroup: Revert "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328654 [15:53:31] !log bking@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:53:32] (03CR) 10Ladsgroup: [C:03+2] Revert "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328654 (owner: 10Ladsgroup) [15:53:41] (03PS10) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [15:53:48] !log bking@cumin2003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['krb1003.eqiad.wmnet'] [15:54:16] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host krb1003.eqiad.wmnet with OS bookworm [15:54:30] !log mpostoronca@deploy1003 dreamyjazz, mpostoronca: Backport for [[gerrit:1328280|ModelToRun: Add getContentPolicyName() before getModelName() changes (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:54:32] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247694 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by bking@cumin2003 for host kr... [15:54:45] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [15:54:49] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [15:55:42] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [15:56:05] (03Merged) 10jenkins-bot: Revert "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328654 (owner: 10Ladsgroup) [15:56:10] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [15:57:25] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [15:57:30] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [15:58:25] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [15:58:35] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [15:59:05] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [15:59:10] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [15:59:50] This is really annoying ^ it's merged and it's pulled but change list is empty [15:59:53] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [16:00:06] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [16:00:26] Amir1: the revert needs to bump the chart forward :( [16:00:46] !log bking@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host krb1003.eqiad.wmnet with OS bookworm [16:00:48] ah, thanks [16:01:00] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247759 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by bking@cumin2003 for host krb100... [16:01:08] I mean the whole point of versioning is to be able to move back and forth [16:01:17] you can do that manually via helm [16:01:51] FIRING: [3x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from swift.discovery.wmnet in codfw #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [16:02:02] (03PS1) 10Ladsgroup: thumbor: Bump chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328657 [16:02:03] that's me [16:02:11] (03CR) 10Ladsgroup: [C:03+2] thumbor: Bump chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328657 (owner: 10Ladsgroup) [16:02:30] (03CR) 10Ladsgroup: thumbor: Bump chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328657 (owner: 10Ladsgroup) [16:02:45] (03PS2) 10Ladsgroup: thumbor: Bump chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328657 [16:02:56] (03CR) 10Ladsgroup: [C:03+2] thumbor: Bump chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328657 (owner: 10Ladsgroup) [16:03:02] (03CR) 10Hnowlan: [C:03+1] thumbor: Bump chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328657 (owner: 10Ladsgroup) [16:03:27] !log mpostoronca@deploy1003 dreamyjazz, mpostoronca: Continuing with deployment [16:04:34] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet [16:04:36] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [16:05:24] (03Merged) 10jenkins-bot: thumbor: Bump chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328657 (owner: 10Ladsgroup) [16:05:33] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [16:05:38] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [16:05:45] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [16:05:50] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [16:06:00] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [16:06:14] pulled it and it's still empty uggh [16:06:17] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [16:06:28] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [16:06:32] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [16:07:40] !log mpostoronca@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328280|ModelToRun: Add getContentPolicyName() before getModelName() changes (T432848)]] (duration: 15m 08s) [16:07:41] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [16:07:45] T432848: EventLogging: Record confidence score results - https://phabricator.wikimedia.org/T432848 [16:07:46] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [16:08:18] (03CR) 10JMeybohm: [C:04-1] team-sre: add serviceops tag (032 comments) [alerts] - 10https://gerrit.wikimedia.org/r/1327076 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [16:09:26] (03CR) 10Blake: team-sre: add serviceops tag [alerts] - 10https://gerrit.wikimedia.org/r/1327076 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [16:09:54] Amir1: we can revert using helm [16:09:59] although I see the diff now [16:10:02] on it already [16:10:06] fceratto@cumin1003 makevm (PID 3037297) is awaiting input [16:10:07] so you can just apply [16:10:12] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [16:10:33] Thanks. Done now [16:10:40] https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency I was getting to this [16:10:46] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [16:11:04] (03PS1) 10Urbanecm: [Growth] eswiki: Increase mentorship to 85% [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328660 (https://phabricator.wikimedia.org/T394867) [16:11:22] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1328635 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [16:14:04] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12247875 (10bking) Reimages are failing; the host doesn't seem to PXE boot properly. I'm updating iDRAC, NIC, a... [16:14:33] TIL about diff too, so no need to apply [16:14:41] 500 rate is right down now <3 [16:14:47] (and spam SAL) [16:15:12] I debug this, sorry for it [16:15:51] (03PS1) 10Kevin Bazira: ml-services: update outlink isvc to encoded length when batching title queries [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328664 (https://phabricator.wikimedia.org/T435586) [16:15:51] (03PS3) 10Majavah: P:wmcs::novaproxy: Manage X-Forwarded-For header in HAProxy logic [puppet] - 10https://gerrit.wikimedia.org/r/1328642 (https://phabricator.wikimedia.org/T429930) [16:15:51] (03PS4) 10Majavah: P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) [16:15:52] (03PS4) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [16:16:51] RESOLVED: [3x] ATSBackendErrorsHigh: ATS: elevated 5xx errors from swift.discovery.wmnet in codfw #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [16:18:35] (03PS4) 10Majavah: P:wmcs::novaproxy: Manage X-Forwarded-For header in HAProxy logic [puppet] - 10https://gerrit.wikimedia.org/r/1328642 (https://phabricator.wikimedia.org/T429930) [16:18:35] (03PS5) 10Majavah: P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) [16:18:35] (03PS5) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [16:19:05] (03PS1) 10Elukey: docker_registry: add /v2/repos/releng to the Releng's location [puppet] - 10https://gerrit.wikimedia.org/r/1328666 (https://phabricator.wikimedia.org/T432829) [16:19:05] (03PS2) 10Kevin Bazira: ml-services: update outlink isvc to measure encoded length when batching title queries [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328664 (https://phabricator.wikimedia.org/T435586) [16:20:16] (03PS1) 10Kamila Součková: Add title-case mapping to support migration to PHP 8.5 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) [16:21:28] (03PS6) 10Majavah: P:wmcs::novaproxy: Maintain map file of backend hosts [puppet] - 10https://gerrit.wikimedia.org/r/1328643 (https://phabricator.wikimedia.org/T429930) [16:21:28] (03PS6) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [16:21:58] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update outlink isvc to measure encoded length when batching title queries [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328664 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [16:24:28] (03Merged) 10jenkins-bot: ml-services: update outlink isvc to measure encoded length when batching title queries [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328664 (https://phabricator.wikimedia.org/T435586) (owner: 10Kevin Bazira) [16:26:15] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [16:31:00] (03PS1) 10Jdlrobson: Support change tags for anonymous users [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328376 [16:34:16] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12248021 (10Andrew) ok -- supermicro servers have an option to optimize powersaving mode for virtual loads... [16:35:28] (03CR) 10Kamila Součková: "This is what came out of the maintenance script. Would either of you be able to double-check that these make sense?" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [16:37:37] (03CR) 10Kamila Součková: "well that looks great in gerrit 😂 looks good on my laptop, I promise 😄" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [16:39:30] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [16:39:39] !log kevinbazira@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' . [16:43:42] (03PS7) 10Majavah: P:wmcs::novaproxy: Use HAProxy mapping to proxy traffic [puppet] - 10https://gerrit.wikimedia.org/r/1328644 (https://phabricator.wikimedia.org/T429930) [16:45:01] (03CR) 10RLazarus: [C:03+1] rest-gateway: Upgrade to envoy 1.39.0-1 in production. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328549 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [16:45:16] (03CR) 10Andrew Bogott: "resolved" [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [16:49:55] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" [16:50:00] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2901.codfw.wmnet - fceratto@cumin1003" [16:50:00] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:50:00] !log fceratto@cumin1003 START - Cookbook sre.dns.wipe-cache db2901.codfw.wmnet on all recursors [16:50:04] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2901.codfw.wmnet on all recursors [16:50:38] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" [16:50:43] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2901.codfw.wmnet - fceratto@cumin1003" [16:50:54] !log fceratto@cumin1003 START - Cookbook sre.hosts.reimage for host db2901.codfw.wmnet with OS trixie [16:50:56] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 25 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [16:51:27] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 25 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328239 (https://phabricator.wikimedia.org/T435258) (owner: 10Bernard Wang) [16:54:03] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host krb1003.eqiad.wmnet with OS bookworm [16:54:12] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12248148 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by bking@cumin2003 for host krb1003.eqiad.wmnet with... [16:56:10] (03PS1) 10DLynch: Update VE core submodule to master (5e8709982) [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328674 (https://phabricator.wikimedia.org/T414328) [16:57:26] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 24 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328674 (https://phabricator.wikimedia.org/T414328) (owner: 10DLynch) [16:59:06] (03CR) 10Hnowlan: [C:03+1] prometheus/jobunavailable: double the alert firing time [alerts] - 10https://gerrit.wikimedia.org/r/1328587 (owner: 10Tiziano Fogli) [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1700) [17:00:05] bd808: A patch you scheduled for MediaWiki infrastructure (UTC late) is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [17:00:05] ryankemper: #bothumor I � Unicode. All rise for Wikidata Query Service weekly deploy deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T1700). [17:04:39] (03PS1) 10JMeybohm: coredns: Double memory limit in codfw and eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328675 (https://phabricator.wikimedia.org/T428573) [17:06:15] (03PS6) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [17:06:57] o/ [17:07:35] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on krb1003.eqiad.wmnet with reason: host reimage [17:07:36] I did not get a chance to log it in the calendar, but I have a scap configuration change to deploy in this infra window [17:07:59] should not conflict with b.d808's planned changes for shellbox [17:08:02] !log fceratto@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on db2901.codfw.wmnet with reason: host reimage [17:09:49] (03CR) 10Scott French: [C:03+2] kubernetes: Enable testwiki httpbb test suite for pretrain [puppet] - 10https://gerrit.wikimedia.org/r/1327168 (https://phabricator.wikimedia.org/T428972) (owner: 10Scott French) [17:10:29] swfrench-wmf: +1. I'm about to start on the patch for mine, but I think we should be out of each other's way. [17:11:27] bd808: sounds good. you'll (eventually) see a noop scap run happen in the background to test things out :) [17:11:40] (03PS1) 10JMeybohm: ml-serve: Bump coredns resources [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328677 (https://phabricator.wikimedia.org/T428573) [17:12:15] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on krb1003.eqiad.wmnet with reason: host reimage [17:14:40] (03PS1) 10BryanDavis: shellbox: Bump to 2026-08-20-124436 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328679 (https://phabricator.wikimedia.org/T421653) [17:15:21] (03CR) 10JMeybohm: [C:03+2] coredns: Double memory limit in codfw and eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328675 (https://phabricator.wikimedia.org/T428573) (owner: 10JMeybohm) [17:15:50] 06SRE, 10LDAP-Access-Requests: Grant Access to wmf, for WMF staff/contractors nda group for gsduser - https://phabricator.wikimedia.org/T435852 (10GSduser) 03NEW [17:16:04] !log jayme@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [17:16:13] !log jayme@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [17:16:13] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2901.codfw.wmnet with reason: host reimage [17:16:53] !log jayme@deploy1003 helmfile [codfw] START helmfile.d/admin 'sync'. [17:18:58] !log jayme@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'sync'. [17:19:21] !log jayme@deploy1003 helmfile [eqiad] START helmfile.d/admin 'sync'. [17:20:04] (03CR) 10BryanDavis: [C:03+2] shellbox: Bump to 2026-08-20-124436 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328679 (https://phabricator.wikimedia.org/T421653) (owner: 10BryanDavis) [17:20:09] !log swfrench@deploy1003 Started scap sync-world: Noop deployment to validate pretrain httpbb check configuration - T428972 [17:20:13] T428972: Configure additional httpbb checks to perform during a deployment - https://phabricator.wikimedia.org/T428972 [17:20:18] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1328637 (https://phabricator.wikimedia.org/T435354) (owner: 10Bking) [17:21:25] !log jayme@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'sync'. [17:21:32] (03CR) 10JMeybohm: [C:03+2] ml-serve: Bump coredns resources [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328677 (https://phabricator.wikimedia.org/T428573) (owner: 10JMeybohm) [17:22:40] !log swfrench@deploy1003 Finished scap sync-world: Noop deployment to validate pretrain httpbb check configuration - T428972 (duration: 02m 30s) [17:22:43] (03Merged) 10jenkins-bot: shellbox: Bump to 2026-08-20-124436 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328679 (https://phabricator.wikimedia.org/T421653) (owner: 10BryanDavis) [17:23:55] !log jayme@deploy1003 helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'. [17:25:20] !log jayme@deploy1003 helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'. [17:25:28] !log jayme@deploy1003 helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'. [17:26:55] !log jayme@deploy1003 helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'. [17:27:14] (03CR) 10BCornwall: [C:03+1] provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1328576 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [17:27:30] 06SRE, 10LDAP-Access-Requests: Grant Access to wmf, for WMF staff/contractors nda group for gsduser - https://phabricator.wikimedia.org/T435852#12248279 (10Aklapper) [17:28:32] 06SRE, 10LDAP-Access-Requests: Grant Access to wmf, for WMF staff/contractors nda group for gsduser - https://phabricator.wikimedia.org/T435852#12248284 (10Aklapper) @GSduser: Hi and welcome! Can you please also [link your LDAP account to your Phabricator account](https://phabricator.wikimedia.org/settings/pan... [17:28:53] (03CR) 10BCornwall: [C:03+1] wikimedia.org: add etherpad-next and lower TTL [dns] - 10https://gerrit.wikimedia.org/r/1328350 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [17:28:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:29:11] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - bking@cumin2003" [17:30:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [17:31:49] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [17:32:15] bking@cumin2003 reimage (PID 301384) is awaiting input [17:32:26] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - bking@cumin2003" [17:32:27] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host krb1003.eqiad.wmnet with OS bookworm [17:32:35] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12248310 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by bking@cumin2003 for host krb1003.eqiad.wmnet with OS... [17:34:38] (03CR) 10Bking: [C:03+2] Kerberos: Acknowledge krb2002 as primary [puppet] - 10https://gerrit.wikimedia.org/r/1328637 (https://phabricator.wikimedia.org/T435354) (owner: 10Bking) [17:36:45] (03CR) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [17:37:17] cmooney@cumin1003 netbox (PID 3101696) is awaiting input [17:38:43] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entry for telxius circuit IP in magru - cmooney@cumin1003" [17:38:47] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entry for telxius circuit IP in magru - cmooney@cumin1003" [17:38:47] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:40:44] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12248324 (10bking) [17:42:23] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/shellbox: apply [17:42:29] !log disable-puppet on A:cp for ATS Lua change - T427666 [17:42:32] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:42:33] T427666: Route testwiki traffic to the Pretrain MVP environment - https://phabricator.wikimedia.org/T427666 [17:43:03] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox: apply [17:43:33] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Bring krb1003 into service - https://phabricator.wikimedia.org/T435854 (10bking) 03NEW [17:43:50] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Bring krb1003 into service - https://phabricator.wikimedia.org/T435854#12248344 (10bking) 05Open→03In progress p:05Triage→03High [17:45:06] (03CR) 10Scott French: [C:03+2] trafficserver: Support testwiki pretrain routing in XWD [puppet] - 10https://gerrit.wikimedia.org/r/1327621 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [17:45:06] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:45:08] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Replace failed KRB master with another physical host - https://phabricator.wikimedia.org/T435818#12248348 (10bking) 05Open→03Resolved `krb1003` is reachable via SSH, so the AC of this ticket is fulfilled. Work to bring thi... [17:45:37] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Metrics: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855 (10cmooney) 03NEW p:05Triage→03High [17:45:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [17:45:45] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/shellbox: apply [17:45:54] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox: apply [17:45:56] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2901.codfw.wmnet with OS trixie [17:45:56] !log fceratto@cumin1003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db2901.codfw.wmnet [17:46:45] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Metrics: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855#12248368 (10cmooney) [17:48:03] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/shellbox: apply [17:48:10] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox: apply [17:48:17] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-constraints: apply [17:48:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:48:55] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply [17:49:02] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-media: apply [17:49:16] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-media: apply [17:49:23] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply [17:49:46] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [17:49:53] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-timeline: apply [17:50:45] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply [17:50:51] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-video: apply [17:51:47] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-video: apply [17:51:57] (03PS7) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [17:51:57] (03PS1) 10Cwhite: profile: copy certificate chain file to more accessible place [puppet] - 10https://gerrit.wikimedia.org/r/1328686 (https://phabricator.wikimedia.org/T350516) [17:54:54] (03PS1) 10Bking: Kerberos: Introduce krb1003 as net-new replica [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) [17:54:57] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [17:55:30] (03PS2) 10Bking: Kerberos: Introduce krb1003 as net-new replica [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) [17:55:31] !log start rolling run-puppet-agent on A:cp after validation on cp4041 - T427666 [17:55:34] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:55:35] T427666: Route testwiki traffic to the Pretrain MVP environment - https://phabricator.wikimedia.org/T427666 [17:57:48] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [17:59:15] !log bd808@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox: apply [18:00:05] !log bd808@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox: apply [18:00:11] !log bd808@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply [18:01:14] !log bd808@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply [18:01:20] !log bd808@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-media: apply [18:01:42] !log bd808@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply [18:01:48] !log bd808@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply [18:02:18] !log bd808@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [18:02:25] !log bd808@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply [18:03:09] !log bd808@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply [18:03:15] !log bd808@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-video: apply [18:04:54] !log bd808@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply [18:06:20] PROBLEM - Check unit status of replicate-krb-database on krb2002 is CRITICAL: CRITICAL: Status of the systemd unit replicate-krb-database https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:08:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:09:58] (03CR) 10Cwhite: [C:03+2] opensearch: configure curator based on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328244 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [18:10:27] !log bd808@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox: apply [18:11:18] !log bd808@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox: apply [18:11:24] !log bd808@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply [18:12:20] !log bd808@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply [18:12:26] !log bd808@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-media: apply [18:13:23] !log bd808@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply [18:13:29] !log bd808@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply [18:13:53] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:13:58] !log bd808@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [18:14:04] !log bd808@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply [18:14:51] !log bd808@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply [18:14:57] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:14:58] !log bd808@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-video: apply [18:16:39] !log bd808@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply [18:31:43] (03CR) 10Bking: "PCC isn't working because `krb1003` is too new; I tried manually updating per https://wikitech.wikimedia.org/wiki/Help:Puppet-compiler#Man" [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [18:32:36] (03PS3) 10Bking: Kerberos: Introduce krb1003 as net-new replica [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) [18:36:59] (03CR) 10Muehlenhoff: [C:03+1] "Looks good. But before you merge, please ensure that krb2002 is now a proper kadminserver. One way to do that is to change your current Ke" [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [18:43:57] jouncebot: nowandnext [18:43:58] No deployments scheduled for the next 1 hour(s) and 16 minute(s) [18:43:58] In 1 hour(s) and 16 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T2000) [18:44:20] (03CR) 10TrainBranchBot: [C:03+2] "Approved by urbanecm@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328660 (https://phabricator.wikimedia.org/T394867) (owner: 10Urbanecm) [18:44:54] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12248577 (10BLiviero-WMF) can you say more about what "optimizing powersaving mode" for virtual instances... [18:45:17] (03Merged) 10jenkins-bot: [Growth] eswiki: Increase mentorship to 85% [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328660 (https://phabricator.wikimedia.org/T394867) (owner: 10Urbanecm) [18:45:33] !log urbanecm@deploy1003 Started scap sync-world: Backport for [[gerrit:1328660|[Growth] eswiki: Increase mentorship to 85% (T394867)]] [18:45:38] T394867: Incrementally increase mentorship at Spanish Wikipedia: increase to 85% of new accounts - https://phabricator.wikimedia.org/T394867 [18:47:33] !log urbanecm@deploy1003 urbanecm: Backport for [[gerrit:1328660|[Growth] eswiki: Increase mentorship to 85% (T394867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:47:58] (03CR) 10Bking: "Thanks for the suggestion, I will try that now." [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [18:48:00] !log urbanecm@deploy1003 urbanecm: Continuing with deployment [18:50:37] (03CR) 10JHathaway: [C:03+1] Kerberos: Introduce krb1003 as net-new replica [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [18:54:21] !log urbanecm@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328660|[Growth] eswiki: Increase mentorship to 85% (T394867)]] (duration: 08m 47s) [18:54:26] T394867: Incrementally increase mentorship at Spanish Wikipedia: increase to 85% of new accounts - https://phabricator.wikimedia.org/T394867 [18:54:49] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [18:57:17] (03PS1) 10Aleksandar Mastilovic: Disable Presto spilling [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) [18:57:51] (03CR) 10CI reject: [V:04-1] Disable Presto spilling [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [18:58:05] (03CR) 10Bking: [C:03+2] "Confirmed that `krb2002` is master via @mmuhlenhoff@wikimedia.org 's suggestion of changing my kerberos pw. Merging..." [puppet] - 10https://gerrit.wikimedia.org/r/1328688 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [18:58:08] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport instability (Aug 2026). - https://phabricator.wikimedia.org/T435810#12248613 (10cmooney) Lumen came back to say they found a problem and are working on it. ` On Mon, 24 Aug 2026 at 19:50, wrote: Trouble shooting the ci... [18:58:57] (03PS2) 10Aleksandar Mastilovic: Disable Presto spilling [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) [18:59:00] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [18:59:45] (03CR) 10Xcollazo: [C:03+1] Disable Presto spilling (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [19:03:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:05:30] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12248640 (10jhathaway) I booted the box into SystemRescue running 12.02 running Linux kernel 6.12.43-1-lts, and I can't detect any performance issues, I... [19:05:54] (03CR) 10Scott French: [C:03+1] docker_registry: add /v2/repos/releng to the Releng's location [puppet] - 10https://gerrit.wikimedia.org/r/1328666 (https://phabricator.wikimedia.org/T432829) (owner: 10Elukey) [19:06:08] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [19:08:00] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport ~~instability~~ outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12248643 (10cmooney) [19:08:09] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12248644 (10cmooney) [19:08:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:09:51] !log bking@krb2002 addprinc -randkey host/krb1003.eqiad.wmnet@WIKIMEDIA T435854 [19:09:55] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:09:56] T435854: Bring krb1003 into service - https://phabricator.wikimedia.org/T435854 [19:10:23] !log bking@krb2002 ktadd -k /tmp/krb1003.keytab host/krb1003.eqiad.wmnet@WIKIMEDIA T435854 [19:10:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:13:53] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:14:57] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:23:53] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:24:57] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:29:57] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:29:57] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [19:30:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [19:32:48] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12248722 (10cmooney) ` On Mon, 24 Aug 2026 at 20:18, wrote: The Lumen field tech is now on site and we are working with them remotely, more status will... [19:34:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:34:33] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:34:49] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:36:35] (03CR) 10Bking: Disable Presto spilling (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [19:36:49] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:37:33] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [19:39:10] RESOLVED: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:39:40] (03CR) 10TrainBranchBot: [C:03+2] "Approved by musikanimal@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328386 (https://phabricator.wikimedia.org/T127682) (owner: 10MusikAnimal) [19:40:35] (03Merged) 10jenkins-bot: CodeMirror: enable JSONC mode [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328386 (https://phabricator.wikimedia.org/T127682) (owner: 10MusikAnimal) [19:40:48] !log musikanimal@deploy1003 Started scap sync-world: Backport for [[gerrit:1328386|CodeMirror: enable JSONC mode (T127682)]] [19:40:50] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Degraded RAID on maps-test2001 - https://phabricator.wikimedia.org/T435311#12248802 (10Jhancock.wm) [19:40:53] T127682: Make code editor understand JSON with comments and trailing commas - https://phabricator.wikimedia.org/T127682 [19:42:37] (03PS1) 10Bking: Kerberos: Remove broken host, add replacement host [puppet] - 10https://gerrit.wikimedia.org/r/1328700 (https://phabricator.wikimedia.org/T435854) [19:42:48] !log musikanimal@deploy1003 musikanimal: Backport for [[gerrit:1328386|CodeMirror: enable JSONC mode (T127682)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [19:43:03] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328700 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [19:43:34] !log musikanimal@deploy1003 musikanimal: Continuing with deployment [19:44:25] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:44:40] RESOLVED: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:44:57] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [19:45:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [19:47:52] !log musikanimal@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328386|CodeMirror: enable JSONC mode (T127682)]] (duration: 07m 04s) [19:47:57] T127682: Make code editor understand JSON with comments and trailing commas - https://phabricator.wikimedia.org/T127682 [19:56:21] RECOVERY - Check unit status of replicate-krb-database on krb2002 is OK: OK: Status of the systemd unit replicate-krb-database https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:58:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:00:04] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: That opportune time for a UTC late backport window deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T2000). [20:00:04] thedj, AaronSchulz, and kemayo: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:32] o/ (I can deploy for myself) [20:01:30] (03PS2) 10Cwhite: profile: copy certificate chain file to more accessible place [puppet] - 10https://gerrit.wikimedia.org/r/1328686 (https://phabricator.wikimedia.org/T350516) [20:01:41] (03PS8) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [20:02:18] Nobody else having shown up so far, I'm going to go ahead and get mine started. [20:02:41] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328674 (https://phabricator.wikimedia.org/T414328) (owner: 10DLynch) [20:03:58] (03PS3) 10Cwhite: profile: copy certificate chain file to more accessible place [puppet] - 10https://gerrit.wikimedia.org/r/1328686 (https://phabricator.wikimedia.org/T350516) [20:03:58] (03PS9) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) [20:05:01] (03Merged) 10jenkins-bot: Update VE core submodule to master (5e8709982) [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1328674 (https://phabricator.wikimedia.org/T414328) (owner: 10DLynch) [20:05:07] (03CR) 10JHathaway: [C:03+1] Kerberos: Remove broken host, add replacement host [puppet] - 10https://gerrit.wikimedia.org/r/1328700 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [20:05:17] 06SRE, 10Beta-Cluster-Infrastructure, 06Traffic: Rename deployment-cache-(text|upload)0x to deployment-cp-(text|upload)0x - https://phabricator.wikimedia.org/T280393#12248905 (10bd808) [20:05:17] !log kemayo@deploy1003 Started scap sync-world: Backport for [[gerrit:1328674|Update VE core submodule to master (5e8709982) (T414328 T433533 T435688)]] [20:05:28] T414328: Cosmetic UX changes to Details Dialog - https://phabricator.wikimedia.org/T414328 [20:05:28] T433533: Empty document placeholder is fragile - https://phabricator.wikimedia.org/T433533 [20:05:29] T435688: DiscussionTools and 2017 wikitext editor interrupts desktop Chinese IME during composition on the first try - https://phabricator.wikimedia.org/T435688 [20:06:03] 06SRE, 10Beta-Cluster-Infrastructure, 06Traffic: Rename deployment-cache-(text|upload)0x to deployment-cp-(text|upload)0x - https://phabricator.wikimedia.org/T280393#12248929 (10bd808) [20:07:20] !log kemayo@deploy1003 kemayo: Backport for [[gerrit:1328674|Update VE core submodule to master (5e8709982) (T414328 T433533 T435688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:07:21] PROBLEM - Check unit status of replicate-krb-database on krb2002 is CRITICAL: CRITICAL: Status of the systemd unit replicate-krb-database https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [20:08:40] (03CR) 10Cwhite: [C:03+2] profile: copy certificate chain file to more accessible place [puppet] - 10https://gerrit.wikimedia.org/r/1328686 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:08:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:09:45] * AaronSchulz waits [20:09:55] !log kemayo@deploy1003 kemayo: Continuing with deployment [20:10:04] (03CR) 10ArielGlenn: "This looks ok to my eyeballs, having installed a font that covers one of the code blocks I had missing. Would you want to do a spot check" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328667 (https://phabricator.wikimedia.org/T432985) (owner: 10Kamila Součková) [20:11:18] (03PS1) 10Arlolra: Set a default for wgUseParsoidLinksUpdate in prod [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328707 (https://phabricator.wikimedia.org/T432760) [20:12:06] 06SRE, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): install1005 running out of disk due to squid log volume from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12248946 (10Snwachukwu) Thanks for the suggestion, @RKemper. @Ottomata had also suggested u... [20:15:12] (03PS1) 10Bking: Kerberos: Don't bail out on a single replica failure [puppet] - 10https://gerrit.wikimedia.org/r/1328709 (https://phabricator.wikimedia.org/T435854) [20:15:28] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 24 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328707 (https://phabricator.wikimedia.org/T432760) (owner: 10Arlolra) [20:16:04] !log kemayo@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328674|Update VE core submodule to master (5e8709982) (T414328 T433533 T435688)]] (duration: 10m 47s) [20:16:12] T414328: Cosmetic UX changes to Details Dialog - https://phabricator.wikimedia.org/T414328 [20:16:13] T433533: Empty document placeholder is fragile - https://phabricator.wikimedia.org/T433533 [20:16:13] T435688: DiscussionTools and 2017 wikitext editor interrupts desktop Chinese IME during composition on the first try - https://phabricator.wikimedia.org/T435688 [20:16:23] AaronSchulz: you're up [20:18:44] ok [20:19:07] (03CR) 10JHathaway: [C:03+1] Kerberos: Don't bail out on a single replica failure [puppet] - 10https://gerrit.wikimedia.org/r/1328709 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [20:22:42] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aaron@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [20:24:15] (03Merged) 10jenkins-bot: Add configurable RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [20:24:27] !log aaron@deploy1003 Started scap sync-world: Backport for [[gerrit:1324819|Add configurable RestModuleOverrides (T434267)]] [20:24:32] T434267: Add $wgRestModuleOverrides config for fragments module - https://phabricator.wikimedia.org/T434267 [20:26:28] !log aaron@deploy1003 aaron, milazg: Backport for [[gerrit:1324819|Add configurable RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:27:07] (03PS1) 10Majavah: Drop irc host from the Beta Cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328711 (https://phabricator.wikimedia.org/T396088) [20:29:27] (03PS3) 10Aleksandar Mastilovic: Disable Presto spilling [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) [20:29:30] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [20:29:45] !log aaron@deploy1003 aaron, milazg: Continuing with deployment [20:30:30] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [20:32:34] (03PS4) 10Aleksandar Mastilovic: Disable Presto spilling [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) [20:32:40] (03CR) 10Aleksandar Mastilovic: [V:03+1 C:03+1] Disable Presto spilling (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [20:34:03] !log aaron@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324819|Add configurable RestModuleOverrides (T434267)]] (duration: 09m 35s) [20:34:08] T434267: Add $wgRestModuleOverrides config for fragments module - https://phabricator.wikimedia.org/T434267 [20:34:33] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aaron@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326410 (https://phabricator.wikimedia.org/T431372) (owner: 10Aaron Schulz) [20:35:42] (03Merged) 10jenkins-bot: Mark all Mathoid-based endpoints as deprecated [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326410 (https://phabricator.wikimedia.org/T431372) (owner: 10Aaron Schulz) [20:35:55] !log aaron@deploy1003 Started scap sync-world: Backport for [[gerrit:1326410|Mark all Mathoid-based endpoints as deprecated (T431372)]] [20:35:59] T431372: Mark Mathoid endpoints as deprecated - https://phabricator.wikimedia.org/T431372 [20:37:30] (03PS1) 10Aleksandar Mastilovic: Disable Presto spilling (production) [puppet] - 10https://gerrit.wikimedia.org/r/1328713 (https://phabricator.wikimedia.org/T435861) [20:37:50] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328713 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [20:37:57] !log aaron@deploy1003 aaron: Backport for [[gerrit:1326410|Mark all Mathoid-based endpoints as deprecated (T431372)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:38:30] (03PS1) 10Cwhite: profile: bugfix: descend into variable to get path [puppet] - 10https://gerrit.wikimedia.org/r/1328715 (https://phabricator.wikimedia.org/T350516) [20:39:44] (03CR) 10Bking: [C:03+2] Kerberos: Don't bail out on a single replica failure [puppet] - 10https://gerrit.wikimedia.org/r/1328709 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [20:40:51] (03CR) 10Majavah: [C:04-1] cloud-vps: kick networkd if network breaks (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [20:44:09] !log aaron@deploy1003 aaron: Continuing with deployment [20:46:03] (03CR) 10Cwhite: [C:03+2] profile: bugfix: descend into variable to get path [puppet] - 10https://gerrit.wikimedia.org/r/1328715 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:47:21] RECOVERY - Check unit status of replicate-krb-database on krb2002 is OK: OK: Status of the systemd unit replicate-krb-database https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [20:47:32] (03CR) 10Bking: [C:03+2] Kerberos: Remove broken host, add replacement host [puppet] - 10https://gerrit.wikimedia.org/r/1328700 (https://phabricator.wikimedia.org/T435854) (owner: 10Bking) [20:48:28] !log aaron@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326410|Mark all Mathoid-based endpoints as deprecated (T431372)]] (duration: 12m 33s) [20:48:33] T431372: Mark Mathoid endpoints as deprecated - https://phabricator.wikimedia.org/T431372 [20:48:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:50:23] ok, done [20:50:40] Is anyone next or can I go? [20:53:34] I guess I'll go [20:53:44] (03CR) 10Bking: [C:03+2] Disable Presto spilling (production) [puppet] - 10https://gerrit.wikimedia.org/r/1328713 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [20:53:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:53:58] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328707 (https://phabricator.wikimedia.org/T432760) (owner: 10Arlolra) [20:55:08] (03Merged) 10jenkins-bot: Set a default for wgUseParsoidLinksUpdate in prod [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328707 (https://phabricator.wikimedia.org/T432760) (owner: 10Arlolra) [20:55:22] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1328707|Set a default for wgUseParsoidLinksUpdate in prod (T432760)]] [20:55:27] T432760: Add configuration options to core for Parsoid by default - https://phabricator.wikimedia.org/T432760 [20:56:44] (03CR) 10Cwhite: prometheus: configure elasticsearch exporter on disable_security_plugin state (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328246 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:57:24] !log arlolra@deploy1003 arlolra: Backport for [[gerrit:1328707|Set a default for wgUseParsoidLinksUpdate in prod (T432760)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:58:11] !log arlolra@deploy1003 arlolra: Continuing with deployment [20:58:21] PROBLEM - Check unit status of replicate-krb-database on krb2002 is CRITICAL: CRITICAL: Status of the systemd unit replicate-krb-database https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [21:00:04] alexsanford, Reedy, sbassett, Maryum, and manfredi: Time to do the Weekly Security deployment window deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T2100). [21:00:43] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12249100 (10cmooney) ` The Lumen NOC has reported that an equipment issue will be corrected during an Urgent Maintenance Network Event on August 25, 2026, between 05:00 GMT and 11:00 GMT... [21:02:29] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328707|Set a default for wgUseParsoidLinksUpdate in prod (T432760)]] (duration: 07m 07s) [21:02:33] T432760: Add configuration options to core for Parsoid by default - https://phabricator.wikimedia.org/T432760 [21:02:43] done [21:04:18] (03PS5) 10Cwhite: mediawiki: enable forward of fatal metrics to statsd exporter [puppet] - 10https://gerrit.wikimedia.org/r/1049625 (https://phabricator.wikimedia.org/T356814) [21:06:31] Hey all - we’d like to get some security patches out today, how are the late backports looking? [21:06:42] (03PS11) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [21:07:17] (03CR) 10CI reject: [V:04-1] cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [21:07:50] (03PS12) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [21:08:25] (03CR) 10CI reject: [V:04-1] cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [21:09:15] (03PS13) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [21:09:50] (03CR) 10CI reject: [V:04-1] cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [21:10:58] (03PS14) 10Andrew Bogott: cloud-vps: kick networkd if network breaks [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) [21:16:14] sbassett: seems to be done, with https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1319945 having been skipped. [21:19:43] (03CR) 10Andrew Bogott: cloud-vps: kick networkd if network breaks (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1318779 (https://phabricator.wikimedia.org/T432426) (owner: 10Andrew Bogott) [21:19:45] AaronSchulz: thanks [21:21:23] !log Deployed security mitigation for T435455 [21:21:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:22:43] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873 (10bking) 03NEW [21:34:08] (03CR) 10Aleksandar Mastilovic: [V:03+1 C:03+1] Disable Presto spilling (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328694 (https://phabricator.wikimedia.org/T435861) (owner: 10Aleksandar Mastilovic) [21:38:08] o/ I'd like to get in line for a scap release deployment. [21:40:26] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Alerting: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801#12249177 (10cmooney) Ok thanks that's a major issue with our stats pipeline then. The cr2-eqord we can ignore, if we have no stats for BGP that's a... [21:41:03] !log Deployed security fix for T435624 [21:41:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:41:37] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12249178 (10bking) 05In progress→03Open a:05bking→03VRiley-WMF @jhathaway Thanks for spending some time with this. If your tests were hitting th... [21:43:17] Nevermind! I will wait until tomorrow morning. Have a lovely day everyone. [21:44:09] !log ryankemper@cumin2003 START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons. [21:44:55] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12249186 (10bking) [21:45:05] !log T435861 [presto] Kicking off rolling restart to pick up `experimental.spill-enabled=false` changes [21:45:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:45:10] T435861: Revert Presto configuration to disable spilling - https://phabricator.wikimedia.org/T435861 [21:49:36] !log T435862 [presto] Rolling restart in progress to activate `experimental.spill-enabled=false `from https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328713 [21:49:41] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:49:42] T435862: With spill_enabled = true, Presto fails queries with CROSS JOIN - https://phabricator.wikimedia.org/T435862 [21:51:04] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:52:40] (03PS1) 10Kimberly Sarabia: Deploy ReaderExperiments in fa,cs,bn [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328721 [21:55:28] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Bring krb1003 into service - https://phabricator.wikimedia.org/T435854#12249285 (10bking) 05In progress→03Resolved We've successfully deployed `krb1003` as a Kerberos replica. There's no need to make... [21:55:45] 06SRE, 06Infrastructure-Foundations, 10netops, 10Observability-Alerting: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801#12249291 (10cmooney) 05Open→03Resolved a:03cmooney Actually I worked it out. We stopped getting stats on July 15 when we upgraded the rou... [21:55:48] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Bring krb1003 into service - https://phabricator.wikimedia.org/T435854#12249290 (10bking) [21:57:09] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12249309 (10bking) [21:57:39] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12249310 (10bking) 05Open→03In progress [22:05:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:17:50] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons. [22:26:23] jouncebot: nowandnext [22:26:23] For the next 0 hour(s) and 33 minute(s): Weekly Security deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T2100) [22:26:23] In 0 hour(s) and 33 minute(s): Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T2300) [22:26:52] Anyone using the security deployment window? I'd like to use scap [22:28:49] Yes, security deploys currently running [22:28:54] !log Deployed security fix for T435822 [22:28:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:33:22] RScout-WMF: Can you ping when done? [22:40:35] I'm running scap for one more patch Dreamy_Jazz and then we should be wrapping up [22:40:54] Thanks [22:47:34] !log Deployed security fix for T435457 [22:47:37] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:47:43] Dreamy_Jazz that should be the last security patch for today [22:51:42] (03PS1) 10BryanDavis: P:mediawiki::php: add 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) [22:54:06] (03PS1) 10Jdlrobson: [beta] Set $wgCentralNoticeProhibitedExperiments [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328722 [22:55:04] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [22:56:00] (03PS2) 10BryanDavis: P:mediawiki::php: add 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) [22:57:38] !log dreamyjazz@deploy1003 dreamyjazz: WikimediaAntiAbusePrivate change for T432848 synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [22:57:44] T432848: EventLogging: Record confidence score results - https://phabricator.wikimedia.org/T432848 [23:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260824T2300) [23:00:28] !log dreamyjazz@deploy1003 dreamyjazz: Continuing with deployment [23:00:59] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12249532 (10RobH) DHL called me back today, they're still trying to get in contact with the local field office for the drop off to find out what is going on with the package... [23:01:59] Dreamy_Jazz: lemme know when you are done so I can go ahead with some deploys. thanks in advance! [23:02:17] Sure, just waiting for scap to sync this out and then will be done [23:04:48] !log dreamyjazz@deploy1003 Synchronized private/PrivateSettings.php: WikimediaAntiAbusePrivate change for T432848 (duration: 08m 59s) [23:04:52] T432848: EventLogging: Record confidence score results - https://phabricator.wikimedia.org/T432848 [23:05:19] Jdlrobson: Over to you [23:05:21] (03PS1) 10Cwhite: logs-api: set authentication headers to be forwarded to OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1328731 (https://phabricator.wikimedia.org/T350516) [23:05:23] (03PS1) 10Cwhite: profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) [23:05:34] thanks Dreamy_Jazz [23:05:55] (03CR) 10CI reject: [V:04-1] logs-api: set authentication headers to be forwarded to OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1328731 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:06:29] (03CR) 10CI reject: [V:04-1] profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:06:59] (03PS2) 10Cwhite: logs-api: set authentication headers to be forwarded to OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1328731 (https://phabricator.wikimedia.org/T350516) [23:07:08] (03PS2) 10Cwhite: profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) [23:07:57] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328722 (owner: 10Jdlrobson) [23:07:57] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328721 (owner: 10Kimberly Sarabia) [23:07:58] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328376 (owner: 10Jdlrobson) [23:08:11] (03PS2) 10Jdlrobson: Support change tags for anonymous users on English Wikivoyage [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328376 [23:08:18] (03CR) 10TrainBranchBot: "Approved by jdlrobson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328376 (owner: 10Jdlrobson) [23:09:09] (03Merged) 10jenkins-bot: [beta] Set $wgCentralNoticeProhibitedExperiments [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328722 (owner: 10Jdlrobson) [23:09:12] (03Merged) 10jenkins-bot: Deploy ReaderExperiments in fa,cs,bn [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328721 (owner: 10Kimberly Sarabia) [23:09:15] (03Merged) 10jenkins-bot: Support change tags for anonymous users on English Wikivoyage [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328376 (owner: 10Jdlrobson) [23:09:29] !log jdlrobson@deploy1003 Started scap sync-world: Backport for [[gerrit:1328722|[beta] Set $wgCentralNoticeProhibitedExperiments]], [[gerrit:1328721|Deploy ReaderExperiments in fa,cs,bn]], [[gerrit:1328376|Support change tags for anonymous users on English Wikivoyage]] [23:10:45] (03PS3) 10Cwhite: profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) [23:11:28] (03CR) 10CI reject: [V:04-1] profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:11:31] !log jdlrobson@deploy1003 ksarabia, jdlrobson: Backport for [[gerrit:1328722|[beta] Set $wgCentralNoticeProhibitedExperiments]], [[gerrit:1328721|Deploy ReaderExperiments in fa,cs,bn]], [[gerrit:1328376|Support change tags for anonymous users on English Wikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [23:11:50] (03PS4) 10Cwhite: profile: configure apache to optionally connect to OpenSearch with TLS [puppet] - 10https://gerrit.wikimedia.org/r/1328732 (https://phabricator.wikimedia.org/T350516) [23:14:08] !log jdlrobson@deploy1003 ksarabia, jdlrobson: Continuing with deployment [23:14:57] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [23:18:27] !log jdlrobson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328722|[beta] Set $wgCentralNoticeProhibitedExperiments]], [[gerrit:1328721|Deploy ReaderExperiments in fa,cs,bn]], [[gerrit:1328376|Support change tags for anonymous users on English Wikivoyage]] (duration: 08m 58s) [23:19:22] FIRING: GnmiInterfaceCountersDrop: ... [23:19:23] lsw1-e8-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-e8-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [23:20:40] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12249586 (10Eevans) Hi @Bethany, I can take care of updating that key for you. Can we first confirm that key out-of-band? The easiest way is to add it one of your on... [23:21:08] all done [23:23:14] 06SRE, 10Wikimedia-Mailing-lists: Create new mailing lists: wikidata-admins@lists.wikimedia.org - https://phabricator.wikimedia.org/T435638#12249588 (10Quiddity) Also it needs a second admin, to start. (To prevent a SPOF (single point of failure)). [23:23:53] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr2-eqiad:et-1/1/5 (Transport: cr2-codfw:et-0/1/4 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [23:25:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [23:27:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:28:47] (03PS2) 10Tim Starling: Enable Produnto on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324965 (https://phabricator.wikimedia.org/T421436) [23:29:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:31:49] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:32:10] FIRING: [5x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:36:39] FIRING: [2x] TransitBGPDown: Transit BGP session down between cr2-eqsin and Hurricane Electric (103.231.152.47) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [23:37:10] FIRING: [7x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:39:46] (03PS1) 10Eevans: Record that a principal was created for user mkrolik-wmf [puppet] - 10https://gerrit.wikimedia.org/r/1328736 (https://phabricator.wikimedia.org/T434877) [23:41:11] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1328737 [23:41:11] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1328737 (owner: 10TrainBranchBot) [23:42:10] RESOLVED: [10x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:42:40] FIRING: [3x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:42:40] FIRING: [6x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:46:46] RESOLVED: [2x] TransitBGPDown: Transit BGP session down between cr2-eqsin and Hurricane Electric (103.231.152.47) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [23:46:46] PROBLEM - OSPF status on cr2-codfw is CRITICAL: OSPFv2: 8/8 UP : OSPFv3: 7/7 UP : 8 v2 P2P interfaces vs. 7 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:47:25] RESOLVED: [11x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:49:07] 06SRE, 10SRE-Access-Requests, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12249642 (10Eevans) Hi @MKrolik-WMF, I created your kerberos principal, and you should have gotten an... [23:49:16] 06SRE, 10SRE-Access-Requests, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12249643 (10Eevans) 05Stalled→03In progress [23:49:35] RECOVERY - OSPF status on cr2-codfw is OK: OSPFv2: 8/8 UP : OSPFv3: 8/8 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:49:38] 06SRE, 10SRE-Access-Requests, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12249644 (10Eevans) p:05Triage→03Medium [23:49:41] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1328737 (owner: 10TrainBranchBot) [23:49:51] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 10/10 UP : OSPFv3: 9/9 UP : 10 v2 P2P interfaces vs. 9 v3 P2P interfaces https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:50:51] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 9/9 UP : OSPFv3: 9/9 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [23:57:26] !log on db1220 x1 wikishared created Produnto tables T421436 [23:57:29] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [23:57:30] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436