[00:01:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:01:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:02:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:04:40] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:05:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:06:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:09:16] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:09:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:11:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:13:48] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:14:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:16:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:16:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:18:50] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:20:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:22:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:25:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:25:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:26:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:28:37] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [00:29:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:29:56] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [00:31:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:31:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [00:32:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:35:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:35:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:38:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:41:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:43:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:44:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:46:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [00:48:07] (03PS5) 10Ryan Kemper: wdqs: Absent unneeded monitors [puppet] - 10https://gerrit.wikimedia.org/r/1327181 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [00:48:07] (03PS2) 10Ryan Kemper: wdqs: remove unneeded monitors [puppet] - 10https://gerrit.wikimedia.org/r/1327182 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [00:48:18] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1327181 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [00:48:19] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1327182 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [00:52:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:52:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:55:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [00:56:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [00:59:02] (03CR) 10Ryan Kemper: [C:03+1] "PS5 only adds wdqs1027 (internal_scholarly) to the Hosts: footer, no code change. PCC now covers both files: the two checks flip to absent" [puppet] - 10https://gerrit.wikimedia.org/r/1327181 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [00:59:14] (03CR) 10Ryan Kemper: [C:03+1] "Same Hosts: footer fix as the parent. PCC shows the 21 check-related resources dropping from the catalog on wdqs1025/wdqs1027, noop elsewh" [puppet] - 10https://gerrit.wikimedia.org/r/1327182 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [01:00:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:01:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:05:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:05:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:06:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:06:23] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d4-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [01:09:01] (03CR) 10Ryan Kemper: [C:03+2] "merging and will let puppet auto run before merging the removal patch" [puppet] - 10https://gerrit.wikimedia.org/r/1327181 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [01:09:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:11:23] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1327234 [01:11:23] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1327234 (owner: 10TrainBranchBot) [01:12:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:12:34] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:19:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:19:34] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:19:56] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1327234 (owner: 10TrainBranchBot) [01:21:06] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:24:06] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:24:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:25:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:26:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:28:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:28:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:32:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:33:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:40:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:40:39] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:41:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:44:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [01:46:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:47:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [01:54:04] (03CR) 10Ryan Kemper: [C:03+2] "absented resources confirmed gone from wdqs*, merging" [puppet] - 10https://gerrit.wikimedia.org/r/1327182 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [01:54:14] (03CR) 10Ryan Kemper: [C:03+2] "doh, missed a merge conflict. fixing" [puppet] - 10https://gerrit.wikimedia.org/r/1327182 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [01:55:37] PROBLEM - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [01:59:35] (03PS3) 10Ryan Kemper: wdqs: remove unneeded monitors [puppet] - 10https://gerrit.wikimedia.org/r/1327182 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [02:00:43] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:01:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:02:03] (03CR) 10Ryan Kemper: [C:03+2] wdqs: remove unneeded monitors [puppet] - 10https://gerrit.wikimedia.org/r/1327182 (https://phabricator.wikimedia.org/T358029) (owner: 10Bking) [02:03:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:06:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:06:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:08:33] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 49s) [02:09:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:10:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:11:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:14:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:16:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:16:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:19:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:19:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:24:03] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:32:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:34:11] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [02:35:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:35:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:37:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:41:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1014.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:41:23] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d4-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [02:42:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:46:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:46:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 19.7% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:46:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:49:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:50:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [02:55:37] RECOVERY - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [02:56:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:57:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [02:57:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [03:00:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [03:00:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [03:01:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [03:01:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [03:04:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [03:05:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [03:06:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [03:06:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [03:14:37] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [03:15:05] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [03:16:05] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [03:16:37] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [03:21:15] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [03:22:40] FIRING: SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:24:44] 10ops-codfw, 06SRE, 06DC-Ops: Degraded RAID on maps-test2001 - https://phabricator.wikimedia.org/T435311#12235123 (10Jhancock.wm) I don't see any disk errors when i log into the server, but i think i see which one might be the issue on this server. However, the server is out of warranty and i do not have th... [04:28:37] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [05:26:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:43:32] (03CR) 10Arnaudb: [C:03+1] wmnet: phabricator -> phab1005.eqiad.wmnet [dns] - 10https://gerrit.wikimedia.org/r/1327141 (https://phabricator.wikimedia.org/T435087) (owner: 10AOkoth) [05:43:51] (03CR) 10Arnaudb: [C:03+1] hiera: make phab1005 phabricator_active_server [puppet] - 10https://gerrit.wikimedia.org/r/1327144 (https://phabricator.wikimedia.org/T435087) (owner: 10AOkoth) [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T0600) [06:00:05] marostegui, Amir1, and federico3: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for Primary database switchover. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T0600). [06:00:05] arnoldokoth, brennen, and andre: gettimeofday() says it's time for Phabricator maintenance window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T0600) [06:00:20] o/ [06:00:25] o/ [06:02:53] (03CR) 10AOkoth: [C:03+2] hiera: make phab1005 phabricator_active_server [puppet] - 10https://gerrit.wikimedia.org/r/1327144 (https://phabricator.wikimedia.org/T435087) (owner: 10AOkoth) [06:08:05] PROBLEM - PHD should be running on phab1004 is CRITICAL: PROCS CRITICAL: 0 processes with regex args php ./phd-daemon, UID = 920 (phd) https://wikitech.wikimedia.org/wiki/Phabricator [06:08:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [06:13:54] andre [06:13:57] phan gone [06:13:58] PROBLEM - PHD should be running on phab1004 is CRITICAL: PROCS CRITICAL: 0 processes with regex args php ./phd-daemon, UID = 920 (phd) [06:14:03] i guess this is source [06:14:09] and i see upstream connect error or disconnect/reset before headers. retried and the latest reset reason: remote connection failure, transport failure reason: delayed connect error: Connection refused [06:14:22] Les4353: Phabricator has a planned maintenance now, see https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/GHNYRAXG3WVZSAEB2FHWOT57K5356KES/ [06:14:37] ohh [06:14:42] Please see the message by jouncebot above. [06:14:44] wait we have mailing lists?? [06:14:52] oh isaw [06:15:11] FIRING: ProbeDown: Service phab1004:443 has failed probes (http_phabricator_wikimedia_org_collab_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#phab1004:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [06:17:15] PROBLEM - PHD should be running on phab1005 is CRITICAL: PROCS CRITICAL: 0 processes with regex args php ./phd-daemon, UID = 920 (phd) https://wikitech.wikimedia.org/wiki/Phabricator [06:17:55] (03CR) 10Klausman: [C:03+1] cirrussearch: revert to default scaling governor [puppet] - 10https://gerrit.wikimedia.org/r/1327187 (https://phabricator.wikimedia.org/T435400) (owner: 10Bking) [06:18:36] !log brennen@deploy1003 Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for T435087 [06:20:22] !log brennen@deploy1003 Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for to pick up config values for T435087 (duration: 01m 46s) [06:20:41] FIRING: ProbeDown: Service phab1005:443 has failed probes (http_phabricator_wikimedia_org_collab_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#phab1005:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [06:21:39] !log brennen@deploy1003 Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for T435087 [06:22:18] !log brennen@deploy1003 Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1005 for T435087 (duration: 00m 39s) [06:23:15] RECOVERY - PHD should be running on phab1005 is OK: PROCS OK: 1 process with regex args php ./phd-daemon, UID = 920 (phd) https://wikitech.wikimedia.org/wiki/Phabricator [06:24:54] (03CR) 10AOkoth: [C:03+2] wmnet: phabricator -> phab1005.eqiad.wmnet [dns] - 10https://gerrit.wikimedia.org/r/1327141 (https://phabricator.wikimedia.org/T435087) (owner: 10AOkoth) [06:25:28] !log aokoth@dns1004 START - running authdns-update [06:27:42] !log aokoth@dns1004 END - running authdns-update [06:33:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [06:40:40] arnoldokoth, andre: mostly stepping away from computer but i'll be awake for a bit yet. please ping if something does come up. [06:41:37] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d4-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [06:47:32] ack [06:48:10] brennen: Ack. [06:53:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [06:58:49] 06SRE, 10Wikimedia-Mailing-lists: Unable to subscribe to mailing lists anonymously: "You are doing that too often. Please try again later." - https://phabricator.wikimedia.org/T435424#12235283 (10Aklapper) [07:00:05] Amir1, urbanecm, and awight: Time to do the UTC morning backport window deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T0700). [07:00:05] No Gerrit patches in the queue for this window AFAICS. [07:17:08] (03PS1) 10Arthur taylor: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327436 (https://phabricator.wikimedia.org/T427589) [07:18:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:19:11] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [07:19:13] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [07:19:13] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [07:22:40] FIRING: SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:31:29] (03PS1) 10Marostegui: db2207: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1327438 (https://phabricator.wikimedia.org/T435271) [07:32:09] (03CR) 10Marostegui: [C:03+2] db2207: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1327438 (https://phabricator.wikimedia.org/T435271) (owner: 10Marostegui) [07:33:44] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:34:03] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [07:49:44] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:50:08] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [07:50:08] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [07:50:08] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.131 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [07:52:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1016:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1016 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [07:55:47] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12235382 (10JMeybohm) 05Open→03Resolved I'm tentatively closing this. Please reopen the task if you prefer the SSH key to be changed or if there are other issu... [07:59:11] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12235394 (10MoritzMuehlenhoff) The initial buildout of the proxies had a memory leak, which can e.g. be seen here: https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&from=2026-07-08T10:2... [08:00:04] andre and brennen: That opportune time for a MediaWiki train - Utc-0+Utc-7 Version deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T0800). [08:00:05] o/ [08:05:22] 10ops-eqiad, 06SRE, 06DC-Ops: krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12235413 (10MoritzMuehlenhoff) >>! In T435354#12233738, @VRiley-WMF wrote: > Upon logging into the unit, there doesn't seem to be any hardware issues. Were you able to log in locally on a tty console... [08:07:31] (03CR) 10JMeybohm: [C:03+1] admin_ng: Use kube-state-metrics 7.3.0 in staging-codfw. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327093 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [08:10:17] (03CR) 10Jelto: [C:03+2] wikikube-staging eqiad: Update to cert-manager 1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327074 (https://phabricator.wikimedia.org/T427402) (owner: 10JMeybohm) [08:10:21] (03CR) 10JMeybohm: [C:03+1] docker_registry: route /v2/dev/.* to the releng S3 bucket [puppet] - 10https://gerrit.wikimedia.org/r/1327146 (https://phabricator.wikimedia.org/T432829) (owner: 10Elukey) [08:10:43] (03PS1) 10TrainBranchBot: group2 to 1.47.0-wmf.16 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327488 (https://phabricator.wikimedia.org/T430835) [08:10:46] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by aklapper@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327488 (https://phabricator.wikimedia.org/T430835) (owner: 10TrainBranchBot) [08:11:44] (03Merged) 10jenkins-bot: group2 to 1.47.0-wmf.16 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327488 (https://phabricator.wikimedia.org/T430835) (owner: 10TrainBranchBot) [08:12:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1016:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1016 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [08:14:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:15:59] (03PS1) 10Slyngshede: IDP: Upgrade CAS to 7.3.8.1 [dns] - 10https://gerrit.wikimedia.org/r/1327490 [08:16:02] (03CR) 10Klausman: [C:03+1] kserve: update images to 0.20 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1327059 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [08:16:08] (03PS1) 10Muehlenhoff: Switch the new URL downloaders to insetup_ferm for the reimage [puppet] - 10https://gerrit.wikimedia.org/r/1327491 (https://phabricator.wikimedia.org/T427282) [08:16:40] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [dns] - 10https://gerrit.wikimedia.org/r/1327490 (owner: 10Slyngshede) [08:16:59] FIRING: KafkaMirrorMakerConsumerMaxLag: Kafka MirrorMaker main-eqiad-to-main-codfw max lag in last 10 minutes - https://wikitech.wikimedia.org/wiki/Kafka/Administration#MirrorMaker - https://grafana.wikimedia.org/d/000000521/kafka-mirrormaker?var-mirror_name=main-eqiad-to-main-codfw - https://alerts.wikimedia.org/?q=alertname%3DKafkaMirrorMakerConsumerMaxLag [08:17:17] (03Abandoned) 10Muehlenhoff: Stop using nftables for the new URL downloader trixie nodes [puppet] - 10https://gerrit.wikimedia.org/r/1327131 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [08:18:16] !log aklapper@deploy1003 rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.16 refs T430835 [08:18:21] T430835: 1.47.0-wmf.16 deployment blockers - https://phabricator.wikimedia.org/T430835 [08:20:35] (03Merged) 10jenkins-bot: wikikube-staging eqiad: Update to cert-manager 1.19.6 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327074 (https://phabricator.wikimedia.org/T427402) (owner: 10JMeybohm) [08:21:21] (03CR) 10Slyngshede: [C:03+2] IDP: Upgrade CAS to 7.3.8.1 [dns] - 10https://gerrit.wikimedia.org/r/1327490 (owner: 10Slyngshede) [08:21:37] !log slyngshede@dns1004 START - running authdns-update [08:23:50] !log slyngshede@dns1004 END - running authdns-update [08:24:18] (03CR) 10Jelto: [V:03+1 C:03+1] "lgtm, build verified locally" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1321574 (https://phabricator.wikimedia.org/T433590) (owner: 10JMeybohm) [08:25:24] (03CR) 10Blake: [C:03+1] trafficserver: Support testwiki pretrain routing in XWD (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1304189 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [08:25:30] (03CR) 10JMeybohm: [V:03+2 C:03+2] Add coredns 1.12 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1321574 (https://phabricator.wikimedia.org/T433590) (owner: 10JMeybohm) [08:26:01] (03CR) 10Blake: [C:03+1] mw-pretrain: Relax x-wikimedia-debug match regex [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327164 (https://phabricator.wikimedia.org/T427668) (owner: 10Scott French) [08:28:37] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:30:33] (03PS4) 10JMeybohm: coredns: Update to 1.12.1, add kubepods support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327115 (https://phabricator.wikimedia.org/T428573) [08:34:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:40:51] (03PS1) 10Tiziano Fogli: kafka-logging: add kafka-logging100[6-8] to eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) [08:47:35] FIRING: DiskSpace: Disk space build2001:9100:/ 2.897% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=build2001 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [08:48:28] (03PS2) 10Tiziano Fogli: kafka-logging: add kafka-logging100[6-8] to eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) [08:48:35] (03CR) 10Tiziano Fogli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [08:49:22] (03CR) 10Lucas Werkmeister (WMDE): [C:03+1] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327436 (https://phabricator.wikimedia.org/T427589) (owner: 10Arthur taylor) [08:49:24] 06SRE, 10SRE-Access-Requests, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12235633 (10JMeybohm) [08:51:57] (03PS1) 10Aklapper: mariadb: add grants to phstats for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327497 (https://phabricator.wikimedia.org/T435442) [08:53:52] (03PS1) 10JMeybohm: Add mkrolik-wmf SSH and analytics-privatedata-users access [puppet] - 10https://gerrit.wikimedia.org/r/1327498 (https://phabricator.wikimedia.org/T434877) [08:54:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:56:09] (03PS2) 10Aklapper: mariadb: add grants to phstats for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327497 (https://phabricator.wikimedia.org/T435442) [08:56:28] (03CR) 10JMeybohm: [C:03+2] Add mkrolik-wmf SSH and analytics-privatedata-users access [puppet] - 10https://gerrit.wikimedia.org/r/1327498 (https://phabricator.wikimedia.org/T434877) (owner: 10JMeybohm) [08:56:59] RESOLVED: KafkaMirrorMakerConsumerMaxLag: Kafka MirrorMaker main-eqiad-to-main-codfw max lag in last 10 minutes - https://wikitech.wikimedia.org/wiki/Kafka/Administration#MirrorMaker - https://grafana.wikimedia.org/d/000000521/kafka-mirrormaker?var-mirror_name=main-eqiad-to-main-codfw - https://alerts.wikimedia.org/?q=alertname%3DKafkaMirrorMakerConsumerMaxLag [08:59:36] jouncebot: nowandnext [08:59:37] For the next 1 hour(s) and 0 minute(s): MediaWiki train - Utc-0+Utc-7 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T0800) [08:59:37] In 1 hour(s) and 0 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1000) [09:01:20] (03CR) 10Arthur taylor: [C:03+2] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327436 (https://phabricator.wikimedia.org/T427589) (owner: 10Arthur taylor) [09:02:13] (03CR) 10Arnaudb: "I think the grants will have to be updated on production" [puppet] - 10https://gerrit.wikimedia.org/r/1327497 (https://phabricator.wikimedia.org/T435442) (owner: 10Aklapper) [09:03:59] (03Merged) 10jenkins-bot: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327436 (https://phabricator.wikimedia.org/T427589) (owner: 10Arthur taylor) [09:04:21] !log arthurtaylor@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [09:04:53] 06SRE, 10SRE-Access-Requests, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12235714 (10JMeybohm) Level 2 access has been granted and should roll out to the infra in the next 30... [09:04:53] !log arthurtaylor@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [09:06:00] (03CR) 10Lucas Werkmeister (WMDE): [C:03+1] RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327219 (https://phabricator.wikimedia.org/T435422) (owner: 10Krinkle) [09:07:31] (03PS3) 10Aklapper: mariadb: add grants to phstats for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327497 (https://phabricator.wikimedia.org/T435442) [09:07:55] !log arthurtaylor@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [09:08:16] !log arthurtaylor@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [09:08:24] !log arthurtaylor@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [09:08:43] !log arthurtaylor@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [09:09:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [09:10:29] (03CR) 10Jelto: [C:03+1] "lgtm" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327115 (https://phabricator.wikimedia.org/T428573) (owner: 10JMeybohm) [09:15:05] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1323943 (owner: 10PipelineBot) [09:15:12] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324530 (owner: 10PipelineBot) [09:15:20] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324531 (owner: 10PipelineBot) [09:15:27] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324550 (owner: 10PipelineBot) [09:15:34] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324574 (owner: 10PipelineBot) [09:15:44] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324771 (owner: 10PipelineBot) [09:15:53] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325431 (owner: 10PipelineBot) [09:16:02] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325433 (owner: 10PipelineBot) [09:16:10] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325489 (owner: 10PipelineBot) [09:19:27] (03CR) 10Clément Goubert: [C:03+1] "LGTM, I should have documented better what "internals" I was talking about because I don't recall exactly what made it problematic at that" [puppet] - 10https://gerrit.wikimedia.org/r/1325935 (https://phabricator.wikimedia.org/T434925) (owner: 10Scott French) [09:20:39] !log imported squid 7.6-2.1for trixie-wikimedia/main T427282 [09:20:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:20:43] T427282: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282 [09:21:28] (03PS1) 10Ryan Kemper: rkemper-kafka: add README for pontoon stack [puppet] - 10https://gerrit.wikimedia.org/r/1327500 (https://phabricator.wikimedia.org/T276088) [09:22:02] (03CR) 10CI reject: [V:04-1] rkemper-kafka: add README for pontoon stack [puppet] - 10https://gerrit.wikimedia.org/r/1327500 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [09:22:17] !log update cert-manager to 1.19.6 on wikikube staging-eqiad - T427402 [09:22:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:22:22] T427402: Update cert-manager to 1.19 - https://phabricator.wikimedia.org/T427402 [09:22:52] !log jelto@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'sync'. [09:23:26] !log jelto@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. [09:23:57] (03PS2) 10Ryan Kemper: rkemper-kafka: add README for pontoon stack [puppet] - 10https://gerrit.wikimedia.org/r/1327500 (https://phabricator.wikimedia.org/T276088) [09:24:35] (03CR) 10JMeybohm: [C:03+2] coredns: Update to 1.12.1, add kubepods support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327115 (https://phabricator.wikimedia.org/T428573) (owner: 10JMeybohm) [09:26:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:33:02] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12235869 (10Marostegui) I've disabled notifications for a few days, in case we need further work on it. [09:33:29] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12235875 (10Marostegui) >>! In T435271#12233874, @Jhancock.wm wrote: > also, should the masters be moved off servers that have Intel Xeon Gold 5317? I have a list in a sheet that should be searchable... [09:33:39] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin2002 from Cumin master firewall rules [puppet] - 10https://gerrit.wikimedia.org/r/1327080 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [09:34:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [09:34:51] (03Merged) 10jenkins-bot: coredns: Update to 1.12.1, add kubepods support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327115 (https://phabricator.wikimedia.org/T428573) (owner: 10JMeybohm) [09:35:53] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin2002 from alertmanager config [puppet] - 10https://gerrit.wikimedia.org/r/1327083 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [09:36:24] (03PS3) 10Muehlenhoff: docker::baseimages: Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327103 (https://phabricator.wikimedia.org/T417389) [09:36:26] 06SRE, 10DNS, 10Domains, 06Traffic-Icebox, 07HTTPS: Merge Wikipedia subdomains into one, to discourage censorship - https://phabricator.wikimedia.org/T215071#12235881 (10Tgr) Domains are first of all trust boundaries on the modern web, and with our different wikis being written by different communities,... [09:48:21] (03CR) 10Muehlenhoff: [C:03+2] docker::baseimages: Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327103 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [09:54:44] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [09:58:12] (03CR) 10Blake: [C:03+2] admin_ng: Use kube-state-metrics 7.3.0 in staging-codfw. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327093 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1000) [10:00:14] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [10:05:54] (03PS1) 10Muehlenhoff: Remove cumin2002 from sync list for firmware [puppet] - 10https://gerrit.wikimedia.org/r/1327503 (https://phabricator.wikimedia.org/T427897) [10:07:18] (03PS1) 10Muehlenhoff: Remove cumin2002 from tcpircbot config [puppet] - 10https://gerrit.wikimedia.org/r/1327504 (https://phabricator.wikimedia.org/T427897) [10:08:00] PROBLEM - Blazegraph Port for wdqs-blazegraph on wdqs1019 is CRITICAL: connect to address 127.0.0.1 and port 9999: Connection refused https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [10:08:56] (03Merged) 10jenkins-bot: admin_ng: Use kube-state-metrics 7.3.0 in staging-codfw. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327093 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [10:09:00] RECOVERY - Blazegraph Port for wdqs-blazegraph on wdqs1019 is OK: TCP OK - 0.000 second response time on 127.0.0.1 port 9999 https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [10:10:14] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [10:13:19] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet [10:16:33] jouncebot: nowandnext [10:16:33] For the next 0 hour(s) and 43 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1000) [10:16:33] In 1 hour(s) and 43 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1200) [10:16:43] !log blake@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [10:17:47] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [10:18:46] !log blake@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [10:20:35] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) [10:20:37] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1901.eqiad.wmnet [10:23:32] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm [10:25:44] (03PS1) 10Btullis: Install amd_rocm version 6.1 on stat1010 after reimage to bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1327505 (https://phabricator.wikimedia.org/T434562) [10:25:49] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 20 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326848 (https://phabricator.wikimedia.org/T433713) (owner: 10Sadiya.mohammed13) [10:26:42] (03CR) 10Btullis: [C:03+2] Install amd_rocm version 6.1 on stat1010 after reimage to bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1327505 (https://phabricator.wikimedia.org/T434562) (owner: 10Btullis) [10:27:14] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet [10:27:16] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [10:31:51] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" [10:31:56] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" [10:31:56] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:31:56] !log fceratto@cumin1003 START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors [10:32:00] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors [10:32:32] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" [10:32:35] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1901.eqiad.wmnet - fceratto@cumin1003" [10:35:34] !log fceratto@cumin1003 START - Cookbook sre.hosts.reimage for host db1901.eqiad.wmnet with OS trixie [10:37:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [10:41:21] PROBLEM - Check whether ferm is active by checking the default input chain on wikikube-worker1262 is CRITICAL: ERROR ferm input drop default policy not set, ferm might not have been started correctly https://wikitech.wikimedia.org/wiki/Monitoring/check_ferm [10:41:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d4-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:47:40] !log fceratto@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage [10:47:52] !log bump space for prometheus k8s-dse in eqiad [10:47:54] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:47:58] !log bump space for prometheus k8s-aux in codfw [10:48:00] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:51:13] PROBLEM - Check whether ferm is active by checking the default input chain on wikikube-worker1068 is CRITICAL: ERROR ferm input drop default policy not set, ferm might not have been started correctly https://wikitech.wikimedia.org/wiki/Monitoring/check_ferm [10:52:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [10:53:46] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1901.eqiad.wmnet with reason: host reimage [10:59:47] (03PS1) 10Mszwarc: UIC: Fix page:page instead of page:other in instrumentation [extensions/CheckUser] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327511 [11:01:45] (03PS1) 10Muehlenhoff: docker::builder:Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327513 (https://phabricator.wikimedia.org/T417389) [11:02:10] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 20 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [extensions/CheckUser] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327511 (owner: 10Mszwarc) [11:02:20] (03CR) 10CI reject: [V:04-1] docker::builder:Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327513 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:02:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [11:04:48] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:05:51] (03PS2) 10Muehlenhoff: docker::builder:Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327513 (https://phabricator.wikimedia.org/T417389) [11:06:25] (03CR) 10CI reject: [V:04-1] docker::builder:Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327513 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:07:37] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1901.eqiad.wmnet with OS trixie [11:07:37] !log fceratto@cumin1003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1901.eqiad.wmnet [11:10:35] (03PS3) 10Muehlenhoff: docker::builder:Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327513 (https://phabricator.wikimedia.org/T417389) [11:11:18] (03PS1) 10Btullis: Add a new partman reuse recipe for hardware RAID two drive hosts [puppet] - 10https://gerrit.wikimedia.org/r/1327516 (https://phabricator.wikimedia.org/T434562) [11:12:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [11:14:19] btullis@cumin1003 reimage (PID 1808650) is awaiting input [11:19:18] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327513 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:21:49] (03CR) 10Btullis: [C:03+2] Add a new partman reuse recipe for hardware RAID two drive hosts [puppet] - 10https://gerrit.wikimedia.org/r/1327516 (https://phabricator.wikimedia.org/T434562) (owner: 10Btullis) [11:22:36] (03PS1) 10Ladsgroup: Enable thumb.wikimedia.org on cswiki and fawiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327518 (https://phabricator.wikimedia.org/T427465) [11:22:40] FIRING: SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:24:11] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host stat1010.eqiad.wmnet with OS bookworm [11:24:25] jouncebot: nowandnext [11:24:25] No deployments scheduled for the next 0 hour(s) and 35 minute(s) [11:24:25] In 0 hour(s) and 35 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1200) [11:24:26] 10ops-eqiad, 06SRE, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Q1:rack/setup/install conf101[0-2] - https://phabricator.wikimedia.org/T435426#12236303 (10Clement_Goubert) @jasmine_ Can you take care of adding the servers to `site.pp` and `preseed.yml` please? [11:25:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327518 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [11:26:00] (03Merged) 10jenkins-bot: Enable thumb.wikimedia.org on cswiki and fawiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327518 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [11:26:34] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1327518|Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] [11:26:40] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [11:29:50] (03CR) 10Hnowlan: sre/cdn: create recording rule in advance of moving to ratio for CDN (032 comments) [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [11:30:49] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1327518|Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:33:00] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [11:33:35] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm [11:35:03] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db1902.eqiad.wmnet [11:37:15] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1327518|Enable thumb.wikimedia.org on cswiki and fawiki (T427465)]] (duration: 10m 40s) [11:37:20] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [11:39:38] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [11:40:53] (03CR) 10Elukey: [C:03+1] docker::builder:Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327513 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:41:11] (03CR) 10Elukey: [C:03+1] Remove cumin2002 from tcpircbot config [puppet] - 10https://gerrit.wikimedia.org/r/1327504 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [11:41:31] (03CR) 10Elukey: [C:03+1] Remove cumin2002 from sync list for firmware [puppet] - 10https://gerrit.wikimedia.org/r/1327503 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [11:43:22] FIRING: [2x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:43:48] PROBLEM - Check unit status of statograph_post on alert1002 is CRITICAL: CRITICAL: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [11:45:08] fceratto@cumin1003 decommission (PID 1871968) is awaiting input [11:46:43] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [11:48:17] (03CR) 10Filippo Giunchedi: "Following up from IRC convo: my proposal is to define a second recording rule with the per-backend availability, then we can use that in t" [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [11:48:22] FIRING: [2x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:48:39] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [11:49:04] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [11:51:39] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [11:52:58] (03PS1) 10Btullis: install_server: Generate the H750 stat host reuse recipe at install time [puppet] - 10https://gerrit.wikimedia.org/r/1327520 (https://phabricator.wikimedia.org/T434562) [11:53:48] RECOVERY - Check unit status of statograph_post on alert1002 is OK: OK: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [11:54:45] (03CR) 10Hnowlan: [C:04-1] "Should this host be added to hieradata/common/profile/rsyslog/kafka_shipper.yaml here or in another commit?" [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [11:55:13] (03CR) 10Btullis: [C:03+2] install_server: Generate the H750 stat host reuse recipe at install time [puppet] - 10https://gerrit.wikimedia.org/r/1327520 (https://phabricator.wikimedia.org/T434562) (owner: 10Btullis) [11:57:02] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [11:57:39] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1902.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [11:57:39] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:57:40] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1902.eqiad.wmnet [12:00:04] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1200) [12:00:05] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin2002 from tcpircbot config [puppet] - 10https://gerrit.wikimedia.org/r/1327504 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [12:00:21] !log btullis@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host stat1010.eqiad.wmnet with OS bookworm [12:01:25] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db1902.eqiad.wmnet [12:01:26] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [12:03:01] (03CR) 10Elukey: [C:03+2] docker_registry: route /v2/dev/.* to the releng S3 bucket [puppet] - 10https://gerrit.wikimedia.org/r/1327146 (https://phabricator.wikimedia.org/T432829) (owner: 10Elukey) [12:04:09] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host stat1010.eqiad.wmnet with OS bookworm [12:04:19] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host stat1009.eqiad.wmnet with OS bookworm [12:05:46] (03CR) 10Marostegui: "Should work out of the box if connecting through the proxy:" [puppet] - 10https://gerrit.wikimedia.org/r/1327497 (https://phabricator.wikimedia.org/T435442) (owner: 10Aklapper) [12:05:50] (03CR) 10Marostegui: [C:03+1] mariadb: add grants to phstats for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327497 (https://phabricator.wikimedia.org/T435442) (owner: 10Aklapper) [12:06:56] fceratto@cumin1003 makevm (PID 1892482) is awaiting input [12:08:13] !log T413390 running CentralAuth:FixRenamedUserGlobalEditCount --wiki=metawiki --since=20250901000000 --until=20260301000000 --fix [12:08:18] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:08:19] T413390: Renaming a user still doubles their edit count (December 2025) - https://phabricator.wikimedia.org/T413390 [12:08:53] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" [12:08:57] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1902.eqiad.wmnet - fceratto@cumin1003" [12:08:57] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:08:57] !log fceratto@cumin1003 START - Cookbook sre.dns.wipe-cache db1902.eqiad.wmnet on all recursors [12:09:01] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1902.eqiad.wmnet on all recursors [12:09:33] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" [12:09:37] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1902.eqiad.wmnet - fceratto@cumin1003" [12:11:23] RESOLVED: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d4-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [12:12:09] !log fceratto@cumin1003 START - Cookbook sre.hosts.reimage for host db1902.eqiad.wmnet with OS trixie [12:12:40] (03PS1) 10Blake: service-catalog: move mw-pretrain to production. [puppet] - 10https://gerrit.wikimedia.org/r/1327524 (https://phabricator.wikimedia.org/T427668) [12:12:40] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin2002 from sync list for firmware [puppet] - 10https://gerrit.wikimedia.org/r/1327503 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [12:14:03] !log move the Docker Registry's /v2/dev/.* prefix to its dedicated S3 backend - T432829 [12:14:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:14:08] T432829: Move Docker images under the /v2/releng prefix to S3 - https://phabricator.wikimedia.org/T432829 [12:15:40] (03CR) 10Muehlenhoff: [C:03+2] docker::builder:Deploy files/directories independent of the systemd timer [puppet] - 10https://gerrit.wikimedia.org/r/1327513 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [12:16:03] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations, 13Patch-For-Review: Move Docker images under the /v2/releng prefix to S3 - https://phabricator.wikimedia.org/T432829#12236574 (10elukey) The /v2/dev/.* prefix has been moved as well! I think we should be good, I'll wait a couple of days for any iss... [12:17:47] (03CR) 10Tiziano Fogli: "I think so, thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [12:21:08] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage [12:22:34] !log fceratto@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage [12:22:35] James_F: o/ I just moved the /dev/.* images in the Docker Registry to a separate S3 bucket. In theory nothing should change from the client's point of view, but I haven't copied $everything (since we never deleted any image from swift). I noticed that you have been using those a lot, if you see any docker pull issues let me know! :) [12:24:04] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1010.eqiad.wmnet with reason: host reimage [12:25:10] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [12:25:35] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage [12:25:56] (03PS1) 10Muehlenhoff: planet: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1327529 (https://phabricator.wikimedia.org/T429175) [12:26:56] (03CR) 10Elukey: "Nice work! I added some thoughts, lemme know!" [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [12:27:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:28:24] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1902.eqiad.wmnet with reason: host reimage [12:28:53] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" [12:29:26] (03PS1) 10Cathal Mooney: Remove INCLUDE for IPv6 reverse entries for drmrs<->eqiad old cct [dns] - 10https://gerrit.wikimedia.org/r/1327531 [12:29:31] (03PS1) 10Muehlenhoff: urldownloaders: Switch the legacy CNAMEs to the bullseye VMs for now [dns] - 10https://gerrit.wikimedia.org/r/1327532 (https://phabricator.wikimedia.org/T429175) [12:29:47] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327529 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [12:29:59] (03CR) 10CI reject: [V:04-1] urldownloaders: Switch the legacy CNAMEs to the bullseye VMs for now [dns] - 10https://gerrit.wikimedia.org/r/1327532 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [12:31:06] (03CR) 10Cathal Mooney: [C:03+2] Remove INCLUDE for IPv6 reverse entries for drmrs<->eqiad old cct [dns] - 10https://gerrit.wikimedia.org/r/1327531 (owner: 10Cathal Mooney) [12:31:36] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad cct - cmooney@cumin1003" [12:31:36] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:31:47] !log cmooney@dns3003 START - running authdns-update [12:32:34] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1009.eqiad.wmnet with reason: host reimage [12:32:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:33:11] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet [12:33:14] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:34:10] !log cmooney@dns3003 END - running authdns-update [12:35:31] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet [12:35:38] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet [12:35:44] (03PS1) 10Clément Goubert: mw-on-k8s: Use statsd exporter for fatal-error.php [puppet] - 10https://gerrit.wikimedia.org/r/1327527 (https://phabricator.wikimedia.org/T435364) [12:37:21] (03PS1) 10AOkoth: m3: add phstats permissions for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327535 (https://phabricator.wikimedia.org/T435442) [12:37:25] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:37:58] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet [12:38:04] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet [12:38:14] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:38:23] (03Abandoned) 10AOkoth: m3: add phstats permissions for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327535 (https://phabricator.wikimedia.org/T435442) (owner: 10AOkoth) [12:40:25] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet [12:41:23] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet [12:42:29] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host db1902.eqiad.wmnet with OS trixie [12:42:29] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=1) for new host db1902.eqiad.wmnet [12:43:14] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:43:47] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet [12:43:59] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet [12:44:43] (03CR) 10Ssingh: [C:03+1] planet: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1327529 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [12:46:23] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet [12:46:32] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet [12:46:35] (03PS2) 10Muehlenhoff: urldownloaders: Switch the legacy CNAMEs to the bullseye VMs for now [dns] - 10https://gerrit.wikimedia.org/r/1327532 (https://phabricator.wikimedia.org/T429175) [12:47:49] FIRING: DiskSpace: Disk space build2001:9100:/ 2.89% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=build2001 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [12:48:57] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet [12:49:18] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [12:50:15] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet [12:52:01] (03CR) 10Ssingh: [C:03+1] urldownloaders: Switch the legacy CNAMEs to the bullseye VMs for now [dns] - 10https://gerrit.wikimedia.org/r/1327532 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [12:53:46] (03CR) 10Ssingh: "Is there a reason we have to switch them to insetup_ferm for the reimage?" [puppet] - 10https://gerrit.wikimedia.org/r/1327491 (https://phabricator.wikimedia.org/T427282) (owner: 10Muehlenhoff) [12:54:02] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet [12:54:12] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet [12:54:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [12:54:47] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" [12:54:48] (03PS1) 10Cathal Mooney: remove include for v6 range used on GTT vlan from drmrs to eqiad [dns] - 10https://gerrit.wikimedia.org/r/1327539 [12:55:48] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove dns entries for IPs formerly used on drmrs<->eqiad GTT vpls - cmooney@cumin1003" [12:55:48] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:56:15] (03CR) 10Muehlenhoff: "While the in-place upgrade to Squid 7.6 went fine, I wasn't sure if there were any further Puppet issues during the installation from a fr" [puppet] - 10https://gerrit.wikimedia.org/r/1327491 (https://phabricator.wikimedia.org/T427282) (owner: 10Muehlenhoff) [12:56:37] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet [12:57:08] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet [12:59:22] (03CR) 10Cathal Mooney: [C:03+2] remove include for v6 range used on GTT vlan from drmrs to eqiad [dns] - 10https://gerrit.wikimedia.org/r/1327539 (owner: 10Cathal Mooney) [12:59:32] !log cmooney@dns3003 START - running authdns-update [12:59:34] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet [12:59:44] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [13:00:04] Lucas_WMDE, urbanecm, and TheresNoTime: OwO what's this, a deployment window?? UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1300). nyaa~ [13:00:05] sadiya_wmde28 and Msz2001: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:10] (03PS1) 10Muehlenhoff: use_linux612_on_bookworm: Bump kernel to 6.12.101 [puppet] - 10https://gerrit.wikimedia.org/r/1327540 [13:00:14] o/ [13:00:17] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet [13:00:30] (03CR) 10Ssingh: [C:03+1] Switch the new URL downloaders to insetup_ferm for the reimage [puppet] - 10https://gerrit.wikimedia.org/r/1327491 (https://phabricator.wikimedia.org/T427282) (owner: 10Muehlenhoff) [13:01:54] o/ [13:01:56] !log cmooney@dns3003 END - running authdns-update [13:01:58] I can deploy [13:02:02] (03CR) 10Muehlenhoff: [C:03+2] urldownloaders: Switch the legacy CNAMEs to the bullseye VMs for now [dns] - 10https://gerrit.wikimedia.org/r/1327532 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:02:08] !log jmm@dns1004 START - running authdns-update [13:03:25] Lucas_WMDE: I am also a deployer, I can deploy my patch myself, but I was waiting not to jump over others in the queue :) (but I don't see Sadiya to be here) [13:03:34] Msz2001: go ahead, I think [13:03:38] hopefully she’ll join soon [13:03:41] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet [13:03:44] Okay, I'll start then [13:03:50] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet [13:03:51] (03CR) 10Btullis: [C:03+1] "Thank you." [puppet] - 10https://gerrit.wikimedia.org/r/1327540 (owner: 10Muehlenhoff) [13:04:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327511 (owner: 10Mszwarc) [13:04:21] !log jmm@dns1004 END - running authdns-update [13:06:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:06:36] (03Merged) 10jenkins-bot: UIC: Fix page:page instead of page:other in instrumentation [extensions/CheckUser] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327511 (owner: 10Mszwarc) [13:06:50] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1327511|UIC: Fix page:page instead of page:other in instrumentation]] [13:07:10] (03PS1) 10Muehlenhoff: Phabricator: Use LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1327541 (https://phabricator.wikimedia.org/T429175) [13:07:15] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet [13:07:44] (03CR) 10CI reject: [V:04-1] Phabricator: Use LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1327541 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:07:53] (03CR) 10Muehlenhoff: [C:03+2] use_linux612_on_bookworm: Bump kernel to 6.12.101 [puppet] - 10https://gerrit.wikimedia.org/r/1327540 (owner: 10Muehlenhoff) [13:08:40] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet [13:08:53] !log mszwarc@deploy1003 mszwarc: Backport for [[gerrit:1327511|UIC: Fix page:page instead of page:other in instrumentation]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:09:33] !log mszwarc@deploy1003 mszwarc: Continuing with deployment [13:09:44] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [13:12:05] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet [13:12:58] (03CR) 10Hnowlan: [C:03+1] "Thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1327527 (https://phabricator.wikimedia.org/T435364) (owner: 10Clément Goubert) [13:13:21] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet [13:13:50] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1327511|UIC: Fix page:page instead of page:other in instrumentation]] (duration: 07m 00s) [13:14:32] Finished deploying [13:14:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [13:14:47] thanks [13:15:31] (03PS2) 10Muehlenhoff: Phabricator: Use LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1327541 (https://phabricator.wikimedia.org/T429175) [13:16:36] (03CR) 10Jelto: [C:03+1] "lgtm once GitLab is behind the CDN" [debs/wmf-laptop] - 10https://gerrit.wikimedia.org/r/1325854 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [13:16:46] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet [13:17:08] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet [13:18:21] (03PS1) 10Majavah: P:wmcs::novaproxy: Make haproxy logrotate more aggresive [puppet] - 10https://gerrit.wikimedia.org/r/1327544 (https://phabricator.wikimedia.org/T429930) [13:19:20] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::novaproxy: Make haproxy logrotate more aggresive [puppet] - 10https://gerrit.wikimedia.org/r/1327544 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:19:50] (03CR) 10Andrew Bogott: [C:03+1] P:wmcs::novaproxy: Make haproxy logrotate more aggresive [puppet] - 10https://gerrit.wikimedia.org/r/1327544 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:20:32] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet [13:20:36] (03CR) 10CI reject: [V:04-1] P:wmcs::novaproxy: Make haproxy logrotate more aggresive [puppet] - 10https://gerrit.wikimedia.org/r/1327544 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:20:39] !log klausman@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet [13:21:19] (03PS2) 10Majavah: P:wmcs::novaproxy: Make haproxy logrotate more aggresive [puppet] - 10https://gerrit.wikimedia.org/r/1327544 (https://phabricator.wikimedia.org/T429930) [13:21:28] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1010.eqiad.wmnet with OS bookworm [13:23:59] !log klausman@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet [13:24:07] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Make haproxy logrotate more aggresive [puppet] - 10https://gerrit.wikimedia.org/r/1327544 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [13:24:25] !log klausman@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker [13:24:27] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet [13:26:29] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1009.eqiad.wmnet with OS bookworm [13:28:32] !log fnegri@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet [13:29:13] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [13:31:15] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [13:32:00] elukey: Thanks! Looks OK for now. And yes, trimming it somehow would be good, if we can work out how. [13:32:35] (03CR) 10FNegri: [C:03+2] "mariadb@s6 is now stopped, merging." [puppet] - 10https://gerrit.wikimedia.org/r/1326830 (https://phabricator.wikimedia.org/T409557) (owner: 10FNegri) [13:33:10] (03CR) 10Aude: [C:04-1] Remove reading list experiment instrumentation (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [13:33:20] PROBLEM - MariaDB Replica IO: s6 on clouddb1025 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [13:33:20] PROBLEM - MariaDB Replica SQL: s6 on clouddb1025 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [13:33:30] PROBLEM - mysqld processes on clouddb1025 is CRITICAL: PROCS CRITICAL: 1 process with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [13:33:52] PROBLEM - MariaDB read only s6 on clouddb1025 is CRITICAL: Could not connect to localhost:3316 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [13:33:52] PROBLEM - MariaDB read only wikireplica-s6 on clouddb1025 is CRITICAL: Could not connect to localhost:3316 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [13:34:35] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet [13:37:10] ^ clouddb1025 is me, I'll downtime it [13:37:16] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327541 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:37:30] RECOVERY - mysqld processes on clouddb1025 is OK: PROCS OK: 1 process with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [13:37:59] (03PS1) 10Cathal Mooney: Add magru HE OSPF ints and remove GTT ints for OSPF adjacency [homer/public] - 10https://gerrit.wikimedia.org/r/1327546 (https://phabricator.wikimedia.org/T424839) [13:38:43] !log fnegri@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on clouddb1025.eqiad.wmnet with reason: Removing s6 from clouddb1025 [13:38:48] (03CR) 10Hnowlan: [C:03+1] centralserver: add utility to delete oldest file based on disk usage (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1321669 (https://phabricator.wikimedia.org/T434690) (owner: 10Cwhite) [13:38:59] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [13:41:00] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [13:41:15] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet [13:41:17] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet [13:41:21] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet [13:45:39] (03CR) 10Ssingh: [C:03+1] hieradata: use pdns v5 cfg flag on dns6002 [puppet] - 10https://gerrit.wikimedia.org/r/1327157 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:49:47] !log fnegri@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1025.eqiad.wmnet [13:50:00] (03CR) 10Muehlenhoff: [C:03+1] "LGTM. When is the migration planned? I can upload a new wmf-laptop deb to apt.wikimedia.org when this is enabled." [debs/wmf-laptop] - 10https://gerrit.wikimedia.org/r/1325854 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [13:50:08] !log fnegri@cumin1003 conftool action : set/weight=100; selector: name=clouddb1025.eqiad.wmnet [13:50:19] PROBLEM - Host wikikube-worker1263 is DOWN: PING CRITICAL - Packet loss = 66%, RTA = 5402.31 ms [13:51:03] (03PS1) 10Majavah: Remove hdfs fuse mounts from clouddumps [puppet] - 10https://gerrit.wikimedia.org/r/1327548 (https://phabricator.wikimedia.org/T434794) [13:51:05] RECOVERY - Host wikikube-worker1263 is UP: PING OK - Packet loss = 0%, RTA = 0.23 ms [13:51:31] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet [13:51:38] (03CR) 10CI reject: [V:04-1] Remove hdfs fuse mounts from clouddumps [puppet] - 10https://gerrit.wikimedia.org/r/1327548 (https://phabricator.wikimedia.org/T434794) (owner: 10Majavah) [13:52:08] (03PS2) 10Majavah: Remove hdfs fuse mounts from clouddumps [puppet] - 10https://gerrit.wikimedia.org/r/1327548 (https://phabricator.wikimedia.org/T434794) [13:52:33] (03PS3) 10Majavah: Remove hdfs fuse mounts from clouddumps [puppet] - 10https://gerrit.wikimedia.org/r/1327548 (https://phabricator.wikimedia.org/T434794) [13:53:23] !log installing apr-util security updates [13:53:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:53:30] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9286/co" [puppet] - 10https://gerrit.wikimedia.org/r/1327548 (https://phabricator.wikimedia.org/T434794) (owner: 10Majavah) [13:55:51] (03CR) 10Filippo Giunchedi: [C:03+1] "LGTM, nice" [puppet] - 10https://gerrit.wikimedia.org/r/1327548 (https://phabricator.wikimedia.org/T434794) (owner: 10Majavah) [13:56:05] !log UTC afternoon backport+config window done [13:56:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:56:19] !log fnegri@cumin1003 START - Cookbook sre.hosts.remove-downtime for clouddb1025.eqiad.wmnet [13:56:20] !log fnegri@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for clouddb1025.eqiad.wmnet [13:56:25] (sadiya_wmde28 needs to register first – I didn’t realize this channel is +r – and will reschedule the config change later) [13:56:45] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [13:56:53] (03PS4) 10CDobbins: hieradata: use cfg for pdns v5 on dns1004 [puppet] - 10https://gerrit.wikimedia.org/r/1327151 (https://phabricator.wikimedia.org/T401832) [13:58:10] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet [13:58:11] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet [13:58:18] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet [13:58:50] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (NOOP 1 CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1327151 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:59:11] (03CR) 10Majavah: [V:03+1 C:03+2] Remove hdfs fuse mounts from clouddumps [puppet] - 10https://gerrit.wikimedia.org/r/1327548 (https://phabricator.wikimedia.org/T434794) (owner: 10Majavah) [14:02:11] cmooney@cumin1003 netbox (PID 2006932) is awaiting input [14:04:13] (03CR) 10Arnaudb: "it happened yesterday and was reverted, I'm still debugging one of the 2 issues that motivated the revert, so: soon I hope!" [debs/wmf-laptop] - 10https://gerrit.wikimedia.org/r/1325854 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [14:05:15] 06SRE, 07SRE-Unowned, 10DNS, 07Kubernetes: 10.67.28.73 reverse DNS showing 2(SERVFAIL) - https://phabricator.wikimedia.org/T428573#12237239 (10JMeybohm) a:03JMeybohm With kubepods enabled we get a `.pod.cluster.local.` response for reverse lookup of pod IPs (with and without service/endpoint): ` jayme@d... [14:05:32] (03CR) 10Muehlenhoff: [C:03+1] "Ok! You can go ahead and merge, I'll deal with the wmf-laptop update when the chage is live" [debs/wmf-laptop] - 10https://gerrit.wikimedia.org/r/1325854 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [14:06:10] !log installing libheif security updates [14:06:12] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:08:26] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1 NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1327154 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [14:08:28] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet [14:09:22] (03CR) 10CDobbins: [C:03+2] hieradata: use pdns v5 cfg flag on dns6002 [puppet] - 10https://gerrit.wikimedia.org/r/1327157 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [14:10:29] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 1 NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1327155 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [14:13:00] !log klausman@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet [14:13:02] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet [14:13:02] !log klausman@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker [14:13:22] FIRING: [3x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:13:31] !log installing util-linux security updates [14:13:34] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:14:15] <_Gerges> Reedy: ping [14:14:27] Hi? [14:14:30] !log bking@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply [14:14:36] !log bking@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply [14:15:10] <_Gerges> @Reedy: I would like to request a license Wikimedia for JetBrains IDEs. It would help me enhance my voluntary contributions to software development within the Wikimedia. [14:18:22] FIRING: [3x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:18:48] (03CR) 10Scott French: [C:03+1] "Thanks, Blake!" [puppet] - 10https://gerrit.wikimedia.org/r/1327524 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:18:50] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499 (10elukey) 03NEW [14:19:31] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" [14:22:35] cmooney@cumin1003 netbox (PID 2006932) is awaiting input [14:22:48] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update entries used on new transport backup eqiad codfw - cmooney@cumin1003" [14:22:48] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:23:17] (03CR) 10Blake: [C:03+2] service-catalog: move mw-pretrain to production. [puppet] - 10https://gerrit.wikimedia.org/r/1327524 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:24:06] FIRING: NetworkDeviceAlarmActive: Alarm active on cr2-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [14:24:19] (03CR) 10Scott French: "Sounds good, and indeed it should be easy to revert if we run into problems upon switching changeprop-jobqueue over. Thanks for the review" [puppet] - 10https://gerrit.wikimedia.org/r/1325935 (https://phabricator.wikimedia.org/T434925) (owner: 10Scott French) [14:24:23] (03PS1) 10Elukey: profile::docker_registry: create the "main" docker distribution instance [puppet] - 10https://gerrit.wikimedia.org/r/1327549 (https://phabricator.wikimedia.org/T435499) [14:24:35] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327549 (https://phabricator.wikimedia.org/T435499) (owner: 10Elukey) [14:25:17] (03CR) 10Bking: [C:03+2] cirrussearch: revert to default scaling governor [puppet] - 10https://gerrit.wikimedia.org/r/1327187 (https://phabricator.wikimedia.org/T435400) (owner: 10Bking) [14:25:50] (03CR) 10Scott French: [C:03+2] hieradata: Tune mw-jobrunner mesh listener before adoption [puppet] - 10https://gerrit.wikimedia.org/r/1325935 (https://phabricator.wikimedia.org/T434925) (owner: 10Scott French) [14:27:16] !log reconfigure eqiad<->codfw bgp settings [14:27:19] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:28:58] (03CR) 10Bking: [C:03+2] "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1327187 (https://phabricator.wikimedia.org/T435400) (owner: 10Bking) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1430) [14:33:21] !log cdobbins@cumin1003 conftool action : set/pooled=yes; selector: name=cp1100.* [14:36:39] FIRING: CoreBGPDown: Core BGP session down between cr1-codfw and cr2-eqiad (208.80.154.197) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [14:37:25] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:41:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-codfw and cr2-eqiad (208.80.154.197) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [14:45:25] ^^^ eh this is due to me / reconfiguration I'm doing ignroe. [14:45:48] (03PS1) 10Dpogorzelski: kserve: 0.20 upstream alignment [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327557 (https://phabricator.wikimedia.org/T433973) [14:47:47] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237508 (10ssingh) [14:48:25] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237509 (10ssingh) Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the maint window of 08:00 UTC. Can you please co... [14:49:08] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db1903.eqiad.wmnet [14:49:10] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [14:51:51] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12237539 (10MoritzMuehlenhoff) [14:52:36] (03CR) 10Ssingh: P:tofurkey enable Tofurkey for MAGRU (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [14:52:51] 06SRE, 06Infrastructure-Foundations, 10netops: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506 (10cmooney) 03NEW p:05Triage→03Medium [14:53:13] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" [14:53:14] (03CR) 10Jgiannelos: "I am not very familiar with this part of the config I will defer to SREs for that." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321523 (owner: 10Mvolz) [14:53:17] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1903.eqiad.wmnet - fceratto@cumin1003" [14:53:17] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:53:17] !log fceratto@cumin1003 START - Cookbook sre.dns.wipe-cache db1903.eqiad.wmnet on all recursors [14:53:20] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1903.eqiad.wmnet on all recursors [14:53:41] (03CR) 10CDanis: [C:03+1] tunnelencabulator: gitlab moved behind the text-lb CDN [debs/wmf-laptop] - 10https://gerrit.wikimedia.org/r/1325854 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [14:53:55] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" [14:53:59] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db1903.eqiad.wmnet - fceratto@cumin1003" [14:54:32] !log fceratto@cumin1003 START - Cookbook sre.hosts.reimage for host db1903.eqiad.wmnet with OS trixie [14:57:15] 06SRE, 06Infrastructure-Foundations, 10netops: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506#12237562 (10cmooney) Hmm.... I removed the config for both FPCs, and requested they go to "offline", however the system alarms have not cleared: ` cmooney@re0.cr2-eqiad> show ch... [15:00:05] andre and brennen: Time to do the Train log triage deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1500). [15:00:26] jouncebot, nah [15:00:47] The "MW New Errors" logstash board is pretty clean [15:01:50] (03PS1) 10PipelineBot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327560 [15:04:18] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237593 (10RobH) >>! In T435406#12237508, @ssingh wrote: > Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the main... [15:05:16] (03PS2) 10Cathal Mooney: Add magru HE OSPF ints and remove GTT ints for OSPF adjacency [homer/public] - 10https://gerrit.wikimedia.org/r/1327546 (https://phabricator.wikimedia.org/T424839) [15:06:38] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237604 (10RobH) [15:07:16] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237605 (10RobH) a:05ssingh→03RobH [15:07:39] !log fceratto@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage [15:07:58] (03PS1) 10Muehlenhoff: Failover irc.wikimedia.org to irc1003 [dns] - 10https://gerrit.wikimedia.org/r/1327561 [15:08:31] (03PS1) 10Elukey: profile::amd_gpu: avoid running amd-smi-gpu-partition every puppet run [puppet] - 10https://gerrit.wikimedia.org/r/1327563 (https://phabricator.wikimedia.org/T420507) [15:12:24] (03CR) 10Dpogorzelski: [C:03+1] profile::amd_gpu: avoid running amd-smi-gpu-partition every puppet run [puppet] - 10https://gerrit.wikimedia.org/r/1327563 (https://phabricator.wikimedia.org/T420507) (owner: 10Elukey) [15:13:13] (03CR) 10Elukey: [C:03+2] profile::amd_gpu: avoid running amd-smi-gpu-partition every puppet run [puppet] - 10https://gerrit.wikimedia.org/r/1327563 (https://phabricator.wikimedia.org/T420507) (owner: 10Elukey) [15:14:28] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1903.eqiad.wmnet with reason: host reimage [15:15:32] (03PS1) 10JMeybohm: Enable CoreDNS kubepods plugin in staging-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327564 (https://phabricator.wikimedia.org/T428573) [15:17:14] !log jayme@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [15:19:31] !log jayme@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [15:19:39] !log jayme@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [15:20:52] (03PS1) 10Hnowlan: tests: add test to verify rsvg use of language codes are as expected [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1327566 (https://phabricator.wikimedia.org/T337139) [15:20:58] (03CR) 10Ladsgroup: [C:03+1] "I will personally strangle this host." [dns] - 10https://gerrit.wikimedia.org/r/1327561 (owner: 10Muehlenhoff) [15:21:48] !log jayme@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [15:21:56] !log jayme@deploy1003 helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'. [15:24:08] !log jayme@deploy1003 helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'. [15:24:16] !log jayme@deploy1003 helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'. [15:26:32] !log jayme@deploy1003 helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'. [15:26:40] !log jayme@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'. [15:28:25] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1903.eqiad.wmnet with OS trixie [15:28:26] !log fceratto@cumin1003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host db1903.eqiad.wmnet [15:28:44] !log jayme@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'. [15:28:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:28:53] !log jayme@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [15:29:15] 06SRE, 06serviceops-deprecated, 10TimedMediaHandler, 05MW-1.47-notes (1.47.0-wmf.16; 2026-08-18): Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419#12237671 (10TheDJ) @ladsgroup @jdforrester-WMF looking at Special:TranscodeStatis... [15:30:29] !log jayme@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [15:30:37] !log jayme@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [15:32:10] !log jayme@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [15:32:17] !log jayme@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'. [15:34:08] (03PS1) 10Jelto: etherpad: add helm chart for etherpad service [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327571 (https://phabricator.wikimedia.org/T435509) [15:34:11] (03PS1) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [15:35:19] !log jayme@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'. [15:35:26] !log jayme@deploy1003 helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'. [15:36:11] (03CR) 10CI reject: [V:04-1] etherpad: add helm chart for etherpad service [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327571 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [15:36:30] (03CR) 10CI reject: [V:04-1] helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [15:37:49] !log jayme@deploy1003 helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'. [15:38:39] (03PS2) 10Jelto: etherpad: add helm chart for etherpad service [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327571 (https://phabricator.wikimedia.org/T435509) [15:39:41] (03CR) 10Cwhite: [C:04-1] "statsd should no longer be used in WMF production. statsd format should never be sent to MW's prometheus-statsd-exporter." [puppet] - 10https://gerrit.wikimedia.org/r/1327527 (https://phabricator.wikimedia.org/T435364) (owner: 10Clément Goubert) [15:41:32] (03PS2) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [15:41:49] (03CR) 10AOkoth: [C:03+2] mariadb: add grants to phstats for phab1005 [puppet] - 10https://gerrit.wikimedia.org/r/1327497 (https://phabricator.wikimedia.org/T435442) (owner: 10Aklapper) [15:43:08] (03CR) 10JMeybohm: [V:03+2 C:03+2] Enable CoreDNS kubepods plugin in staging-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327564 (https://phabricator.wikimedia.org/T428573) (owner: 10JMeybohm) [15:44:01] (03CR) 10CI reject: [V:04-1] helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [15:44:18] PROBLEM - Recursive DNS on 2a02:ec80:700:1:195:200:68:4 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [15:45:16] RECOVERY - Recursive DNS on 2a02:ec80:700:1:195:200:68:4 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [15:46:52] (03CR) 10Lerickson: WIP: Enable posting to eventgate from the proxy. (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326893 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [15:47:00] (03PS4) 10Lerickson: Enable posting to eventgate from the proxy. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326893 (https://phabricator.wikimedia.org/T433375) [15:52:10] (03CR) 10Gmodena: [C:03+1] Enable posting to eventgate from the proxy. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326893 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [15:53:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:55:27] !log jayme@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [15:55:43] (03CR) 10Gmodena: [C:03+1] Add egress to the service mesh. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327186 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [15:56:28] !log jayme@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [15:59:21] !log amastilovic@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply [16:00:05] jhathaway and rzl: OwO what's this, a deployment window?? Puppet request window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1600). nyaa~ [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:00:23] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12237989 (10Bethany) Thank you @JMeybohm. Can you please update it to the one provided in this task? Thanks so much. [16:00:33] !log amastilovic@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply [16:05:07] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host stat1011.eqiad.wmnet with OS bookworm [16:05:26] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host stat1011 [16:05:35] !log bking@cumin2003 START - Cookbook sre.dns.netbox [16:09:28] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" [16:09:43] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host stat1011 - bking@cumin2003" [16:09:43] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:09:44] !log bking@cumin2003 START - Cookbook sre.dns.wipe-cache stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [16:09:47] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) stat1011.eqiad.wmnet 14.36.64.10.in-addr.arpa 4.1.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [16:09:48] !log bking@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host stat1011 [16:10:20] !log bking@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host stat1011 [16:10:20] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host stat1011 [16:14:51] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=dns6002.* [reason: depooling for trixie upgrade] [16:15:46] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host dns6002.wikimedia.org with OS trixie [16:19:18] PROBLEM - BFD status on asw1-b13-drmrs.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [16:21:10] FIRING: [2x] BFDdown: BFD session down between asw1-b13-drmrs and 185.15.58.37 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=asw1-b13-drmrs:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:24:24] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12238154 (10JMeybohm) 05Resolved→03Open [16:31:58] 06SRE, 06serviceops-deprecated, 10TimedMediaHandler, 05MW-1.47-notes (1.47.0-wmf.16; 2026-08-18): Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419#12238181 (10Ladsgroup) > Invalid maximum framerate value: '30.000300003' https:/... [16:32:01] (03CR) 10Scott French: "Thanks for the review!" [puppet] - 10https://gerrit.wikimedia.org/r/1304189 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [16:34:07] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on dns6002.wikimedia.org with reason: host reimage [16:36:25] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db2901.codfw.wmnet [16:36:43] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage [16:38:46] (03PS1) 10Federico Ceratto: site.pp: Set MariaDB role for db190[1-3] db290[1-2] [puppet] - 10https://gerrit.wikimedia.org/r/1327586 (https://phabricator.wikimedia.org/T435059) [16:39:48] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns6002.wikimedia.org with reason: host reimage [16:43:38] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on stat1011.eqiad.wmnet with reason: host reimage [16:43:54] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [16:45:24] (03PS1) 10CDobbins: geo-maps: update US mapping [dns] - 10https://gerrit.wikimedia.org/r/1327587 (https://phabricator.wikimedia.org/T402512) [16:47:49] FIRING: DiskSpace: Disk space build2001:9100:/ 2.89% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=build2001 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [16:49:21] fceratto@cumin1003 decommission (PID 2122137) is awaiting input [16:51:45] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db2901.codfw.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [16:51:54] !log disable-puppet on A:cp for ATS Lua change - T427666 [16:51:59] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:52:00] T427666: Route testwiki traffic to the Pretrain MVP environment - https://phabricator.wikimedia.org/T427666 [16:52:56] jouncebot: nowandnext [16:52:57] For the next 0 hour(s) and 7 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1600) [16:52:57] In 0 hour(s) and 7 minute(s): Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1700) [16:52:57] In 0 hour(s) and 7 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1700) [16:53:00] (03CR) 10Scott French: [C:03+2] trafficserver: Support testwiki pretrain routing in XWD [puppet] - 10https://gerrit.wikimedia.org/r/1304189 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [16:53:43] (03PS1) 10Ladsgroup: Make sure fpsmax is an int value [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327590 (https://phabricator.wikimedia.org/T318419) [16:54:17] (03CR) 10Ladsgroup: [C:03+2] Make sure fpsmax is an int value [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327590 (https://phabricator.wikimedia.org/T318419) (owner: 10Ladsgroup) [16:54:50] fceratto@cumin1003 decommission (PID 2122137) is awaiting input [16:58:29] PROBLEM - Recursive DNS on 2a02:ec80:600:2:185:15:58:37 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [16:58:29] PROBLEM - Recursive DNS on 185.15.58.37 is CRITICAL: DNS_QUERY CRITICAL - query timed out https://wikitech.wikimedia.org/wiki/DNS [17:00:05] bd808: May I have your attention please! Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker). (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1700) [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1700) [17:00:11] 06SRE, 06Infrastructure-Foundations, 10netops: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506#12238341 (10cmooney) @VRiley was able to unseat the cards, which means the FPC slots now show as 'empty' rather than 'offline' ` cmooney@re0.cr2-eqiad> show chassis fpc... [17:02:27] RECOVERY - Recursive DNS on 185.15.58.37 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [17:02:27] RECOVERY - Recursive DNS on 2a02:ec80:600:2:185:15:58:37 is OK: DNS_QUERY OK - Success https://wikitech.wikimedia.org/wiki/DNS [17:04:06] RESOLVED: NetworkDeviceAlarmActive: Alarm active on cr2-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [17:04:41] (03PS1) 10Scott French: Revert "trafficserver: Support testwiki pretrain routing in XWD" [puppet] - 10https://gerrit.wikimedia.org/r/1327592 (https://phabricator.wikimedia.org/T427666) [17:05:51] (03CR) 10Blake: [C:03+1] Revert "trafficserver: Support testwiki pretrain routing in XWD" [puppet] - 10https://gerrit.wikimedia.org/r/1327592 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [17:07:24] (03CR) 10Gergő Tisza: [C:03+1] Remove sending email to legal team about rejected requests [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325920 (https://phabricator.wikimedia.org/T374053) (owner: 10Neriah) [17:07:41] (03CR) 10Scott French: [C:03+2] Revert "trafficserver: Support testwiki pretrain routing in XWD" [puppet] - 10https://gerrit.wikimedia.org/r/1327592 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [17:08:28] (03PS2) 10Ladsgroup: Thumbor: Add the wikimedia_thumbor.filter.format filter [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322507 (https://phabricator.wikimedia.org/T430528) (owner: 10Samwilson) [17:09:10] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327590 (https://phabricator.wikimedia.org/T318419) (owner: 10Ladsgroup) [17:10:58] (03PS1) 10BryanDavis: developer-portal: Bump container to 2026-08-20-122835-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327594 (https://phabricator.wikimedia.org/T433106) [17:13:19] RECOVERY - BFD status on asw1-b13-drmrs.mgmt is OK: UP: 6 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [17:15:26] (03Merged) 10jenkins-bot: Make sure fpsmax is an int value [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327590 (https://phabricator.wikimedia.org/T318419) (owner: 10Ladsgroup) [17:15:42] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1327590|Make sure fpsmax is an int value (T318419)]] [17:15:46] T318419: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419 [17:16:03] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [17:16:10] RESOLVED: [2x] BFDdown: BFD session down between asw1-b13-drmrs and 185.15.58.37 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=asw1-b13-drmrs:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:16:19] Amir1: mind if I sneak in a helmfile-only mediawiki change when you're done? [17:16:28] (should be quick, modulo CI) [17:16:33] sure thing! [17:17:38] (03CR) 10BryanDavis: [C:03+2] developer-portal: Bump container to 2026-08-20-122835-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327594 (https://phabricator.wikimedia.org/T433106) (owner: 10BryanDavis) [17:17:45] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1327590|Make sure fpsmax is an int value (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:19:53] (03Merged) 10jenkins-bot: developer-portal: Bump container to 2026-08-20-122835-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327594 (https://phabricator.wikimedia.org/T433106) (owner: 10BryanDavis) [17:20:06] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [17:24:19] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1327590|Make sure fpsmax is an int value (T318419)]] (duration: 08m 37s) [17:24:24] T318419: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419 [17:24:33] swfrench-wmf: the floor is yours! [17:24:42] Amir1: great, thanks! [17:24:53] (03CR) 10Scott French: [C:03+2] mw-pretrain: Relax x-wikimedia-debug match regex [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327164 (https://phabricator.wikimedia.org/T427668) (owner: 10Scott French) [17:26:35] (03PS3) 10Muehlenhoff: Phabricator: Use LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1327541 (https://phabricator.wikimedia.org/T429175) [17:27:24] (03Merged) 10jenkins-bot: mw-pretrain: Relax x-wikimedia-debug match regex [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327164 (https://phabricator.wikimedia.org/T427668) (owner: 10Scott French) [17:27:59] !log bd808@deploy1003 helmfile [staging] START helmfile.d/services/developer-portal: apply [17:29:07] !log swfrench@deploy1003 helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [17:29:22] !log swfrench@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [17:29:23] !log bd808@deploy1003 helmfile [staging] DONE helmfile.d/services/developer-portal: apply [17:29:30] !log bd808@deploy1003 helmfile [codfw] START helmfile.d/services/developer-portal: apply [17:29:49] !log bd808@deploy1003 helmfile [codfw] DONE helmfile.d/services/developer-portal: apply [17:30:25] !log bd808@deploy1003 helmfile [eqiad] START helmfile.d/services/developer-portal: apply [17:30:39] !log swfrench@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply [17:30:46] !log bd808@deploy1003 helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply [17:30:48] !log swfrench@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply [17:31:38] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns6002.wikimedia.org with OS trixie [17:32:22] (03PS7) 10Bernard Wang: Remove reading list experiment instrumentation [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [17:32:38] (03CR) 10Bernard Wang: Remove reading list experiment instrumentation (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [17:32:41] (03PS8) 10Bernard Wang: Remove reading list experiment instrumentation [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [17:34:11] Amir1: all done on my end (in case you have other patches) [17:34:27] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host stat1011.eqiad.wmnet with OS bookworm [17:34:34] thanks. No deploys for now [17:35:04] (03CR) 10Marostegui: "I just realised...is the .yaml not needed?" [puppet] - 10https://gerrit.wikimedia.org/r/1327586 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [17:37:35] (03CR) 10Ssingh: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327541 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [17:40:03] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238568 (10ssingh) @RobH, @cmooney: @SLyngshede-WMF will be handling this from Traffic, just as an FYI. [17:41:50] (03CR) 10Ssingh: [C:03+1] Phabricator: Use LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1327541 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [17:42:18] (03CR) 10Ladsgroup: [C:03+2] Thumbor: Add the wikimedia_thumbor.filter.format filter [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322507 (https://phabricator.wikimedia.org/T430528) (owner: 10Samwilson) [17:44:55] (03PS9) 10Hashar: zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) [17:45:02] (03CR) 10Hashar: zookeeper: fix log4j addition when tls is enabled (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [17:45:05] (03Merged) 10jenkins-bot: Thumbor: Add the wikimedia_thumbor.filter.format filter [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322507 (https://phabricator.wikimedia.org/T430528) (owner: 10Samwilson) [17:47:56] 06SRE, 06serviceops-deprecated, 10TimedMediaHandler, 05MW-1.47-notes (1.47.0-wmf.16; 2026-08-18): Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419#12238646 (10Ladsgroup) > Invalid maximum framerate value: '30' Progress :/ [17:48:38] (03PS1) 10Krinkle: Render overline of \bar with stretchy=false [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327596 (https://phabricator.wikimedia.org/T435456) [17:50:21] (03PS1) 10Krinkle: Make \Omicron non upright (like \Chi) [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327598 (https://phabricator.wikimedia.org/T434428) [17:51:13] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [17:51:23] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [17:52:40] (03PS2) 10Krinkle: Render overline of \bar with stretchy=false [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327596 (https://phabricator.wikimedia.org/T435456) [17:52:44] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 20 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324752 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [17:52:52] 10ops-eqiad, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12238677 (10wiki_willy) a:03VRiley-WMF Adding ops-eqiad tag and assigning to Valerie. ++ @RobH - can you confirm... [17:56:45] !log cdobbins@cumin1003 START - Cookbook sre.hosts.remove-downtime for dns6002.wikimedia.org [17:56:46] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dns6002.wikimedia.org [17:57:23] !log cdobbins@cumin1003 conftool action : set/pooled=yes; selector: name=dns6002.* [reason: depooling for trixie upgrade] [17:58:32] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [17:59:04] !log sukhe@dns1004 START - running authdns-update [17:59:20] 10ops-eqiad, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12238697 (10RobH) >>! In T435537#12238677, @wiki_willy wrote: > Adding ops-eqiad tag and assigning to Valerie. > >... [18:00:04] andre and brennen: MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1800). Please do the needful. [18:01:13] !log sukhe@dns1004 END - running authdns-update [18:01:58] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [18:04:41] o/ nothing for this window. [18:05:19] (03CR) 10Ottomata: "Nice!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326332 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [18:05:23] (03PS2) 10Lerickson: Add egress to the service mesh. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327186 (https://phabricator.wikimedia.org/T433375) [18:05:24] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [18:06:23] (03CR) 10Ottomata: "FWIW, we will need a similar chart for upcoming reference count (AKA article features count) work for the Attribution API." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326332 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [18:07:30] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [18:10:53] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238743 (10cmooney) >>! In T435406#12237508, @ssingh wrote: > Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the m... [18:11:55] (03CR) 10Lerickson: [C:03+2] Add egress to the service mesh. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327186 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [18:14:45] (03Merged) 10jenkins-bot: Add egress to the service mesh. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327186 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [18:18:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 13.43% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:18:22] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:19:14] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 20 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325920 (https://phabricator.wikimedia.org/T374053) (owner: 10Neriah) [18:22:25] (03PS1) 10Ladsgroup: Avoid casting fpxmax to string [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327614 (https://phabricator.wikimedia.org/T318419) [18:22:38] jouncebot: nowandnext [18:22:38] For the next 1 hour(s) and 37 minute(s): MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T1800) [18:22:38] In 1 hour(s) and 37 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T2000) [18:22:52] (03CR) 10Ladsgroup: [C:03+2] Avoid casting fpxmax to string [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327614 (https://phabricator.wikimedia.org/T318419) (owner: 10Ladsgroup) [18:26:19] Amir1: can I go after you? [18:28:10] sure [18:28:27] RECOVERY - OSPF status on cr1-magru is OK: OSPFv2: 4/4 UP : OSPFv3: 4/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [18:28:28] (03CR) 10Krinkle: "The script already supports dogstatsd. There is a dogstatsd_host option." [puppet] - 10https://gerrit.wikimedia.org/r/1327527 (https://phabricator.wikimedia.org/T435364) (owner: 10Clément Goubert) [18:30:13] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 10/10 UP : OSPFv3: 10/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [18:35:17] (03Merged) 10jenkins-bot: Avoid casting fpxmax to string [extensions/TimedMediaHandler] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327614 (https://phabricator.wikimedia.org/T318419) (owner: 10Ladsgroup) [18:36:57] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [18:37:09] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [18:37:25] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:37:45] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [18:37:54] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [18:39:07] (03CR) 10Lerickson: [C:03+2] Enable posting to eventgate from the proxy. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326893 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [18:41:33] (03Merged) 10jenkins-bot: Enable posting to eventgate from the proxy. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326893 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [18:41:34] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238851 (10ssingh) >>! In T435406#12238743, @cmooney wrote: >>>! In T435406#12237508, @ssingh wrote: >> Thanks, @RobH. We will then aim to de... [18:43:14] (03CR) 10BCornwall: [C:03+1] geo-maps: update US mapping [dns] - 10https://gerrit.wikimedia.org/r/1327587 (https://phabricator.wikimedia.org/T402512) (owner: 10CDobbins) [18:43:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:43:31] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1327614|Avoid casting fpxmax to string (T318419)]] [18:43:37] T318419: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419 [18:44:38] (03CR) 10Ssingh: "@bblack@wikimedia.org, @cdanis@wikimedia.org: Just as an FYI, this is the first pull from Airflow to the puppetservers and data is generat" [dns] - 10https://gerrit.wikimedia.org/r/1327587 (https://phabricator.wikimedia.org/T402512) (owner: 10CDobbins) [18:45:37] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1327614|Avoid casting fpxmax to string (T318419)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:46:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 17.8% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:46:18] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [18:46:58] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [18:49:54] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 06Traffic, 13Patch-For-Review: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12238881 (10ssingh) 05Open→03Resolved a:03ssingh urldownloaders are now behind LVS as a low-traffic servic... [18:50:26] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [18:50:59] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1327614|Avoid casting fpxmax to string (T318419)]] (duration: 07m 28s) [18:51:05] T318419: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419 [18:52:06] Krinkle: I'm done :D [18:52:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327598 (https://phabricator.wikimedia.org/T434428) (owner: 10Krinkle) [18:52:22] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327596 (https://phabricator.wikimedia.org/T435456) (owner: 10Krinkle) [18:52:24] Amir1: ack [18:59:44] 06SRE, 06Data-Engineering, 10Observability-Logging, 10Wikimedia-Logstash, and 2 others: Produce ECS formatted logstash logs to Event Platform, allowing them to be queried in the WMF Data Lake with SQL - https://phabricator.wikimedia.org/T291645#12238908 (10Ottomata) > We now have a means for any applic... [19:05:12] (03Merged) 10jenkins-bot: Make \Omicron non upright (like \Chi) [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327598 (https://phabricator.wikimedia.org/T434428) (owner: 10Krinkle) [19:05:14] (03Merged) 10jenkins-bot: Render overline of \bar with stretchy=false [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327596 (https://phabricator.wikimedia.org/T435456) (owner: 10Krinkle) [19:05:31] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1327598|Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596|Render overline of \bar with stretchy=false (T435456)]] [19:05:38] T434428: \Omega rendered in italic rather than roman in Clientside SVG and MathML modes - https://phabricator.wikimedia.org/T434428 [19:05:38] T435456: Render overline of \bar with stretchy=false - https://phabricator.wikimedia.org/T435456 [19:07:34] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1327598|Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596|Render overline of \bar with stretchy=false (T435456)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [19:10:13] 06SRE, 06serviceops-deprecated, 10TimedMediaHandler, 05MW-1.47-notes (1.47.0-wmf.16; 2026-08-18), 13Patch-For-Review: Upgrade Wikimedia production's ffmpeg to 4.4 or later so we can use the fpsmax flag - https://phabricator.wikimedia.org/T318419#12238946 (10Ladsgroup) 05Open→03Resolved a:03Ladsg... [19:16:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:17:38] (03CR) 10DLynch: "I'm not familiar enough with the needs of other services to say whether it'd be useful to push this particular one in that direction -- I " [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326332 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [19:19:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:20:07] !log krinkle@deploy1003 krinkle: Continuing with deployment [19:21:10] (03PS1) 10Scott French: trafficserver: Support testwiki pretrain routing in XWD [puppet] - 10https://gerrit.wikimedia.org/r/1327621 (https://phabricator.wikimedia.org/T427666) [19:21:10] (03CR) 10Scott French: "As explained in the commit message, this takes a more relaxed approach - i.e., although the backends associated with a given scope are sti" [puppet] - 10https://gerrit.wikimedia.org/r/1327621 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [19:24:15] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 21.48% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:24:25] 10ops-eqiad, 06SRE, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Q1:rack/setup/install conf101[0-2] - https://phabricator.wikimedia.org/T435426#12238981 (10Scott_French) @jasmine_ let me know if you have any questions. This will look roughly identical to what we've done for `conf200[789]`. [19:24:25] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1327598|Make \Omicron non upright (like \Chi) (T434428)]], [[gerrit:1327596|Render overline of \bar with stretchy=false (T435456)]] (duration: 18m 54s) [19:24:32] T434428: \Omega rendered in italic rather than roman in Clientside SVG and MathML modes - https://phabricator.wikimedia.org/T434428 [19:24:32] T435456: Render overline of \bar with stretchy=false - https://phabricator.wikimedia.org/T435456 [19:27:34] FIRING: [2x] DiskSpace: Disk space build2001:9100:/ 2.886% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [19:29:15] RESOLVED: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 20.09% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:30:01] (03CR) 10Ssingh: [C:03+1] hieradata: use pdns v5 cfg flag on dns5004 [puppet] - 10https://gerrit.wikimedia.org/r/1327156 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [19:35:26] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543 (10cmooney) 03NEW p:05Triage→03High [19:36:56] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12239020 (10VRiley-WMF) I think I figured out what was going on with cloudvirt1054. It doesn't seem to be booting up properly. I'm currently trying to troubleshoot... [19:42:23] (03PS2) 10Cwhite: beta-logs: enable security plugin [puppet] - 10https://gerrit.wikimedia.org/r/1326957 (https://phabricator.wikimedia.org/T350516) [19:44:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 21.19% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:45:00] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 21.19% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:46:10] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module wmflib [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) [19:49:05] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts in module wmflib [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [19:53:20] (03CR) 10Ottomata: "As far as I can tell...this is just naming. But, I _think_ we can probably easily rename the chart later if this turns about to be true. " [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326332 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [19:54:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 17.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:54:41] (03PS2) 10JHathaway: Puppet 8: Replace legacy facts in module wmflib [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) [19:56:33] (03PS1) 10Cwhite: profile: fix file_ensure ternary operator in security_plugin.pp [puppet] - 10https://gerrit.wikimedia.org/r/1327625 (https://phabricator.wikimedia.org/T350516) [19:56:47] (03CR) 10Ryan Kemper: kafka: converge topic config from hieradata (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [19:57:39] (03PS2) 10Cwhite: profile: fix file_ensure ternary operator in security_plugin.pp [puppet] - 10https://gerrit.wikimedia.org/r/1327625 (https://phabricator.wikimedia.org/T350516) [19:59:02] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12239070 (10VRiley-WMF) @jcrespo I was able to look further into this and it does seem that ipmi is enabled, and it does seem the passwords are there as well. I could try to reprovision this... [20:00:04] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:00:04] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor Q:Why did functions stop calling each other? A:They had arguments. Rise for UTC late backport window . (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T2000). [20:00:04] Reedy and tgr: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:39] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12239071 (10VRiley-WMF) Hey @MoritzMuehlenhoff @Dzahn I didn't know if you were able to check the preseed issue I am having with this unit? [20:00:52] (03PS4) 10Reedy: InitialiseSettings: Enable 2FA banners on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324752 (https://phabricator.wikimedia.org/T428103) [20:00:56] (03CR) 10Reedy: [C:03+2] InitialiseSettings: Enable 2FA banners on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324752 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [20:01:23] (03CR) 10Reedy: [C:03+2] Remove sending email to legal team about rejected requests [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325920 (https://phabricator.wikimedia.org/T374053) (owner: 10Neriah) [20:01:51] (03CR) 10Scott French: [C:03+1] profile::docker_registry: create the "main" docker distribution instance [puppet] - 10https://gerrit.wikimedia.org/r/1327549 (https://phabricator.wikimedia.org/T435499) (owner: 10Elukey) [20:02:25] (03CR) 10Cwhite: [C:03+2] profile: fix file_ensure ternary operator in security_plugin.pp [puppet] - 10https://gerrit.wikimedia.org/r/1327625 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:03:15] (03Merged) 10jenkins-bot: InitialiseSettings: Enable 2FA banners on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324752 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [20:03:31] (03Merged) 10jenkins-bot: Remove sending email to legal team about rejected requests [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325920 (https://phabricator.wikimedia.org/T374053) (owner: 10Neriah) [20:04:17] (03CR) 10Bernard Wang: Remove reading list experiment instrumentation (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [20:07:07] (03PS1) 10CDanis: Bug fixes and UX fixes [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1327626 [20:07:37] (03CR) 10CDanis: [V:03+2 C:03+2] Bug fixes and UX fixes [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1327626 (owner: 10CDanis) [20:08:15] !log cdanis@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" [20:08:17] !log cdanis@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 [20:09:07] !log cdanis@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: bug fixes & UX fixes - cdanis@cumin1003 [20:09:09] !log cdanis@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "bug fixes & UX fixes - cdanis@cumin1003" [20:10:22] (03PS1) 10Cwhite: profile: ensure security dir is a directory [puppet] - 10https://gerrit.wikimedia.org/r/1327627 (https://phabricator.wikimedia.org/T350516) [20:10:34] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [20:11:19] PROBLEM - Squid on install1005 is CRITICAL: connect to address 208.80.154.134 and port 8080: Connection refused https://wikitech.wikimedia.org/wiki/HTTP_proxy [20:13:19] (03CR) 10Cwhite: [C:03+2] profile: ensure security dir is a directory [puppet] - 10https://gerrit.wikimedia.org/r/1327627 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:13:22] FIRING: [5x] ProbeDown: Ripe Atlas anchor atlas1001:80 is not returning HTTP 200 OK on port 80 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:14:11] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [20:17:25] FIRING: [4x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:21:45] distractions [20:23:29] !log reedy@deploy1003 Started scap sync-world: Backport for [[gerrit:1324752|InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920|Remove sending email to legal team about rejected requests (T374053)]] [20:23:36] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [20:23:36] T374053: Rejected vanish requests are not sending notification to WMF Legal - https://phabricator.wikimedia.org/T374053 [20:24:05] thx Reedy. Not testable in production AFAICT. [20:24:17] !log reedy@deploy1003 sync-world failed: Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.15,1.47.0-wmf.16,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/med [20:24:18] iawiki-multiversion --singleversion-image-basename docker-registry.discovery.wmnet/restricted/mediawiki-singleversion --webserver-image-name docker-registry.discovery.wmnet/restricted/mediawiki-webserver --latest-tag latest --label vnd.wikimedia.builder.name=scap --label vnd.wikimedia.builder.version=4.283.0 --label vnd.wikimedia.scap.stage_dir=/srv/mediawiki-staging --label vnd.wikimedia.scap.build_state_dir=/srv/mediawi [20:24:18] ki-staging/scap/image-build' returned non-zero exit status 1. (scap version: 4.283.0) (duration: 00m 49s) [20:24:20] Yeah, I kinda presumed that might be the case :) [20:24:25] oh [20:27:07] Reedy: From ~reedy/scap-image-build-and-push-log: ` Could not connect to webproxy:8080 (208.80.154.134). - connect (111: Connection refused)` [20:28:44] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239173 (10cmooney) [20:29:27] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239180 (10cmooney) [20:34:42] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239204 (10cmooney) I've sent a mail to the HE noc (cc'd noc@wikimedia) to raise a ticket about this. Let's see what they say. [20:36:29] (03PS1) 10Lerickson: Bump wdqs-proxy 0.4.0 -> 0.5.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327630 (https://phabricator.wikimedia.org/T433935) [20:38:13] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 9/10 UP : OSPFv3: 9/10 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [20:38:27] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 4/5 UP : OSPFv3: 4/5 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [20:39:06] (03PS1) 10Cwhite: beta-logs: disable management of internal_users file [puppet] - 10https://gerrit.wikimedia.org/r/1327631 (https://phabricator.wikimedia.org/T350516) [20:40:10] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [20:40:24] (03CR) 10Cwhite: [C:03+2] beta-logs: disable management of internal_users file [puppet] - 10https://gerrit.wikimedia.org/r/1327631 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [20:41:15] (03PS1) 10Eevans: linked-artifacts: add edit_suggestions_counts config (staging) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327632 (https://phabricator.wikimedia.org/T431499) [20:41:17] (03PS1) 10Eevans: linked-artifacts: add edit_suggestions_counts config (production) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327633 (https://phabricator.wikimedia.org/T431499) [20:41:26] FIRING: RoutinatorRRDPErrors: Routinator RRDP fetching issue in eqiad - https://wikitech.wikimedia.org/wiki/RPKI#RRDP_status - https://grafana.wikimedia.org/d/UwUa77GZk/rpki - https://alerts.wikimedia.org/?q=alertname%3DRoutinatorRRDPErrors [20:47:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:51:19] RECOVERY - Squid on install1005 is OK: TCP OK - 0.000 second response time on 208.80.154.134 port 8080 https://wikitech.wikimedia.org/wiki/HTTP_proxy [20:51:48] !log reedy@deploy1003 Started scap sync-world: Backport for [[gerrit:1324752|InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920|Remove sending email to legal team about rejected requests (T374053)]] [20:51:54] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [20:51:55] T374053: Rejected vanish requests are not sending notification to WMF Legal - https://phabricator.wikimedia.org/T374053 [20:52:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:52:35] FIRING: [2x] DiskSpace: Disk space build2001:9100:/ 2.886% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [20:53:22] FIRING: [5x] ProbeDown: Ripe Atlas anchor atlas1001:80 is not returning HTTP 200 OK on port 80 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:53:58] !log reedy@deploy1003 neriah, reedy: Backport for [[gerrit:1324752|InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920|Remove sending email to legal team about rejected requests (T374053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:54:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 22.18% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:54:43] !log reedy@deploy1003 neriah, reedy: Continuing with deployment [20:54:49] PROBLEM - Check unit status of statograph_post on alert1002 is CRITICAL: CRITICAL: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [20:57:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:58:22] FIRING: [7x] ProbeDown: Ripe Atlas anchor atlas1001:80 is not returning HTTP 200 OK on port 80 - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:59:06] !log reedy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324752|InitialiseSettings: Enable 2FA banners on remaining private wikis (T428103)]], [[gerrit:1325920|Remove sending email to legal team about rejected requests (T374053)]] (duration: 07m 18s) [20:59:13] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [20:59:14] T374053: Rejected vanish requests are not sending notification to WMF Legal - https://phabricator.wikimedia.org/T374053 [20:59:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 24.19% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:00:04] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260820T2100) [21:02:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:02:43] preparing to do a security deploy for one patch [21:04:49] RECOVERY - Check unit status of statograph_post on alert1002 is OK: OK: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [21:07:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:08:48] (03PS1) 10Eevans: cassandra: add linked_artifacts.edit_suggestions_count GRANTS [puppet] - 10https://gerrit.wikimedia.org/r/1327637 (https://phabricator.wikimedia.org/T431499) [21:09:39] preparing to run scap [21:11:00] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 21.63% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:11:26] RESOLVED: RoutinatorRRDPErrors: Routinator RRDP fetching issue in eqiad - https://wikitech.wikimedia.org/wiki/RPKI#RRDP_status - https://grafana.wikimedia.org/d/UwUa77GZk/rpki - https://alerts.wikimedia.org/?q=alertname%3DRoutinatorRRDPErrors [21:13:57] (03CR) 10Jdlrobson: [C:03+1] "When can we backport safely?" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [21:17:39] !log Deployed security fix for T433020 [21:17:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:19:07] scap deploy complete [21:19:34] (03CR) 10JHathaway: [C:03+2] admin::hashuser: fix gid check [puppet] - 10https://gerrit.wikimedia.org/r/1326900 (owner: 10JHathaway) [21:21:00] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 24.02% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:22:21] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations: install1005 running out of disk due squid logs from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555 (10Scott_French) 03NEW [21:22:25] FIRING: [4x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:22:51] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations: install1005 running out of disk due squid logs from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12239451 (10Scott_French) p:05Triage→03High [21:23:15] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in module interface [puppet] - 10https://gerrit.wikimedia.org/r/1326909 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:24:07] 06SRE, 06Infrastructure-Foundations, 10Puppet-Infrastructure, 13Patch-For-Review: Fix remaining scoped legacy fact usage - https://phabricator.wikimedia.org/T435225#12239456 (10jhathaway) [21:26:02] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327219 (https://phabricator.wikimedia.org/T435422) (owner: 10Krinkle) [21:27:09] (03Merged) 10jenkins-bot: RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327219 (https://phabricator.wikimedia.org/T435422) (owner: 10Krinkle) [21:27:20] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1327219|RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] [21:27:25] T435422: 10% of flame graph samples attributed to unknown_unknown - https://phabricator.wikimedia.org/T435422 [21:29:26] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1327219|RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:31:23] !log vriley@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1054 [21:31:29] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations: install1005 running out of disk due squid logs from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12239472 (10Scott_French) [21:31:35] !log vriley@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1054 [21:31:39] !log krinkle@deploy1003 krinkle: Continuing with deployment [21:31:54] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:32:20] (03PS3) 10JHathaway: Puppet 8: Replace legacy facts in module wmflib [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) [21:32:28] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:35:50] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1327219|RunSingleJob: Define MW_ENTRY_POINT for flamegraph sample attribution (T435422)]] (duration: 08m 30s) [21:35:56] T435422: 10% of flame graph samples attributed to unknown_unknown - https://phabricator.wikimedia.org/T435422 [21:36:14] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:36:36] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie [21:36:53] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12239476 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1054.eqiad.wmnet with OS trixie [21:37:07] (03PS1) 10Lerickson: Update the Qlever image to include named subqueries. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327642 (https://phabricator.wikimedia.org/T433880) [21:38:31] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in misc modules [puppet] - 10https://gerrit.wikimedia.org/r/1327643 (https://phabricator.wikimedia.org/T435225) [21:38:50] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327643 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [21:48:55] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations: install1005 running out of disk due squid logs from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12239510 (10Scott_French) In case it's useful for correlation with specific workloads, if I sample just the first 2 minutes of... [21:51:43] !log vriley@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage [21:52:51] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12239511 (10VRiley-WMF) [21:57:15] 06SRE, 06Infrastructure-Foundations, 10netops: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239515 (10cmooney) Ticket ID HE#7216121 [21:58:23] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1054.eqiad.wmnet with reason: host reimage [22:04:24] (03PS1) 10Cwhite: Revert "beta-logs: enable security plugin" [puppet] - 10https://gerrit.wikimedia.org/r/1327647 [22:08:03] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1327624 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [22:08:19] (03PS1) 10Cwhite: beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) [22:08:54] (03CR) 10CI reject: [V:04-1] beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:09:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 22.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:10:58] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12239559 (10VRiley-WMF) Okay, so... After troubleshooting it for a while, I thought it may be the CPU that was acting up. However, after reseating all the compone... [22:11:08] (03PS2) 10Cwhite: beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) [22:11:42] (03CR) 10CI reject: [V:04-1] beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:11:45] (03PS3) 10Cwhite: beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) [22:13:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:14:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 22.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:14:35] !log vriley@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" [22:14:57] !log vriley@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003" [22:14:59] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1054.eqiad.wmnet with OS trixie [22:15:13] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12239572 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1054.eqiad.wmnet with OS trixie completed: - clo... [22:15:21] (03PS4) 10Cwhite: beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) [22:16:27] (03PS1) 10Cwhite: logstash: add security-plugin required fields to output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) [22:21:52] (03PS4) 10DLynch: editcheck-headless: Add chart and deployment for the technical pilot [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326332 (https://phabricator.wikimedia.org/T434109) [22:23:16] (03PS5) 10Cwhite: beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) [22:23:16] (03PS2) 10Cwhite: logstash: add security-plugin required fields to output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) [22:26:17] (03CR) 10Cwhite: [C:03+2] Revert "beta-logs: enable security plugin" [puppet] - 10https://gerrit.wikimedia.org/r/1327647 (owner: 10Cwhite) [22:29:13] (03PS1) 10Krinkle: RunSingleJob: Add ProfilingContext::init() [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327652 (https://phabricator.wikimedia.org/T435422) [22:35:08] (03PS5) 10DLynch: editcheck-headless: Add chart and deployment for the technical pilot [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326332 (https://phabricator.wikimedia.org/T434109) [22:45:57] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations: install1005 running out of disk due squid logs from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12239640 (10Scott_French) This started happening again, now hitting commons.wikimedia.org:443 via webproxy. Some lucky tailing... [22:52:06] (03PS6) 10Cwhite: beta-logs: add options required for logstash output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327648 (https://phabricator.wikimedia.org/T350516) [22:52:06] (03PS3) 10Cwhite: logstash: add security-plugin required fields to output plugin [puppet] - 10https://gerrit.wikimedia.org/r/1327650 (https://phabricator.wikimedia.org/T350516) [23:00:23] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [23:09:31] (03PS1) 10Ladsgroup: upload: Drop profile::cache::upload::upload_webp_hits_threshold [puppet] - 10https://gerrit.wikimedia.org/r/1327656 (https://phabricator.wikimedia.org/T431150) [23:09:38] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations: install1005 running out of disk due squid logs from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12239691 (10Scott_French) p:05High→03Medium Dropping to Medium, as the trigger workload has been stopped. @Snwachukwu was... [23:11:40] (03PS1) 10Ladsgroup: cache: Reduce the webp threshold to 80 [puppet] - 10https://gerrit.wikimedia.org/r/1327657 (https://phabricator.wikimedia.org/T431150) [23:23:04] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12239709 (10RobH) 05Open→03Resolved [23:26:46] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12239713 (10RobH) [23:29:14] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327652 (https://phabricator.wikimedia.org/T435422) (owner: 10Krinkle) [23:30:46] (03Merged) 10jenkins-bot: RunSingleJob: Add ProfilingContext::init() [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1327652 (https://phabricator.wikimedia.org/T435422) (owner: 10Krinkle) [23:30:59] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1327652|RunSingleJob: Add ProfilingContext::init() (T435422)]] [23:31:04] T435422: 10% of flame graph samples attributed to unknown_unknown - https://phabricator.wikimedia.org/T435422 [23:33:05] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1327652|RunSingleJob: Add ProfilingContext::init() (T435422)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [23:38:48] !log krinkle@deploy1003 krinkle: Continuing with deployment [23:41:32] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1327661 [23:41:32] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1327661 (owner: 10TrainBranchBot) [23:43:04] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1327652|RunSingleJob: Add ProfilingContext::init() (T435422)]] (duration: 12m 06s) [23:43:09] T435422: 10% of flame graph samples attributed to unknown_unknown - https://phabricator.wikimedia.org/T435422 [23:50:26] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1327661 (owner: 10TrainBranchBot)