[00:04:24] (03PS1) 10DLynch: ve.ui.CodeMirrorAction: load ::highlight() styles only when used [extensions/CodeMirror] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324427 (https://phabricator.wikimedia.org/T434403) [00:04:39] (03PS1) 10DLynch: ve.ui.CodeMirrorAction: load ::highlight() styles only when used [extensions/CodeMirror] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324428 (https://phabricator.wikimedia.org/T434403) [00:05:59] (03CR) 10CI reject: [V:04-1] ve.ui.CodeMirrorAction: load ::highlight() styles only when used [extensions/CodeMirror] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324428 (https://phabricator.wikimedia.org/T434403) (owner: 10DLynch) [00:08:06] (03PS1) 10DLynch: Bump mediawiki/mediawiki-codesniffer to v52.0.0 [extensions/CodeMirror] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324429 [00:10:24] (03PS2) 10DLynch: ve.ui.CodeMirrorAction: load ::highlight() styles only when used [extensions/CodeMirror] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324428 (https://phabricator.wikimedia.org/T434403) [00:13:48] I'm going to do a quick out-of-window backport for a VE performance issue. There doesn't seem to be anything else on the schedule for the next ~6 hours, so I don't think I'm stepping on anyone's toes. [00:15:07] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy1003 using scap backport" [extensions/CodeMirror] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324427 (https://phabricator.wikimedia.org/T434403) (owner: 10DLynch) [00:15:08] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy1003 using scap backport" [extensions/CodeMirror] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324429 (owner: 10DLynch) [00:15:08] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy1003 using scap backport" [extensions/CodeMirror] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324428 (https://phabricator.wikimedia.org/T434403) (owner: 10DLynch) [00:16:45] (03Merged) 10jenkins-bot: ve.ui.CodeMirrorAction: load ::highlight() styles only when used [extensions/CodeMirror] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324427 (https://phabricator.wikimedia.org/T434403) (owner: 10DLynch) [00:16:47] (03Merged) 10jenkins-bot: Bump mediawiki/mediawiki-codesniffer to v52.0.0 [extensions/CodeMirror] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324429 (owner: 10DLynch) [00:16:48] (03Merged) 10jenkins-bot: ve.ui.CodeMirrorAction: load ::highlight() styles only when used [extensions/CodeMirror] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324428 (https://phabricator.wikimedia.org/T434403) (owner: 10DLynch) [00:17:37] !log kemayo@deploy1003 Started scap sync-world: Backport for [[gerrit:1324427|ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429|Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428|ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] [00:17:41] T434403: VE slowed down in Chrome by "Recalculate Style" - https://phabricator.wikimedia.org/T434403 [00:19:39] !log kemayo@deploy1003 kemayo: Backport for [[gerrit:1324427|ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429|Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428|ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [00:21:22] !log kemayo@deploy1003 kemayo: Continuing with deployment [00:25:32] !log kemayo@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324427|ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]], [[gerrit:1324429|Bump mediawiki/mediawiki-codesniffer to v52.0.0]], [[gerrit:1324428|ve.ui.CodeMirrorAction: load ::highlight() styles only when used (T434403)]] (duration: 07m 55s) [00:25:37] T434403: VE slowed down in Chrome by "Recalculate Style" - https://phabricator.wikimedia.org/T434403 [00:30:22] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321224 (https://phabricator.wikimedia.org/T426102) (owner: 10Daimona Eaytoy) [00:35:48] All done. [00:40:27] RECOVERY - Dell PowerEdge or Supermicro Broadcom RAID Controller on an-presto1013 is OK: communication: 0 OK : controller: 0 OK : physical_disk: 0 OK : virtual_disk: 0 OK : bbu: 0 OK : enclosure: 0 OK https://wikitech.wikimedia.org/wiki/PERCCli%23Monitoring [00:52:55] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:12:26] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1324446 [01:12:26] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1324446 (owner: 10TrainBranchBot) [01:13:40] FIRING: SystemdUnitFailed: wmf_auto_restart_rsyslog.service on ml-serve2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:14:44] FIRING: [3x] CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in cloudelastic - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [01:24:13] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1324446 (owner: 10TrainBranchBot) [02:00:36] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:07:21] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 45s) [02:07:55] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:14:44] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:52:40] 06SRE, 10DNS: [Update DNS Record Request] - wikimedia.org - Add TXT verification for Mentimeter - https://phabricator.wikimedia.org/T434620#12206358 (10Peachey88) [02:52:55] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:54:44] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:25:38] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [03:41:13] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-tool1008.eqiad.wmnet with OS bookworm [03:53:03] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage [03:58:24] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-tool1008.eqiad.wmnet with reason: host reimage [04:16:24] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-tool1008.eqiad.wmnet with OS bookworm [04:40:44] !log T434494 reimaged `an-tool1008.eqiad.wmnet` to bookworm; yarn.wikimedia.org is back up [04:40:47] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [04:40:48] T434494: Migrate production hadoop cluster to bookworm - https://phabricator.wikimedia.org/T434494 [05:13:40] FIRING: SystemdUnitFailed: wmf_auto_restart_rsyslog.service on ml-serve2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:14:44] FIRING: [3x] CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in cloudelastic - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [05:16:43] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324530 [05:17:00] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324531 [05:53:50] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5004.eqsin.wmnet [05:55:06] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti5004.eqsin.wmnet [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T0600) [06:03:28] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5004.eqsin.wmnet [06:03:35] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5004.eqsin.wmnet [06:04:21] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324550 [06:06:09] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5005.eqsin.wmnet [06:12:12] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti5005.eqsin.wmnet [06:14:53] (03PS2) 10Trueg: WDQS: Qlever env variables are not all upper-case'd [deployment-charts] - 10https://gerrit.wikimedia.org/r/1323289 (https://phabricator.wikimedia.org/T434321) [06:18:14] (03CR) 10Trueg: [C:03+2] WDQS: Qlever env variables are not all upper-case'd [deployment-charts] - 10https://gerrit.wikimedia.org/r/1323289 (https://phabricator.wikimedia.org/T434321) (owner: 10Trueg) [06:20:18] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5005.eqsin.wmnet [06:20:32] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5005.eqsin.wmnet [06:20:32] (03Merged) 10jenkins-bot: WDQS: Qlever env variables are not all upper-case'd [deployment-charts] - 10https://gerrit.wikimedia.org/r/1323289 (https://phabricator.wikimedia.org/T434321) (owner: 10Trueg) [06:23:01] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5006.eqsin.wmnet [06:27:07] jmm@cumin2003 drain-node (PID 935293) is awaiting input [06:28:03] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti5006.eqsin.wmnet [06:35:24] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1011.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [06:36:22] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5006.eqsin.wmnet [06:36:24] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [06:36:29] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5006.eqsin.wmnet [06:38:37] !log failover ganeti master in eqsin to ganeti5004 [06:38:39] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:41:34] PROBLEM - ganeti-wconfd running on ganeti5007 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 110 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [06:46:12] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12206463 (10jcrespo) [06:48:18] (03PS1) 10Arnaudb: gitlab: alert on backup, restore and replica data staleness [alerts] - 10https://gerrit.wikimedia.org/r/1324322 (https://phabricator.wikimedia.org/T425441) [06:48:28] (03CR) 10Arnaudb: [C:03+2] gitlab: alert on backup, restore and replica data staleness [alerts] - 10https://gerrit.wikimedia.org/r/1324322 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [06:51:25] (03Merged) 10jenkins-bot: gitlab: alert on backup, restore and replica data staleness [alerts] - 10https://gerrit.wikimedia.org/r/1324322 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [06:53:53] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12206469 (10jcrespo) Added to nda LDAP group. [06:54:08] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [06:54:13] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti5007.eqsin.wmnet [06:54:24] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [06:54:44] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:56:08] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [06:56:24] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [06:56:57] (03PS4) 10Jcrespo: admin: Add chlod to production access and deployment rights [puppet] - 10https://gerrit.wikimedia.org/r/1322077 (https://phabricator.wikimedia.org/T433791) (owner: 10Chlod Alejandro) [06:56:57] (03CR) 10Jcrespo: [C:03+1] admin: Add chlod to production access and deployment rights [puppet] - 10https://gerrit.wikimedia.org/r/1322077 (https://phabricator.wikimedia.org/T433791) (owner: 10Chlod Alejandro) [06:57:28] (03CR) 10CI reject: [V:04-1] admin: Add chlod to production access and deployment rights [puppet] - 10https://gerrit.wikimedia.org/r/1322077 (https://phabricator.wikimedia.org/T433791) (owner: 10Chlod Alejandro) [06:58:44] (03CR) 10Jcrespo: [C:03+1] "@Moritz or @Simon, this is ready for review, no further blockers. User was just added to the NDA group." [puppet] - 10https://gerrit.wikimedia.org/r/1322077 (https://phabricator.wikimedia.org/T433791) (owner: 10Chlod Alejandro) [06:58:57] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti5007.eqsin.wmnet [07:00:05] Amir1, urbanecm, and awight: UTC morning backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T0700). Please do the needful. [07:00:05] No Gerrit patches in the queue for this window AFAICS. [07:05:08] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324574 [07:05:10] (03CR) 10Arnaudb: [C:03+2] gitlab: point gitlab-replica-b at the CDN [dns] - 10https://gerrit.wikimedia.org/r/1324275 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [07:05:21] !log arnaudb@dns1006 START - running authdns-update [07:07:08] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti5007.eqsin.wmnet [07:07:15] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti5007.eqsin.wmnet [07:07:21] !log arnaudb@dns1006 END - running authdns-update [07:11:11] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3008.esams.wmnet [07:11:51] (03CR) 10Filippo Giunchedi: [C:03+1] Remove obsolete ms-be.cfg Partman recipe [puppet] - 10https://gerrit.wikimedia.org/r/1324354 (https://phabricator.wikimedia.org/T156955) (owner: 10Muehlenhoff) [07:12:48] (03CR) 10Filippo Giunchedi: [C:03+1] debian: guard against missing systemd package [debs/pint] - 10https://gerrit.wikimedia.org/r/1319039 (owner: 10Hnowlan) [07:12:54] !log btullis@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet [07:13:32] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet [07:13:39] (03PS1) 10Jcrespo: mariadb: Setup db2250 additional s5 instance to replace db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324577 (https://phabricator.wikimedia.org/T434532) [07:13:44] !log btullis@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-master1003.eqiad.wmnet [07:14:14] jmm@cumin2003 drain-node (PID 944859) is awaiting input [07:14:24] !log btullis@cumin1003 START - Cookbook sre.hosts.reboot-single for host an-master1003.eqiad.wmnet [07:15:14] (03PS3) 10Arnaudb: gitlab: point gitlab-replica-a at the CDN [dns] - 10https://gerrit.wikimedia.org/r/1324280 (https://phabricator.wikimedia.org/T425441) [07:15:19] (03PS2) 10Jcrespo: mariadb: Setup db2250 additional s5 instance to replace db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324577 (https://phabricator.wikimedia.org/T434532) [07:15:21] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324577 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [07:15:27] (03PS3) 10Arnaudb: gitlab: point gitlab.wikimedia.org at the CDN [dns] - 10https://gerrit.wikimedia.org/r/1324281 (https://phabricator.wikimedia.org/T425441) [07:15:34] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm [07:16:25] (03PS3) 10Jcrespo: mariadb: Setup db2250 additional s5 instance to replace db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324577 (https://phabricator.wikimedia.org/T434532) [07:16:29] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324577 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [07:16:35] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti3008.esams.wmnet [07:16:45] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm [07:17:40] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1011.eqiad.wmnet with OS bookworm [07:18:52] !log btullis@cumin1003 END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-master1003.eqiad.wmnet [07:18:57] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-master1003.eqiad.wmnet [07:19:22] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm [07:20:33] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324582 [07:22:57] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-web1001.eqiad.wmnet with OS bookworm [07:24:49] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3008.esams.wmnet [07:25:13] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3008.esams.wmnet [07:25:23] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3007.esams.wmnet [07:25:38] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [07:27:34] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti3007.esams.wmnet [07:29:09] FIRING: KubernetesAPILatency: High Kubernetes API latency (GET clusterinformations) on k8s-staging@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s-staging&var-latency_percentile=0.95&var-verb=GET - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [07:31:24] 06SRE, 06DC-Ops, 10decommission-hardware: decommission an-test-master100[1-2] - https://phabricator.wikimedia.org/T433495#12206533 (10BTullis) [07:31:54] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631 (10MoritzMuehlenhoff) 03NEW [07:32:43] (03CR) 10CWilliams: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1324577 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [07:33:43] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage [07:34:09] RESOLVED: KubernetesAPILatency: High Kubernetes API latency (GET clusterinformations) on k8s-staging@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s-staging&var-latency_percentile=0.95&var-verb=GET - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [07:34:51] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1322077 (https://phabricator.wikimedia.org/T433791) (owner: 10Chlod Alejandro) [07:35:47] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3007.esams.wmnet [07:35:53] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3007.esams.wmnet [07:38:02] (03PS1) 10Arnaudb: gitlab: tighten replica data staleness threshold to 30h [alerts] - 10https://gerrit.wikimedia.org/r/1324587 (https://phabricator.wikimedia.org/T425441) [07:38:10] !log btullis@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-master1003.eqiad.wmnet with OS bookworm [07:38:11] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1011.eqiad.wmnet with reason: host reimage [07:39:37] (03CR) 10Arnaudb: [C:03+2] gitlab: tighten replica data staleness threshold to 30h [alerts] - 10https://gerrit.wikimedia.org/r/1324587 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [07:39:59] (03PS1) 10CWilliams: mariadb: Add db1272 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1324588 (https://phabricator.wikimedia.org/T407942) [07:41:46] (03Merged) 10jenkins-bot: gitlab: tighten replica data staleness threshold to 30h [alerts] - 10https://gerrit.wikimedia.org/r/1324587 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [07:45:13] (03PS6) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [07:47:20] (03PS1) 10CWilliams: mariadb: Productionize db1274 [puppet] - 10https://gerrit.wikimedia.org/r/1324593 (https://phabricator.wikimedia.org/T407942) [07:48:08] (03PS7) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [07:50:34] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-master1003.eqiad.wmnet with OS bookworm [07:51:05] (03PS1) 10Filippo Giunchedi: hieradata: enable dumps-nfs.w.o usage in production [puppet] - 10https://gerrit.wikimedia.org/r/1324597 (https://phabricator.wikimedia.org/T432212) [07:54:30] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db2213 to s5 master [puppet] - 10https://gerrit.wikimedia.org/r/1324599 (https://phabricator.wikimedia.org/T434635) [07:57:53] (03PS8) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [07:58:50] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage [07:59:09] 06SRE, 06Infrastructure-Foundations: May 2026 SRE reboots - https://phabricator.wikimedia.org/T426720#12206656 (10LSobanski) 05Open→03In progress [07:59:37] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12206664 (10LSobanski) 05Open→03In progress [08:00:05] brennen and jnuche: Time to snap out of that daydream and deploy MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot). (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T0800). [08:00:16] (03PS1) 10Tiziano Fogli: rsyslog/opensearch: filter out safepoint messages [puppet] - 10https://gerrit.wikimedia.org/r/1324600 (https://phabricator.wikimedia.org/T434502) [08:01:43] (03CR) 10Kosta Harlan: [C:03+1] WikimediaAntiAbuse: Enable personal info for enwiki with no display [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324306 (https://phabricator.wikimedia.org/T431292) (owner: 10Dreamy Jazz) [08:04:10] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-web1001.eqiad.wmnet with reason: host reimage [08:04:56] (03PS9) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [08:07:08] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage [08:07:45] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12206676 (10LSobanski) @MoritzMuehlenhoff, can you suggest a time for a depool and firmware upgrade? [08:08:13] btullis@cumin1003 reimage (PID 2730982) is awaiting input [08:08:16] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12206678 (10LSobanski) 05Open→03In progress [08:09:48] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12206688 (10LSobanski) p:05Triage→03Medium [08:12:11] (03CR) 10MVernon: "Hi," [puppet] - 10https://gerrit.wikimedia.org/r/1324354 (https://phabricator.wikimedia.org/T156955) (owner: 10Muehlenhoff) [08:13:16] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-master1003.eqiad.wmnet with reason: host reimage [08:13:48] (03PS2) 10Filippo Giunchedi: hieradata: enable dumps-nfs.w.o usage in production [puppet] - 10https://gerrit.wikimedia.org/r/1324597 (https://phabricator.wikimedia.org/T432212) [08:15:10] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops, 06Machine-Learning-Team (Q1 FY2026-27): eqiad row A&B host migration details request for Machine Learning - https://phabricator.wikimedia.org/T432649#12206713 (10Gehel) Waiting for the DC switch to start working on it. [08:15:42] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops, 06Machine-Learning-Team (Q1 FY2026-27): eqiad row A&B host migration details request for Machine Learning - https://phabricator.wikimedia.org/T432649#12206714 (10Gehel) Spreadsheet has been updated [08:19:24] 06SRE, 06collaboration-services, 06Traffic-Icebox, 10Wikimedia-Planet: mixed-content issues on planet.wikimedia.org - https://phabricator.wikimedia.org/T141480#12206729 (10LSobanski) 05Open→03Declined We won't be working on this. [08:19:39] (03CR) 10Filippo Giunchedi: "PCC https://puppet-compiler.wmflabs.org/output/1324597/9204/" [puppet] - 10https://gerrit.wikimedia.org/r/1324597 (https://phabricator.wikimedia.org/T432212) (owner: 10Filippo Giunchedi) [08:20:10] (03PS10) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [08:25:16] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1011.eqiad.wmnet with OS bookworm [08:25:53] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1012.eqiad.wmnet with OS bookworm [08:27:02] (03CR) 10Filippo Giunchedi: "Context for this change is enabling use of dumps-nfs.w.o: i.e. move to a single nfs read only mount for dumps. Failover then happens on th" [puppet] - 10https://gerrit.wikimedia.org/r/1324597 (https://phabricator.wikimedia.org/T432212) (owner: 10Filippo Giunchedi) [08:27:25] (03CR) 10Marostegui: [C:03+1] mariadb: Add db1272 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1324588 (https://phabricator.wikimedia.org/T407942) (owner: 10CWilliams) [08:27:49] (03CR) 10Marostegui: [C:03+1] mariadb: Productionize db1274 [puppet] - 10https://gerrit.wikimedia.org/r/1324593 (https://phabricator.wikimedia.org/T407942) (owner: 10CWilliams) [08:27:53] (03PS11) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [08:28:53] !log marostegui@cumin1003 dbctl commit (dc=all): 'Depool ms3 T434288', diff saved to https://phabricator.wikimedia.org/P95993 and previous config saved to /var/cache/conftool/dbconfig/20260812-082852-marostegui.json [08:28:58] T434288: Switchover ms1, ms2 and ms3 master - https://phabricator.wikimedia.org/T434288 [08:29:32] (03PS1) 10Marostegui: db1268: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1324604 (https://phabricator.wikimedia.org/T434288) [08:29:57] (03PS12) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [08:31:15] (03CR) 10Marostegui: [C:03+2] db1268: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1324604 (https://phabricator.wikimedia.org/T434288) (owner: 10Marostegui) [08:32:08] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db2252.codfw.wmnet,db[1153,1268].eqiad.wmnet with reason: Switching over ms3 [08:32:34] (03PS1) 10Marostegui: instances.yaml: Add db1268 [puppet] - 10https://gerrit.wikimedia.org/r/1324605 (https://phabricator.wikimedia.org/T434288) [08:33:34] (03PS13) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [08:33:59] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1268 [puppet] - 10https://gerrit.wikimedia.org/r/1324605 (https://phabricator.wikimedia.org/T434288) (owner: 10Marostegui) [08:35:51] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-web1001.eqiad.wmnet with OS bookworm [08:35:58] (03CR) 10Jcrespo: [C:03+2] mariadb: Setup db2250 additional s5 instance to replace db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324577 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [08:37:23] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1268 to dbctl T434288', diff saved to https://phabricator.wikimedia.org/P95994 and previous config saved to /var/cache/conftool/dbconfig/20260812-083722-marostegui.json [08:37:29] T434288: Switchover ms1, ms2 and ms3 master - https://phabricator.wikimedia.org/T434288 [08:37:35] !log Failover ms3 T434288 [08:37:38] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:38:05] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-master1003.eqiad.wmnet with OS bookworm [08:38:17] !log marostegui@cumin1003 dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95995 and previous config saved to /var/cache/conftool/dbconfig/20260812-083816-marostegui.json [08:38:19] (03PS4) 10Blake: mw-pretrain: Add a jobrunner-canary values.yaml. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321588 (https://phabricator.wikimedia.org/T427668) [08:38:31] (03CR) 10Blake: mw-pretrain: Add a jobrunner-canary values.yaml. (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321588 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [08:39:58] (03CR) 10Jcrespo: [C:03+2] admin: Add chlod to production access and deployment rights [puppet] - 10https://gerrit.wikimedia.org/r/1322077 (https://phabricator.wikimedia.org/T433791) (owner: 10Chlod Alejandro) [08:40:06] (03PS1) 10Marostegui: db1153: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1324654 (https://phabricator.wikimedia.org/T434288) [08:40:24] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm [08:41:37] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-launcher1003.eqiad.wmnet with OS bookworm [08:42:12] PROBLEM - mysqld processes on db2250 is CRITICAL: PROCS CRITICAL: 1 process with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [08:42:16] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage [08:43:04] jynus: ^ [08:43:04] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm [08:43:30] yeah, it is being setup, notifications are disabled but puppet is ongoing [08:44:06] (03PS2) 10Marostegui: db1153: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1324654 (https://phabricator.wikimedia.org/T434288) [08:44:06] (03PS1) 10Marostegui: installserver: Do not format db1268 [puppet] - 10https://gerrit.wikimedia.org/r/1324664 (https://phabricator.wikimedia.org/T434288) [08:45:08] !log fceratto@cumin1003 START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 [08:45:45] (03CR) 10Marostegui: [C:03+2] db1153: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1324654 (https://phabricator.wikimedia.org/T434288) (owner: 10Marostegui) [08:46:42] (03CR) 10Marostegui: [C:03+2] installserver: Do not format db1268 [puppet] - 10https://gerrit.wikimedia.org/r/1324664 (https://phabricator.wikimedia.org/T434288) (owner: 10Marostegui) [08:50:08] (03PS1) 10Marostegui: db1278: Enable notifications and add it to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1324667 (https://phabricator.wikimedia.org/T407942) [08:50:09] btullis@cumin1003 reimage (PID 2748494) is awaiting input [08:51:16] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1012.eqiad.wmnet with reason: host reimage [08:53:25] (03CR) 10Marostegui: [C:03+2] db1278: Enable notifications and add it to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1324667 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [08:53:44] 07Puppet, 06Collaboration-Services, 10Gerrit, 06Infrastructure-Foundations, 13Patch-For-Review: Change puppet-merge git origin to use gerrit.discovery.wmnet instead of gerrit.wikimedia.org - https://phabricator.wikimedia.org/T420184#12206826 (10LSobanski) a:03ABran-WMF [08:55:23] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1278 to dbctl T434288', diff saved to https://phabricator.wikimedia.org/P95996 and previous config saved to /var/cache/conftool/dbconfig/20260812-085521-marostegui.json [08:55:29] T434288: Switchover ms1, ms2 and ms3 master - https://phabricator.wikimedia.org/T434288 [08:55:56] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1278: Pool in x1 [08:57:58] (03PS14) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [08:58:04] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage [09:03:12] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-launcher1003.eqiad.wmnet with reason: host reimage [09:04:02] fceratto@cumin1003 sanitize-wiki (PID 2750463) is awaiting input [09:04:14] (03PS1) 10Giuseppe Lavagetto: Do not escape special metacharacters in regexes [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1324668 [09:05:15] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Extend NEL headers to sites not fronted by CDN - https://phabricator.wikimedia.org/T303725#12206896 (10LSobanski) No longer applicable for GitLab, untagging Collab. [09:06:52] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Extend NEL headers to sites not fronted by CDN - https://phabricator.wikimedia.org/T303725#12206898 (10LSobanski) [09:07:36] (03PS15) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [09:07:47] 06SRE, 10Infrastructure Security, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 07SecTeam-Processed, and 2 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12206901 (10LSobanski) [09:08:59] (03CR) 10Giuseppe Lavagetto: [V:03+2 C:03+2] Do not escape special metacharacters in regexes [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1324668 (owner: 10Giuseppe Lavagetto) [09:09:52] (03PS16) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [09:10:44] (03PS17) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [09:12:30] (03PS1) 10Marostegui: installserver: Do not reimage db1278 [puppet] - 10https://gerrit.wikimedia.org/r/1324671 [09:13:06] !log oblivian@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" [09:13:08] !log oblivian@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 [09:13:40] FIRING: SystemdUnitFailed: wmf_auto_restart_rsyslog.service on ml-serve2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:13:47] (03PS18) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [09:14:01] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: fix regexp escaping bug - oblivian@cumin1003 [09:14:02] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: fix regexp escaping bug - oblivian@cumin1003" [09:14:44] FIRING: [3x] CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in cloudelastic - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [09:15:33] (03PS1) 10Jcrespo: dbbackups: Pool db2250:s5 to replace db2201:s5 for backups [puppet] - 10https://gerrit.wikimedia.org/r/1324672 (https://phabricator.wikimedia.org/T434532) [09:15:59] (03PS2) 10Jcrespo: dbbackups: Pool db2250:s5 to replace db2201:s5 for backups [puppet] - 10https://gerrit.wikimedia.org/r/1324672 (https://phabricator.wikimedia.org/T434532) [09:16:02] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3006.esams.wmnet [09:16:02] (03CR) 10Jcrespo: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324672 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [09:17:27] (03CR) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [09:17:36] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [09:18:03] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti3006.esams.wmnet [09:18:31] (03CR) 10Marostegui: [C:03+2] installserver: Do not reimage db1278 [puppet] - 10https://gerrit.wikimedia.org/r/1324671 (owner: 10Marostegui) [09:19:15] (03PS2) 10Muehlenhoff: Remove obsolete ms-be.cfg Partman recipe [puppet] - 10https://gerrit.wikimedia.org/r/1324354 (https://phabricator.wikimedia.org/T156955) [09:20:03] (03CR) 10Muehlenhoff: Remove obsolete ms-be.cfg Partman recipe (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324354 (https://phabricator.wikimedia.org/T156955) (owner: 10Muehlenhoff) [09:23:07] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm [09:23:24] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm [09:24:12] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1012.eqiad.wmnet with OS bookworm [09:24:59] (03CR) 10Jcrespo: "last step on setup of replacement" [puppet] - 10https://gerrit.wikimedia.org/r/1324672 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [09:26:09] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti3006.esams.wmnet [09:26:14] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3006.esams.wmnet [09:29:40] !log failover ganeti master in esams to ganeti3008 [09:29:41] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:30:54] (03PS2) 10CWilliams: mariadb: Add db1272 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1324588 (https://phabricator.wikimedia.org/T407942) [09:31:34] (03PS2) 10CWilliams: mariadb: Productionize db1274 [puppet] - 10https://gerrit.wikimedia.org/r/1324593 (https://phabricator.wikimedia.org/T407942) [09:31:38] PROBLEM - ganeti-wconfd running on ganeti3005 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 110 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [09:32:58] 06SRE, 10DNS, 06Traffic: [Update DNS Record Request] - wikimedia.org - Add TXT verification for Mentimeter - https://phabricator.wikimedia.org/T434620#12206949 (10Aklapper) > Hey SRE. @bcampbell: For future reference, please add a corresponding project tag if you'd like to contact SRE, as Phabricator is use... [09:33:00] (03CR) 10MVernon: "Hi," [puppet] - 10https://gerrit.wikimedia.org/r/1324354 (https://phabricator.wikimedia.org/T156955) (owner: 10Muehlenhoff) [09:33:36] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1014.eqiad.wmnet with OS bookworm [09:34:15] !log btullis@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-presto1013.eqiad.wmnet with OS bookworm [09:36:31] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12206952 (10jcrespo) 05In progress→03Resolved Access and deployment rights, toghether with LDAP nda grants have been deployed. It can take up to 30 minutes to propagate to all.... [09:37:26] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm [09:37:32] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet [09:40:38] jmm@cumin2003 drain-node (PID 972104) is awaiting input [09:41:04] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-launcher1003.eqiad.wmnet with OS bookworm [09:41:05] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1278: Pool in x1 [09:43:13] (03CR) 10CWilliams: [C:03+2] mariadb: Productionize db1274 [puppet] - 10https://gerrit.wikimedia.org/r/1324593 (https://phabricator.wikimedia.org/T407942) (owner: 10CWilliams) [09:44:10] FIRING: KubernetesAPILatency: High Kubernetes API latency (GET clusterinformations) on k8s-staging@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s-staging&var-latency_percentile=0.95&var-verb=GET - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [09:44:43] (03PS3) 10CWilliams: mariadb: Add db1272 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1324588 (https://phabricator.wikimedia.org/T407942) [09:45:59] btullis@cumin1003 reimage (PID 2792594) is awaiting input [09:48:02] (03CR) 10CWilliams: [C:03+2] mariadb: Add db1272 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1324588 (https://phabricator.wikimedia.org/T407942) (owner: 10CWilliams) [09:49:09] RESOLVED: KubernetesAPILatency: High Kubernetes API latency (GET clusterinformations) on k8s-staging@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s-staging&var-latency_percentile=0.95&var-verb=GET - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [09:49:47] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage [09:50:20] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti3005.esams.wmnet [09:50:37] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm [09:53:16] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db2203 to s1 master [puppet] - 10https://gerrit.wikimedia.org/r/1324677 (https://phabricator.wikimedia.org/T434644) [09:53:30] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1014.eqiad.wmnet with reason: host reimage [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1000) [10:00:16] (03CR) 10Blake: [C:03+2] kubernetes: Add a jobrunner-canary deployment for mw-pretrain. [puppet] - 10https://gerrit.wikimedia.org/r/1324296 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [10:00:59] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: Primary switchover s1 T434644 [10:01:03] T434644: Switchover s1 master (db2212 -> db2203) - https://phabricator.wikimedia.org/T434644 [10:01:35] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Set db2203 with weight 0 T434644', diff saved to https://phabricator.wikimedia.org/P96001 and previous config saved to /var/cache/conftool/dbconfig/20260812-100134-cwilliams.json [10:02:18] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1013.eqiad.wmnet with OS bookworm [10:03:56] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1013.eqiad.wmnet with OS bookworm [10:05:34] (03PS2) 10Gerrit maintenance bot: mariadb: Promote db2203 to s1 master [puppet] - 10https://gerrit.wikimedia.org/r/1324677 (https://phabricator.wikimedia.org/T434644) [10:06:21] (03CR) 10CWilliams: [C:03+2] mariadb: Promote db2203 to s1 master [puppet] - 10https://gerrit.wikimedia.org/r/1324677 (https://phabricator.wikimedia.org/T434644) (owner: 10Gerrit maintenance bot) [10:07:00] btullis@cumin1003 reimage (PID 2801475) is awaiting input [10:08:01] !log Starting s1 codfw failover from db2212 to db2203 - T434644 [10:08:05] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:08:05] T434644: Switchover s1 master (db2212 -> db2203) - https://phabricator.wikimedia.org/T434644 [10:08:09] (03CR) 10Abijeet Patro: [C:04-1] "Should be deployed only after the dependent config patch has been deployed on production" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324031 (https://phabricator.wikimedia.org/T429122) (owner: 10Abijeet Patro) [10:08:50] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Promote db2203 to s1 primary T434644', diff saved to https://phabricator.wikimedia.org/P96002 and previous config saved to /var/cache/conftool/dbconfig/20260812-100849-cwilliams.json [10:09:42] !log powercycle ganeti3005 [10:09:44] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:10:23] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:10:48] PROBLEM - Host ganeti3005 is DOWN: PING CRITICAL - Packet loss = 100% [10:10:54] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depool db2212 T434644', diff saved to https://phabricator.wikimedia.org/P96003 and previous config saved to /var/cache/conftool/dbconfig/20260812-101053-cwilliams.json [10:11:53] !log blake@deploy1003 Started scap sync-world: Non-deployment scap run to populate new release values for T427668 [10:11:57] T427668: Turn up the Pretrain MVP environment - https://phabricator.wikimedia.org/T427668 [10:12:26] !log blake@deploy1003 Stopping before sync operations [10:12:53] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [10:13:06] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [10:13:14] PROBLEM - orchestrator resolve cache non-FQDNs on dborch1002 is CRITICAL: CRITICAL: 2 non-FQDN entries in orchestrator resolve cache: https://wikitech.wikimedia.org/wiki/Orchestrator [10:13:16] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12207074 (10jcrespo) 05Open→03Resolved a:05ARamirez_WMF→03Arnoldokoth We waited for a week and we didn't receive confirmation- it is ok, maybe the person is unavail... [10:13:50] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:13:52] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2212.codfw.wmnet with reason: Maintenance [10:15:10] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [10:15:23] RESOLVED: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:18:21] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage [10:20:22] (03CR) 10Clément Goubert: [C:03+1] "Overall change looks good to me, one small nit and coordination with ServiceOps for deployment left :)" [puppet] - 10https://gerrit.wikimedia.org/r/1315141 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [10:22:44] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646 (10MoritzMuehlenhoff) 03NEW [10:22:47] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12207109 (10MoritzMuehlenhoff) p:05Triage→03High [10:23:33] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1013.eqiad.wmnet with reason: host reimage [10:24:52] !log cwilliams@cumin1003 START - Cookbook sre.mysql.depool depool db1159: Clone source for db1274 [10:25:21] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Clone source for db1274 [10:25:33] (03PS21) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) [10:28:41] (03CR) 10CI reject: [V:04-1] cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [10:29:24] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 2 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12207146 (10MoritzMuehlenhoff) Looks good, for eqiad please use A, B, C and for codfw B, C, D [10:29:47] !log cmooney@cumin1003 START - Cookbook sre.network.peering with action 'configure' for AS: 396993 [10:31:01] !log cmooney@cumin1003 END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 396993 [10:33:10] !log cwilliams@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for db1274.eqiad.wmnet [10:33:34] !log cwilliams@cumin1003 END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1274.eqiad.wmnet [10:38:04] !log cwilliams@cumin1003 START - Cookbook sre.mysql.clone of db1159.eqiad.wmnet onto db1274.eqiad.wmnet [10:39:58] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [10:41:47] (03CR) 10Mvolz: [C:03+2] citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322203 (owner: 10PipelineBot) [10:41:52] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1014.eqiad.wmnet with OS bookworm [10:43:11] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1013.eqiad.wmnet with OS bookworm [10:44:19] (03Merged) 10jenkins-bot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322203 (owner: 10PipelineBot) [10:45:18] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 [10:46:50] (03CR) 10Slyngshede: [C:03+2] P:tofurkey Add tofurkey (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1260730 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [10:48:30] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db2212: Security update [10:50:15] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1016.eqiad.wmnet with OS bookworm [10:50:16] 10ops-codfw, 06SRE, 10Data-Persistence-Backup, 10database-backups, and 3 others: db2201 memory failure (was: both mysql instances crashed in the last few days) - https://phabricator.wikimedia.org/T434532#12207195 (10jcrespo) [10:50:23] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12207196 (10Chlod) Thank you, @jcrespo! Can confirm that I'm able to connect to the bastion and deployment hosts. Per https://wikitech.wikimedia.org/wiki/How_to_deploy_code#Deployme... [10:51:04] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm [10:51:17] * _joe_ I'm having problems with my IRC bouncer, but I'm around o/ [10:54:39] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1009.eqiad.wmnet with OS bookworm [10:54:44] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:54:47] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 2 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12207206 (10Clement_Goubert) [10:55:03] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12207207 (10jcrespo) I added you myself to the last group, as well as the WMF-NDA group here on phabricator, but please sync with release engineering team for coordination, on this... [10:55:11] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 2 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12207209 (10Clement_Goubert) 05Open→03In progress a:03Clement_Goubert [11:00:05] mvolz: gettimeofday() says it's time for Services – Citoid / Zotero. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1100) [11:00:38] btullis@cumin1003 reimage (PID 2848640) is awaiting input [11:00:41] !log fceratto@cumin1003 START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis testwiki in section s3 [11:00:47] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Chlod Alejandro - https://phabricator.wikimedia.org/T433791#12207215 (10Chlod) Gotcha, thanks again! [11:02:36] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6003.drmrs.wmnet [11:03:07] (03PS1) 10Marostegui: mariadb: Productionize db1273 [puppet] - 10https://gerrit.wikimedia.org/r/1324688 (https://phabricator.wikimedia.org/T407942) [11:03:57] (03CR) 10Marostegui: [C:03+2] mariadb: Productionize db1273 [puppet] - 10https://gerrit.wikimedia.org/r/1324688 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [11:04:04] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning [11:05:11] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db1159: Pool db1159.eqiad.wmnet in after cloning [11:05:40] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [11:06:00] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [11:07:01] !log marostegui@cumin1003 START - Cookbook sre.mysql.clone of db1158.eqiad.wmnet onto db1273.eqiad.wmnet [11:07:04] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 [11:07:15] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage [11:08:36] jmm@cumin2003 drain-node (PID 988409) is awaiting input [11:08:50] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1158: Depool db1158.eqiad.wmnet to then clone it to db1273.eqiad.wmnet - marostegui@cumin1003 [11:09:08] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1009.eqiad.wmnet with OS bookworm [11:10:34] !log jmm@cumin2003 END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti3005.esams.wmnet [11:10:34] !log jmm@cumin2003 END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti3005.esams.wmnet [11:11:12] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage [11:11:28] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti6003.drmrs.wmnet [11:11:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:12:14] PROBLEM - MariaDB Replica IO: s7 on db1155 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db1158.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db1158.eqiad.wmnet (111 Connection refused) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [11:14:24] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1016.eqiad.wmnet with reason: host reimage [11:16:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:16:50] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/citoid: apply [11:17:19] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/citoid: apply [11:17:26] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1009.eqiad.wmnet with reason: host reimage [11:17:37] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6003.drmrs.wmnet [11:17:42] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6003.drmrs.wmnet [11:18:05] db1155 is expected, I thought I had downtimed it... [11:18:17] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 22 hosts with reason: Cloning [11:18:18] [13:04:04] <+logmsgbot> !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on 20 hosts with reason: Cloning [11:18:18] ^that was the downtime [11:18:21] so not sure why it alerted [11:18:50] FIRING: [2x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:19:28] fceratto@cumin1003 sanitize-wiki (PID 2857327) is awaiting input [11:19:45] (03CR) 10CWilliams: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1324672 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [11:20:18] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/citoid: apply [11:20:48] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/citoid: apply [11:22:54] !log failover ganeti master in drmrs01 to ganeti6003 [11:22:57] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:25:31] PROBLEM - ganeti-wconfd running on ganeti6001 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 110 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [11:25:43] (03CR) 10Jcrespo: [C:03+2] dbbackups: Pool db2250:s5 to replace db2201:s5 for backups [puppet] - 10https://gerrit.wikimedia.org/r/1324672 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [11:25:56] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321917 (owner: 10PipelineBot) [11:26:04] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322199 (owner: 10PipelineBot) [11:26:11] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322201 (owner: 10PipelineBot) [11:26:15] (03PS1) 10Slyngshede: P:tofurkey enable Tofurkey for MAGRU [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) [11:27:17] (03CR) 10Blake: [C:03+2] envoy-future: Update envoy-future to 1.39.0. [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1324312 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [11:27:20] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [11:27:29] (03CR) 10Blake: [V:03+2 C:03+2] envoy-future: Update envoy-future to 1.39.0. [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1324312 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [11:27:46] (03PS1) 10Jcrespo: dbbackups: Reenable notifications for db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324690 (https://phabricator.wikimedia.org/T434532) [11:28:44] (03Abandoned) 10Jcrespo: dbbackups: Reenable notifications for db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324690 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [11:29:00] jouncebot: nowandnext [11:29:00] For the next 0 hour(s) and 30 minute(s): Services – Citoid / Zotero (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1100) [11:29:00] In 1 hour(s) and 30 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1300) [11:31:17] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324306 (https://phabricator.wikimedia.org/T431292) (owner: 10Dreamy Jazz) [11:32:14] (03Merged) 10jenkins-bot: WikimediaAntiAbuse: Enable personal info for enwiki with no display [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324306 (https://phabricator.wikimedia.org/T431292) (owner: 10Dreamy Jazz) [11:32:43] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1324306|WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] [11:32:47] T431292: Tag revisions with private change tag for likely oversightable/revision deletable content - https://phabricator.wikimedia.org/T431292 [11:33:05] Dreamy_Jazz: I'm done if you have something to do [11:33:21] Yeah, currently using scap [11:33:26] Thanks [11:33:29] +1 [11:33:37] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6001.drmrs.wmnet [11:33:41] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2212: Security update [11:34:47] !log dreamyjazz@deploy1003 dreamyjazz: Backport for [[gerrit:1324306|WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:35:23] (03PS2) 10Slyngshede: P:tofurkey enable Tofurkey for MAGRU [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) [11:37:08] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [11:38:24] jmm@cumin2003 drain-node (PID 994141) is awaiting input [11:38:48] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti6001.drmrs.wmnet [11:39:02] !log dreamyjazz@deploy1003 dreamyjazz: Continuing with deployment [11:42:15] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-presto1015.eqiad.wmnet with OS bookworm [11:43:09] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324306|WikimediaAntiAbuse: Enable personal info for enwiki with no display (T431292)]] (duration: 10m 26s) [11:43:13] T431292: Tag revisions with private change tag for likely oversightable/revision deletable content - https://phabricator.wikimedia.org/T431292 [11:45:04] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6001.drmrs.wmnet [11:45:22] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6001.drmrs.wmnet [11:46:42] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6004.drmrs.wmnet [11:47:47] (03Restored) 10Jcrespo: dbbackups: Reenable notifications for db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324690 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [11:48:00] (03CR) 10Muehlenhoff: "Sure,not rush at all. This is just a cleanup I noticed when looking into XFS usage across the fleet. Just drop me a note this is good to m" [puppet] - 10https://gerrit.wikimedia.org/r/1324354 (https://phabricator.wikimedia.org/T156955) (owner: 10Muehlenhoff) [11:48:01] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1016.eqiad.wmnet with OS bookworm [11:48:36] (03PS2) 10Jcrespo: dbbackups: Disable notifications for db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324690 (https://phabricator.wikimedia.org/T434532) [11:49:33] I should be done [11:49:45] jmm@cumin2003 drain-node (PID 995396) is awaiting input [11:50:24] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Pool db1159.eqiad.wmnet in after cloning [11:50:39] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1159.eqiad.wmnet onto db1274.eqiad.wmnet [11:52:55] (03PS1) 10Majavah: P:toolforge::prometheus: Remove unconditional team: wmcs label [puppet] - 10https://gerrit.wikimedia.org/r/1324696 [11:53:26] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1015.eqiad.wmnet with OS bookworm [11:53:42] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti6004.drmrs.wmnet [11:56:22] btullis@cumin1003 reimage (PID 2900163) is awaiting input [11:56:41] (03CR) 10FNegri: [C:03+1] P:toolforge::prometheus: Remove unconditional team: wmcs label [puppet] - 10https://gerrit.wikimedia.org/r/1324696 (owner: 10Majavah) [11:57:26] (03PS4) 10Clément Goubert: rdb-lock: Prepare adding VMs for new cluster [puppet] - 10https://gerrit.wikimedia.org/r/1324692 (https://phabricator.wikimedia.org/T434188) [11:58:22] (03CR) 10Filippo Giunchedi: [C:03+2] prometheus: add wikitech-static metrics [puppet] - 10https://gerrit.wikimedia.org/r/1312959 (https://phabricator.wikimedia.org/T362397) (owner: 10Filippo Giunchedi) [12:00:00] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6004.drmrs.wmnet [12:00:22] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6004.drmrs.wmnet [12:01:18] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1017.eqiad.wmnet with OS bookworm [12:02:25] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1009.eqiad.wmnet with OS bookworm [12:04:26] !log failover ganeti master in drmrs02 to ganeti6004 [12:04:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:06:29] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9218/co" [puppet] - 10https://gerrit.wikimedia.org/r/1324696 (owner: 10Majavah) [12:06:43] PROBLEM - ganeti-wconfd running on ganeti6002 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 110 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [12:06:55] (03CR) 10Majavah: [V:03+1 C:03+2] P:toolforge::prometheus: Remove unconditional team: wmcs label [puppet] - 10https://gerrit.wikimedia.org/r/1324696 (owner: 10Majavah) [12:07:49] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage [12:10:27] 10ops-eqiad, 06SRE, 06DC-Ops: wmcs PXE issues - https://phabricator.wikimedia.org/T433538#12207383 (10fgiunchedi) Ok thank you @Jhancock.wm ! There's ongoing ceph work (e.g. T429387) and once that settles we'll let you know [12:10:48] fceratto@cumin1003 sanitize-wiki (PID 2857327) is awaiting input [12:11:32] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1015.eqiad.wmnet with reason: host reimage [12:18:06] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage [12:19:21] (03CR) 10Jcrespo: [C:03+2] dbbackups: Disable notifications for db2201 [puppet] - 10https://gerrit.wikimedia.org/r/1324690 (https://phabricator.wikimedia.org/T434532) (owner: 10Jcrespo) [12:20:15] RECOVERY - MariaDB Replica IO: s7 on db1155 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [12:21:03] (03PS1) 10Daimona Eaytoy: EventDetailsParticipantsModule: populate cache with non-local users [extensions/CampaignEvents] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324700 (https://phabricator.wikimedia.org/T434597) [12:21:17] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [extensions/CampaignEvents] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324700 (https://phabricator.wikimedia.org/T434597) (owner: 10Daimona Eaytoy) [12:22:45] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1158: Pool db1158.eqiad.wmnet in after cloning [12:24:17] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1017.eqiad.wmnet with reason: host reimage [12:24:46] (03PS3) 10Filippo Giunchedi: team-wmcs: add wikitech alerts [alerts] - 10https://gerrit.wikimedia.org/r/1312960 (https://phabricator.wikimedia.org/T362397) [12:24:56] (03PS1) 10Bartosz Wójtowicz: ml-services: Image bump for revise-tone-task-generator. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324702 (https://phabricator.wikimedia.org/T433319) [12:26:11] (03PS4) 10Clément Goubert: rdb-lock: Create puppet role redis::lock::instance [puppet] - 10https://gerrit.wikimedia.org/r/1324693 (https://phabricator.wikimedia.org/T434188) [12:26:27] (03PS4) 10Clément Goubert: site.pp: Configure rdb-lock instances [puppet] - 10https://gerrit.wikimedia.org/r/1324694 (https://phabricator.wikimedia.org/T434188) [12:28:04] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1324692 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [12:28:31] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack, 13Patch-For-Review: Clarify when spicerack/cumin is going to retry - https://phabricator.wikimedia.org/T433698#12207450 (10fgiunchedi) >>! In T433698#12178581, @Mahveotm wrote: > Hi @fgiunchedi, I’d like to work on this. I'm propose retaining the dela... [12:30:03] (03CR) 10AikoChou: [C:03+1] ml-services: Image bump for revise-tone-task-generator. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324702 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [12:30:38] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Image bump for revise-tone-task-generator. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324702 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [12:31:36] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti6002.drmrs.wmnet [12:32:48] (03Merged) 10jenkins-bot: ml-services: Image bump for revise-tone-task-generator. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324702 (https://phabricator.wikimedia.org/T433319) (owner: 10Bartosz Wójtowicz) [12:33:58] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [12:34:36] !log bwojtowicz@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [12:34:42] jmm@cumin2003 drain-node (PID 1004950) is awaiting input [12:34:48] (03PS4) 10Dreamy Jazz: Register the mediawiki.wikimedia_antiabuse.content_policy_score stream [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321579 (https://phabricator.wikimedia.org/T432848) (owner: 10Mpostoronca) [12:35:09] !log bwojtowicz@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [12:35:14] (03CR) 10Filippo Giunchedi: "Metric is now live and this is ready to go, let me know what you think !" [alerts] - 10https://gerrit.wikimedia.org/r/1312960 (https://phabricator.wikimedia.org/T362397) (owner: 10Filippo Giunchedi) [12:36:22] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' . [12:41:40] 10SRE-SLO: Sloth: Automated report generation broken due to Grafana Image Renderer - https://phabricator.wikimedia.org/T432970#12207520 (10tappof) [12:44:40] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1017.eqiad.wmnet with OS bookworm [12:47:35] (03PS1) 10Kosta Harlan: Backport all changes from wmf/1.47.0-wmf.15 [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324707 [12:50:13] jmm@cumin2003 drain-node (PID 1004950) is awaiting input [12:53:04] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Degraded RAID on an-presto1013 - https://phabricator.wikimedia.org/T433030#12207630 (10VRiley-WMF) 05Open→03Resolved The drive seems to have been fully rebuilt and is now in a healthy state [12:55:16] (03PS1) 10Slyngshede: C:tofurkey make secret binary [puppet] - 10https://gerrit.wikimedia.org/r/1324709 (https://phabricator.wikimedia.org/T427465) [12:55:58] (03CR) 10CI reject: [V:04-1] C:tofurkey make secret binary [puppet] - 10https://gerrit.wikimedia.org/r/1324709 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [12:56:05] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti6002.drmrs.wmnet [12:57:51] (03PS1) 10Atsuko: dag-processor [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324711 (https://phabricator.wikimedia.org/T433388) [12:58:30] (03Abandoned) 10Atsuko: dag-processor [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324711 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [12:59:15] (03PS2) 10Atsuko: dag-processor [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) [12:59:17] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1015.eqiad.wmnet with OS bookworm [12:59:48] (03PS3) 10Atsuko: airflow: dag-processor for Airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: That opportune time for a UTC afternoon backport window deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1300). [13:00:05] Daimona: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:17] I can’t deploy, in a meeting, sorry [13:00:41] (03PS3) 10Federico Ceratto: mysql.multiinstance_reboot_test: switch to mocker [cookbooks] - 10https://gerrit.wikimedia.org/r/1314730 [13:00:48] o/ I'm here but not a deployer [13:01:02] I will do the table creation myself but need a deployer for everything else [13:01:47] (03PS2) 10Slyngshede: C:tofurkey make secret binary [puppet] - 10https://gerrit.wikimedia.org/r/1324709 (https://phabricator.wikimedia.org/T427465) [13:02:22] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti6002.drmrs.wmnet [13:02:27] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti6002.drmrs.wmnet [13:02:28] (03PS3) 10Slyngshede: C:tofurkey make secret binary [puppet] - 10https://gerrit.wikimedia.org/r/1324709 (https://phabricator.wikimedia.org/T355446) [13:03:33] (03CR) 10CI reject: [V:04-1] mysql.multiinstance_reboot_test: switch to mocker [cookbooks] - 10https://gerrit.wikimedia.org/r/1314730 (owner: 10Federico Ceratto) [13:04:53] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324709 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:07:00] Any deployers around? [13:07:43] (03CR) 10Clément Goubert: [C:03+2] rdb-lock: Prepare adding VMs for new cluster [puppet] - 10https://gerrit.wikimedia.org/r/1324692 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [13:07:59] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Pool db1158.eqiad.wmnet in after cloning [13:08:08] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1158.eqiad.wmnet onto db1273.eqiad.wmnet [13:09:36] (03PS4) 10Atsuko: airflow: dag-processor for Airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) [13:09:47] (03PS2) 10Anzx: thwiki: reinstate temporary wiki25 logos [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324713 (https://phabricator.wikimedia.org/T431094) [13:10:12] (03CR) 10Anzx: "thwiki logos were also removed in this patch, restoring it through I4cdcbf368eacfb7fad8ac7695bb26cf4b6ac41ea" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323350 (https://phabricator.wikimedia.org/T415307) (owner: 10Superpes15) [13:10:53] (03CR) 10TChin: [C:03+2] [eventgate] Bump to v1.32.0, make eventgate-analytics auto-refresh stream configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324331 (https://phabricator.wikimedia.org/T430154) (owner: 10TChin) [13:11:00] Daimona: I can help you [13:11:22] RECOVERY - orchestrator resolve cache non-FQDNs on dborch1002 is OK: OK: all orchestrator resolve cache entries are FQDNs https://wikitech.wikimedia.org/wiki/Orchestrator [13:11:25] What do you need first? [13:11:57] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324707 (owner: 10Kosta Harlan) [13:12:08] (03CR) 10Btullis: airflow: dag-processor for Airflow 3 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [13:12:12] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2027.codfw.wmnet [13:12:14] !log cgoubert@cumin2003 START - Cookbook sre.ganeti.makevm for new host rdb-lock1001.eqiad.wmnet [13:12:16] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [13:12:35] kostajh: i have 1 patch that needs deploying, can i add it calendar if you can deploy - https://gerrit.wikimedia.org/r/1324713 [13:13:01] anzx: ok [13:13:11] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324713 (https://phabricator.wikimedia.org/T431094) (owner: 10Anzx) [13:13:29] kostajh: added to calendar [13:13:40] FIRING: SystemdUnitFailed: wmf_auto_restart_rsyslog.service on ml-serve2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:13:54] Thanks kostajh! I think the order in the calendar is good [13:14:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324713 (https://phabricator.wikimedia.org/T431094) (owner: 10Anzx) [13:14:06] (03CR) 10Ssingh: [C:03+1] C:tofurkey make secret binary [puppet] - 10https://gerrit.wikimedia.org/r/1324709 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:14:07] (03Merged) 10jenkins-bot: [eventgate] Bump to v1.32.0, make eventgate-analytics auto-refresh stream configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324331 (https://phabricator.wikimedia.org/T430154) (owner: 10TChin) [13:14:16] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2027.codfw.wmnet [13:14:28] !log tchin@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply [13:14:37] I'll take care of the data migration, so I "only" need the config change and the two backports [13:14:43] Daimona: ok, I'll start with your config patch, then will do https://gerrit.wikimedia.org/r/c/1321224/ and pause for table creation [13:14:44] FIRING: [3x] CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in cloudelastic - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [13:15:04] !log tchin@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply [13:15:12] !log tchin@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply [13:15:31] (03Merged) 10jenkins-bot: thwiki: reinstate temporary wiki25 logos [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324713 (https://phabricator.wikimedia.org/T431094) (owner: 10Anzx) [13:15:40] !log tchin@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply [13:15:51] (03PS5) 10Atsuko: airflow: dag-processor for Airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) [13:15:52] !log kharlan@deploy1003 Started scap sync-world: Backport for [[gerrit:1324713|thwiki: reinstate temporary wiki25 logos (T431094)]] [13:15:56] T431094: thwiki: change to Wikipedia 25 logo - https://phabricator.wikimedia.org/T431094 [13:16:04] (03CR) 10Atsuko: airflow: dag-processor for Airflow 3 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [13:16:04] Thanks! [13:16:41] (03PS6) 10Reedy: InitialiseSettings: Enable 2FA warnings on more private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324338 (https://phabricator.wikimedia.org/T428103) [13:16:50] (03PS6) 10Atsuko: airflow: dag-processor for Airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) [13:17:09] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" [13:17:14] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" [13:17:14] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:17:14] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock1001.eqiad.wmnet on all recursors [13:17:17] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1001.eqiad.wmnet on all recursors [13:17:39] !log tchin@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply [13:17:45] !log cgoubert@cumin2003 START - Cookbook sre.ganeti.makevm for new host rdb-lock2001.codfw.wmnet [13:17:46] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" [13:17:47] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [13:17:51] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1001.eqiad.wmnet - cgoubert@cumin2003" [13:17:57] !log kharlan@deploy1003 anzx, kharlan: Backport for [[gerrit:1324713|thwiki: reinstate temporary wiki25 logos (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:18:03] looking [13:18:13] 06SRE, 10DNS, 06Traffic: [Update DNS Record Request] - wikimedia.org - Add TXT verification for Mentimeter - https://phabricator.wikimedia.org/T434620#12207756 (10ssingh) a:03CDobbins [13:18:38] !log cgoubert@cumin2003 START - Cookbook sre.hosts.reimage for host rdb-lock1001.eqiad.wmnet with OS trixie [13:18:41] !log tchin@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply [13:18:52] kostajh: looks good ok to continue [13:18:54] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12207759 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cgoubert@cumin2003 for host rdb-lock1001.eqiad.wmn... [13:18:58] !log kharlan@deploy1003 anzx, kharlan: Continuing with deployment [13:19:37] !log tchin@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply [13:20:15] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2027.codfw.wmnet [13:20:23] (03PS4) 10Federico Ceratto: mysql.multiinstance_reboot_test: switch to mocker [cookbooks] - 10https://gerrit.wikimedia.org/r/1314730 [13:20:27] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2027.codfw.wmnet [13:20:29] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1018.eqiad.wmnet with OS bookworm [13:20:40] Daimona: you have the DBA approval for this, right? [13:20:42] !log tchin@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply [13:21:11] (03CR) 10Bking: [C:03+1] rsyslog/opensearch: filter out safepoint messages [puppet] - 10https://gerrit.wikimedia.org/r/1324600 (https://phabricator.wikimedia.org/T434502) (owner: 10Tiziano Fogli) [13:21:23] Daimona: also, can I backport https://gerrit.wikimedia.org/r/c/mediawiki/extensions/CampaignEvents/+/1324700 together with the other wmf.15 patch? [13:21:45] (03PS1) 10Dreamy Jazz: Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" [extensions/CheckUser] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324719 (https://phabricator.wikimedia.org/T366938) [13:22:03] (03PS1) 10Dreamy Jazz: Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324720 (https://phabricator.wikimedia.org/T366938) [13:22:59] We had approval for the rename. I do not have specific approval for the rename plan, I suppose DBAs could do it with a RENAME command directly, my plan is to duplicate the table [13:23:05] !log kharlan@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324713|thwiki: reinstate temporary wiki25 logos (T431094)]] (duration: 07m 13s) [13:23:09] T431094: thwiki: change to Wikipedia 25 logo - https://phabricator.wikimedia.org/T431094 [13:23:10] (03CR) 10EarlyWarningBot: "[Failed command](https://integration.wikimedia.org/ci/job/quibble-vendor-mysql-php83/100969/consoleFull): `composer --ansi install --no-pr" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324720 (https://phabricator.wikimedia.org/T366938) (owner: 10Dreamy Jazz) [13:23:14] And yes they can be backported at the same time [13:23:20] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" [13:23:31] (03CR) 10Btullis: [C:03+1] "Really nice." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [13:23:35] btullis@cumin1003 reimage (PID 2962644) is awaiting input [13:23:42] (03CR) 10Atsuko: [C:03+2] airflow: dag-processor for Airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [13:23:43] btullis@cumin1003 reimage (PID 2962738) is awaiting input [13:23:49] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" [13:23:49] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:23:49] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock2001.codfw.wmnet on all recursors [13:23:50] FIRING: [2x] ProbeDown: Service ganeti2027:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:23:52] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2001.codfw.wmnet on all recursors [13:24:00] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1275938 (https://phabricator.wikimedia.org/T424016) (owner: 10Yahya) [13:24:16] (03CR) 10Federico Ceratto: "This should be ready for review." [cookbooks] - 10https://gerrit.wikimedia.org/r/1314730 (owner: 10Federico Ceratto) [13:24:23] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" [13:24:28] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2001.codfw.wmnet - cgoubert@cumin2003" [13:24:36] !log cgoubert@cumin2003 START - Cookbook sre.hosts.reimage for host rdb-lock2001.codfw.wmnet with OS trixie [13:24:46] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12207804 (10jcrespo) @VRiley-WMF I am unable to connect to either the ssh or https management point of this server. I can do it with, e.g. unrelated db1244 with no issue. I don't know if it i... [13:24:49] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12207803 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cgoubert@cumin2003 for host rdb-lock2001.codfw.wmn... [13:24:59] (03Merged) 10jenkins-bot: Enable campaignEvents on bdwikimedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1275938 (https://phabricator.wikimedia.org/T424016) (owner: 10Yahya) [13:25:12] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Managing sanitization for wikis testwiki in section s3 [13:25:19] !log kharlan@deploy1003 Started scap sync-world: Backport for [[gerrit:1275938|Enable campaignEvents on bdwikimedia (T424016)]] [13:25:23] T424016: Enable CampaignEvents extension on bd.wikimedia.org - https://phabricator.wikimedia.org/T424016 [13:26:27] (03CR) 10Ssingh: P:tofurkey enable Tofurkey for MAGRU (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:26:42] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply [13:26:44] (03CR) 10Neriah: "You will need to schedule this patch for deployment following the process at https://wikitech.wikimedia.org/wiki/Backport_windows#How_to_s" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311141 (https://phabricator.wikimedia.org/T432311) (owner: 10NguoiDungKhongDinhDanh) [13:26:53] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply [13:26:55] kostajh: please run for purging logos https://www.irccloud.com/pastebin/P909qrlQ/ [13:27:02] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2028.codfw.wmnet [13:27:07] (03Merged) 10jenkins-bot: airflow: dag-processor for Airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324678 (https://phabricator.wikimedia.org/T433388) (owner: 10Atsuko) [13:27:12] jouncebot: nowandnext [13:27:12] For the next 0 hour(s) and 32 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1300) [13:27:12] In 0 hour(s) and 32 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1400) [13:27:23] Dreamy_Jazz: still deploying things here [13:27:24] !log kharlan@deploy1003 kharlan, yahya: Backport for [[gerrit:1275938|Enable campaignEvents on bdwikimedia (T424016)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:27:39] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [extensions/CheckUser] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324719 (https://phabricator.wikimedia.org/T366938) (owner: 10Dreamy Jazz) [13:27:47] Daimona: the config patch can be verified via mwdebug now [13:28:21] (03CR) 10Superpes15: "Ooops Thanks! don't know how it happened lol" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323350 (https://phabricator.wikimedia.org/T415307) (owner: 10Superpes15) [13:28:32] marostegui / Amir1 are either of you around for a quick consult? [13:28:43] Thank you, verified via https://bd.wikimedia.org/wiki/%E0%A6%AC%E0%A6%BF%E0%A6%B6%E0%A7%87%E0%A6%B7:AllEvents that the extension is installed there [13:28:51] !log kharlan@deploy1003 kharlan, yahya: Continuing with deployment [13:29:08] (03PS1) 10Dreamy Jazz: Bump mediawiki/mediawiki-codesniffer to v52.0.0 [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324722 (https://phabricator.wikimedia.org/T434187) [13:29:24] (03PS2) 10Dreamy Jazz: Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324720 (https://phabricator.wikimedia.org/T366938) [13:29:51] (03CR) 10Superpes15: "This is actually weird... I only edited two lines and Also I run Tox only for the single project..." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323350 (https://phabricator.wikimedia.org/T415307) (owner: 10Superpes15) [13:29:56] anzx: done [13:30:01] t [13:30:10] kostajh: thanks for deploying [13:31:12] jmm@cumin2003 drain-node (PID 1016884) is awaiting input [13:31:30] (03PS1) 10Bking: apifeatureusage: update logstash repo for bookworm compatibility [puppet] - 10https://gerrit.wikimedia.org/r/1324723 (https://phabricator.wikimedia.org/T433890) [13:31:47] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324720 (https://phabricator.wikimedia.org/T366938) (owner: 10Dreamy Jazz) [13:32:30] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324722 (https://phabricator.wikimedia.org/T434187) (owner: 10Dreamy Jazz) [13:32:36] !log cgoubert@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage [13:32:39] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1019.eqiad.wmnet with OS bookworm [13:32:47] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-presto1020.eqiad.wmnet with OS bookworm [13:32:54] !log kharlan@deploy1003 Finished scap sync-world: Backport for [[gerrit:1275938|Enable campaignEvents on bdwikimedia (T424016)]] (duration: 07m 35s) [13:32:55] I can do mine last [13:32:59] T424016: Enable CampaignEvents extension on bd.wikimedia.org - https://phabricator.wikimedia.org/T424016 [13:33:35] (Will be AFK for a few moments, so ping me when it's my turn) [13:33:45] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy1003 using scap backport" [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324707 (owner: 10Kosta Harlan) [13:34:29] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324723 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [13:35:24] (03Merged) 10jenkins-bot: Backport all changes from wmf/1.47.0-wmf.15 [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324707 (owner: 10Kosta Harlan) [13:35:32] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2028.codfw.wmnet [13:35:46] !log kharlan@deploy1003 Started scap sync-world: Backport for [[gerrit:1324707|Backport all changes from wmf/1.47.0-wmf.15]] [13:36:49] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage [13:38:50] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1001.eqiad.wmnet with reason: host reimage [13:41:18] (03CR) 10Bking: [C:03+2] apifeatureusage: update logstash repo for bookworm compatibility [puppet] - 10https://gerrit.wikimedia.org/r/1324723 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [13:41:52] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1018.eqiad.wmnet with reason: host reimage [13:41:53] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2028.codfw.wmnet [13:41:59] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2028.codfw.wmnet [13:43:08] !log cgoubert@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage [13:43:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [13:43:50] FIRING: [2x] ProbeDown: Service ganeti2028:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:44:08] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host apifeatureusage2001.codfw.wmnet with OS bookworm [13:46:05] (03PS1) 10Phuedx: EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324725 (https://phabricator.wikimedia.org/T429898) [13:47:33] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Configuring db1272 for s3 pooling', diff saved to https://phabricator.wikimedia.org/P96021 and previous config saved to /var/cache/conftool/dbconfig/20260812-134732-cwilliams.json [13:48:23] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2029.codfw.wmnet [13:48:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [13:49:26] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2001.codfw.wmnet with reason: host reimage [13:49:27] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage [13:49:55] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage [13:50:49] (03CR) 10Slyngshede: [C:03+2] C:tofurkey make secret binary [puppet] - 10https://gerrit.wikimedia.org/r/1324709 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [13:52:03] (03PS1) 10Jforrester: wikifunctions: Upgrade evaluators from 2026-08-04-215640 to 2026-08-11-201322 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324726 (https://phabricator.wikimedia.org/T432522) [13:52:05] (03PS1) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-08-04-203437 to 2026-08-11-210638 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324727 (https://phabricator.wikimedia.org/T433506) [13:52:21] (03PS1) 10Phuedx: eventgate-analytics-external: Enable testKitchen transform [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324728 (https://phabricator.wikimedia.org/T429898) [13:52:37] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1020.eqiad.wmnet with reason: host reimage [13:52:49] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2029.codfw.wmnet [13:53:18] !log kharlan@deploy1003 kharlan: Backport for [[gerrit:1324707|Backport all changes from wmf/1.47.0-wmf.15]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:53:33] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1001.eqiad.wmnet with OS trixie [13:53:33] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1001.eqiad.wmnet [13:53:50] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12207936 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cgoubert@cumin2003 for host rdb-lock1001.eqiad.wmnet w... [13:54:03] !log btullis@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on archiva1002.wikimedia.org with reason: Upgrading in-place [13:56:29] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-presto1019.eqiad.wmnet with reason: host reimage [13:57:40] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade evaluators from 2026-08-04-215640 to 2026-08-11-201322 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324726 (https://phabricator.wikimedia.org/T432522) (owner: 10Jforrester) [13:57:54] (03PS1) 10CWilliams: mariadb: Enabling notifications for db1272 [puppet] - 10https://gerrit.wikimedia.org/r/1324729 (https://phabricator.wikimedia.org/T407942) [13:58:13] !log cgoubert@cumin2003 START - Cookbook sre.ganeti.makevm for new host rdb-lock1002.eqiad.wmnet [13:58:15] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [13:58:55] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2029.codfw.wmnet [13:59:21] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2029.codfw.wmnet [14:00:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1400) [14:00:13] (03Merged) 10jenkins-bot: wikifunctions: Upgrade evaluators from 2026-08-04-215640 to 2026-08-11-201322 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324726 (https://phabricator.wikimedia.org/T432522) (owner: 10Jforrester) [14:00:59] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2030.codfw.wmnet [14:02:36] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage [14:02:42] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" [14:03:31] I'm still deploying a patch [14:03:37] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2001.codfw.wmnet with OS trixie [14:03:37] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2001.codfw.wmnet [14:03:47] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12207991 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cgoubert@cumin2003 for host rdb-lock2001.codfw.wmnet w... [14:04:07] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:04:32] !log kharlan@deploy1003 kharlan: Continuing with deployment [14:04:58] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:05:03] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" [14:05:03] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:05:03] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock1002.eqiad.wmnet on all recursors [14:05:06] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1002.eqiad.wmnet on all recursors [14:05:35] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" [14:05:39] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1002.eqiad.wmnet - cgoubert@cumin2003" [14:05:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [14:05:53] !log cgoubert@cumin2003 START - Cookbook sre.ganeti.makevm for new host rdb-lock2002.codfw.wmnet [14:05:56] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [14:06:01] !log cgoubert@cumin2003 START - Cookbook sre.hosts.reimage for host rdb-lock1002.eqiad.wmnet with OS trixie [14:06:12] (03PS1) 10Jgiannelos: prv: Enable parsoid rendering for 5 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324731 [14:06:15] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12207994 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cgoubert@cumin2003 for host rdb-lock1002.eqiad.wmn... [14:06:34] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:06:36] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2030.codfw.wmnet [14:06:55] (03CR) 10Ottomata: [C:03+1] EventStreamConfig: Mark product_metrics.web_base and .web_base_with_ip as Test Kitchen streams [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324725 (https://phabricator.wikimedia.org/T429898) (owner: 10Phuedx) [14:07:12] (03PS2) 10Jgiannelos: prv: Enable parsoid rendering for 5 wikisource wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324731 [14:07:57] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie [14:08:11] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12208010 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [14:08:34] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:09:23] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:09:29] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on apifeatureusage2001.codfw.wmnet with reason: host reimage [14:09:44] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-08-04-203437 to 2026-08-11-210638 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324727 (https://phabricator.wikimedia.org/T433506) (owner: 10Jforrester) [14:10:41] (03CR) 10Jforrester: [C:03+2] wikifunctions: Lower ORCHESTRATOR_HEAP_SIZE for staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321658 (https://phabricator.wikimedia.org/T433506) (owner: 10Jforrester) [14:11:30] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:12:02] cgoubert@cumin2003 makevm (PID 1029247) is awaiting input [14:12:10] (03Merged) 10jenkins-bot: wikifunctions: Upgrade orchestrator from 2026-08-04-203437 to 2026-08-11-210638 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324727 (https://phabricator.wikimedia.org/T433506) (owner: 10Jforrester) [14:12:39] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2030.codfw.wmnet [14:12:46] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2030.codfw.wmnet [14:12:52] bking@cumin2003 reimage (PID 1018975) is awaiting input [14:13:07] (03Merged) 10jenkins-bot: wikifunctions: Lower ORCHESTRATOR_HEAP_SIZE for staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321658 (https://phabricator.wikimedia.org/T433506) (owner: 10Jforrester) [14:13:51] FIRING: [2x] ProbeDown: Service ganeti2030:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:14:02] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" [14:14:06] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" [14:14:06] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:14:06] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock2002.codfw.wmnet on all recursors [14:14:09] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2002.codfw.wmnet on all recursors [14:14:14] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:14:36] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2045.codfw.wmnet [14:14:37] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:14:42] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" [14:14:47] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2002.codfw.wmnet - cgoubert@cumin2003" [14:14:52] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:15:25] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [14:15:35] !log cgoubert@cumin2003 START - Cookbook sre.hosts.reimage for host rdb-lock2002.codfw.wmnet with OS trixie [14:15:53] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12208045 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cgoubert@cumin2003 for host rdb-lock2002.codfw.wmn... [14:16:33] !log installing Linux 6.1.180 on Bookworm hosts [14:16:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:16:38] !log kharlan@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324707|Backport all changes from wmf/1.47.0-wmf.15]] (duration: 40m 51s) [14:16:43] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:16:52] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:17:08] bking@cumin2003 reimage (PID 1018975) is awaiting input [14:17:27] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:18:00] !log cgoubert@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage [14:18:52] jouncebot: nowandnext [14:18:52] For the next 0 hour(s) and 41 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1400) [14:18:53] In 0 hour(s) and 11 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1430) [14:19:11] jforrester: Can I use scap? [14:19:17] The backport window ran over [14:19:34] jmm@cumin2003 drain-node (PID 1030524) is awaiting input [14:19:42] Actually they don't seem to be in this channel [14:20:30] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2045.codfw.wmnet [14:20:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [14:21:13] Going to use scap as it seems they are done [14:21:30] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324719 (https://phabricator.wikimedia.org/T366938) (owner: 10Dreamy Jazz) [14:21:31] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324720 (https://phabricator.wikimedia.org/T366938) (owner: 10Dreamy Jazz) [14:21:31] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324722 (https://phabricator.wikimedia.org/T434187) (owner: 10Dreamy Jazz) [14:22:33] !log fceratto@cumin1003 START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis testwiki in section s3 [14:23:06] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1019.eqiad.wmnet with OS bookworm [14:23:29] (03Merged) 10jenkins-bot: Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" [extensions/CheckUser] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324719 (https://phabricator.wikimedia.org/T366938) (owner: 10Dreamy Jazz) [14:23:43] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1002.eqiad.wmnet with reason: host reimage [14:23:48] (03CR) 10Andrew Bogott: [C:03+2] aptrepo: Add k8s v1.33, remove k8s v1.31 [puppet] - 10https://gerrit.wikimedia.org/r/1324360 (https://phabricator.wikimedia.org/T408785) (owner: 10Andrew Bogott) [14:25:39] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2045.codfw.wmnet [14:25:46] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2045.codfw.wmnet [14:26:31] (03PS1) 10Bking: apifeatureusage: explicitly set java version [puppet] - 10https://gerrit.wikimedia.org/r/1324732 (https://phabricator.wikimedia.org/T433890) [14:27:10] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2046.codfw.wmnet [14:29:21] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324732 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [14:29:26] Going to use scap as it seems they are done [14:29:30] jouncebot: nowandnext [14:29:30] For the next 0 hour(s) and 30 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1400) [14:29:30] In 0 hour(s) and 0 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1430) [14:30:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1400) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1430) [14:30:32] (03CR) 10Scott French: [C:03+1] "Thanks, Blake!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321588 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:31:11] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1020.eqiad.wmnet with OS bookworm [14:31:34] jmm@cumin2003 drain-node (PID 1036805) is awaiting input [14:32:24] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet [14:32:28] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-presto1018.eqiad.wmnet with OS bookworm [14:32:29] (03Merged) 10jenkins-bot: Bump mediawiki/mediawiki-codesniffer to v52.0.0 [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324722 (https://phabricator.wikimedia.org/T434187) (owner: 10Dreamy Jazz) [14:32:34] (03Merged) 10jenkins-bot: Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" [extensions/CheckUser] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324720 (https://phabricator.wikimedia.org/T366938) (owner: 10Dreamy Jazz) [14:33:03] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1324719|Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720|Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722|Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] [14:33:06] !log cgoubert@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage [14:33:11] T366938: Reduce relying on database locks - https://phabricator.wikimedia.org/T366938 [14:33:12] T434596: Timeout while waiting for lock in UserAgentClientHintsManager::insertMappingRows - https://phabricator.wikimedia.org/T434596 [14:33:12] T434187: CI runs for repos with codesniffer 51 fail due to PKSA-rdkp-vv9z-mjkg - https://phabricator.wikimedia.org/T434187 [14:34:08] PROBLEM - Host aux-k8s-etcd2003 is DOWN: PING CRITICAL - Packet loss = 100% [14:34:20] PROBLEM - Host logstash2023 is DOWN: PING CRITICAL - Packet loss = 100% [14:34:20] PROBLEM - Host kubestagemaster2005 is DOWN: PING CRITICAL - Packet loss = 100% [14:34:30] PROBLEM - Host dse-k8s-etcd2001 is DOWN: PING CRITICAL - Packet loss = 100% [14:34:57] (03PS1) 10Kosta Harlan: WikimediaAntiAbuse: Enable personal info tag display on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324733 (https://phabricator.wikimedia.org/T431292) [14:35:42] (03PS2) 10Kosta Harlan: WikimediaAntiAbuse: Enable personal info tag display on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324733 (https://phabricator.wikimedia.org/T431292) [14:36:57] !log dreamyjazz@deploy1003 dreamyjazz: Backport for [[gerrit:1324719|Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720|Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722|Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] synced to the testservers (see https://wikitech. [14:36:57] wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:37:48] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1002.eqiad.wmnet with OS trixie [14:37:48] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1002.eqiad.wmnet [14:37:50] !log dreamyjazz@deploy1003 dreamyjazz: Continuing with deployment [14:38:00] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12208157 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cgoubert@cumin2003 for host rdb-lock1002.eqiad.wmnet w... [14:38:02] ^ these are expected, Ganeti reboots [14:38:50] FIRING: [5x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:38:57] FIRING: KubernetesCalicoDown: kubestagemaster2005.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=codfw%20prometheus%2Fk8s-staging&var-instance=kubestagemaster2005.codfw.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [14:39:20] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2002.codfw.wmnet with reason: host reimage [14:39:30] !log cgoubert@cumin2003 START - Cookbook sre.ganeti.makevm for new host rdb-lock1003.eqiad.wmnet [14:39:32] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [14:43:09] (03PS2) 10Bking: apifeatureusage: Use correct java version [puppet] - 10https://gerrit.wikimedia.org/r/1324732 (https://phabricator.wikimedia.org/T433890) [14:43:40] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324732 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [14:44:05] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324719|Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324720|Partial revert "Use LockManager service instead of Database::getScopedLockAndFlush()" (T366938 T434596)]], [[gerrit:1324722|Bump mediawiki/mediawiki-codesniffer to v52.0.0 (T434187)]] (duration: 11m 02s) [14:44:13] T366938: Reduce relying on database locks - https://phabricator.wikimedia.org/T366938 [14:44:13] T434596: Timeout while waiting for lock in UserAgentClientHintsManager::insertMappingRows - https://phabricator.wikimedia.org/T434596 [14:44:13] T434187: CI runs for repos with codesniffer 51 fail due to PKSA-rdkp-vv9z-mjkg - https://phabricator.wikimedia.org/T434187 [14:44:38] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" [14:45:03] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1272.eqiad.wmnet with reason: Enabling notifications [14:45:20] (03CR) 10CWilliams: [C:03+2] mariadb: Enabling notifications for db1272 [puppet] - 10https://gerrit.wikimedia.org/r/1324729 (https://phabricator.wikimedia.org/T407942) (owner: 10CWilliams) [14:45:27] (03PS2) 10CWilliams: mariadb: Enabling notifications for db1272 [puppet] - 10https://gerrit.wikimedia.org/r/1324729 (https://phabricator.wikimedia.org/T407942) [14:45:38] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" [14:45:38] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [14:45:38] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock1003.eqiad.wmnet on all recursors [14:45:41] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock1003.eqiad.wmnet on all recursors [14:46:09] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" [14:46:13] (03PS1) 10Andrew Bogott: apt: remove bullseye routinator [puppet] - 10https://gerrit.wikimedia.org/r/1324734 [14:46:13] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock1003.eqiad.wmnet - cgoubert@cumin2003" [14:46:38] (03CR) 10CWilliams: [C:03+2] mariadb: Enabling notifications for db1272 [puppet] - 10https://gerrit.wikimedia.org/r/1324729 (https://phabricator.wikimedia.org/T407942) (owner: 10CWilliams) [14:47:23] !log cgoubert@cumin2003 START - Cookbook sre.hosts.reimage for host rdb-lock1003.eqiad.wmnet with OS trixie [14:47:31] (03CR) 10Tiziano Fogli: [C:03+2] rsyslog/opensearch: filter out safepoint messages [puppet] - 10https://gerrit.wikimedia.org/r/1324600 (https://phabricator.wikimedia.org/T434502) (owner: 10Tiziano Fogli) [14:47:38] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12208186 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cgoubert@cumin2003 for host rdb-lock1003.eqiad.wmn... [14:51:21] (03CR) 10Blake: [C:03+2] mw-pretrain: Add a jobrunner-canary values.yaml. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321588 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:52:23] PROBLEM - Host syslog.anycast.wmnet is DOWN: PING CRITICAL - Packet loss = 100% [14:52:49] PROBLEM - Host ganeti2046 is DOWN: PING CRITICAL - Packet loss = 100% [14:53:51] FIRING: [6x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [14:54:03] (03Merged) 10jenkins-bot: mw-pretrain: Add a jobrunner-canary values.yaml. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1321588 (https://phabricator.wikimedia.org/T427668) (owner: 10Blake) [14:54:44] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:55:12] !log powercycle ganeti2046 [14:55:16] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:55:35] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [14:56:18] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [14:56:26] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply [14:56:48] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply [14:57:02] (03PS7) 10Reedy: InitialiseSettings: Enable 2FA warnings on more private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324338 (https://phabricator.wikimedia.org/T428103) [14:57:08] (03CR) 10Reedy: [C:03+2] InitialiseSettings: Enable 2FA warnings on more private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324338 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [14:57:19] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2002.codfw.wmnet with OS trixie [14:57:19] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2002.codfw.wmnet [14:57:33] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12208234 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cgoubert@cumin2003 for host rdb-lock2002.codfw.wmnet w... [14:57:37] !log cgoubert@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage [14:57:38] !log cgoubert@cumin2003 START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet [14:57:40] FIRING: KubernetesAPINotScrapable: k8s-aux@codfw is failing to scrape the k8s api - https://phabricator.wikimedia.org/T343529 - TODO - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPINotScrapable [14:57:40] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [14:58:03] (03PS1) 10Tiziano Fogli: rsyslog/opensearch: filter out safepoint messages [puppet] - 10https://gerrit.wikimedia.org/r/1324748 (https://phabricator.wikimedia.org/T434502) [14:59:04] (03Merged) 10jenkins-bot: InitialiseSettings: Enable 2FA warnings on more private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324338 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [15:00:32] (03PS1) 10Muehlenhoff: Fix routinator import config [puppet] - 10https://gerrit.wikimedia.org/r/1324749 [15:00:42] !log reedy@deploy1003 Started scap sync-world: Backport for [[gerrit:1324338|InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] [15:00:47] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [15:01:19] (03CR) 10Tiziano Fogli: [C:03+2] rsyslog/opensearch: filter out safepoint messages [puppet] - 10https://gerrit.wikimedia.org/r/1324748 (https://phabricator.wikimedia.org/T434502) (owner: 10Tiziano Fogli) [15:01:24] (03CR) 10Majavah: [C:03+1] Fix routinator import config [puppet] - 10https://gerrit.wikimedia.org/r/1324749 (owner: 10Muehlenhoff) [15:02:17] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [15:02:55] RESOLVED: KubernetesAPINotScrapable: k8s-aux@codfw is failing to scrape the k8s api - https://phabricator.wikimedia.org/T343529 - TODO - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPINotScrapable [15:02:56] !log reedy@deploy1003 reedy: Backport for [[gerrit:1324338|InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:03:03] (03CR) 10Andrew Bogott: [C:03+1] Fix routinator import config [puppet] - 10https://gerrit.wikimedia.org/r/1324749 (owner: 10Muehlenhoff) [15:03:19] 10ops-codfw, 06SRE, 06DC-Ops: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681 (10MoritzMuehlenhoff) 03NEW [15:03:21] (03Abandoned) 10Andrew Bogott: apt: remove bullseye routinator [puppet] - 10https://gerrit.wikimedia.org/r/1324734 (owner: 10Andrew Bogott) [15:03:28] 10ops-codfw, 06SRE, 06DC-Ops: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12208297 (10MoritzMuehlenhoff) p:05Triage→03High [15:03:40] (03CR) 10Muehlenhoff: [C:03+2] Fix routinator import config [puppet] - 10https://gerrit.wikimedia.org/r/1324749 (owner: 10Muehlenhoff) [15:03:42] !log reedy@deploy1003 reedy: Continuing with deployment [15:03:52] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock1003.eqiad.wmnet with reason: host reimage [15:03:54] (03PS9) 10Tsevener: Point Test Wiki to new docroot [puppet] - 10https://gerrit.wikimedia.org/r/1315141 (https://phabricator.wikimedia.org/T432412) [15:04:03] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [15:04:03] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:04:04] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors [15:04:06] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors [15:04:25] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [15:07:45] !log reedy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324338|InitialiseSettings: Enable 2FA warnings on more private wikis (T428103)]] (duration: 07m 02s) [15:07:48] jouncebot nowandnext [15:07:48] No deployments scheduled for the next 1 hour(s) and 52 minute(s) [15:07:48] In 1 hour(s) and 52 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1700) [15:07:49] RECOVERY - Host syslog.anycast.wmnet is UP: PING OK - Packet loss = 0%, RTA = 0.22 ms [15:07:51] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [15:08:06] going to do a backport for a train blocker. [15:08:11] (03PS10) 10Tsevener: Point Test Wiki to new docroot [puppet] - 10https://gerrit.wikimedia.org/r/1315141 (https://phabricator.wikimedia.org/T432412) [15:09:15] (03CR) 10TrainBranchBot: [C:03+2] "Approved by brennen@deploy1003 using scap backport" [extensions/CampaignEvents] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324700 (https://phabricator.wikimedia.org/T434597) (owner: 10Daimona Eaytoy) [15:10:49] cdobbins@cumin1003 reimage (PID 3002927) is awaiting input [15:11:33] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [15:11:37] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [15:11:37] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:11:38] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors [15:11:41] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors [15:11:53] !log cgoubert@cumin2003 END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet [15:12:02] !log cgoubert@cumin2003 START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet [15:12:04] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [15:13:58] (03CR) 10Bking: [C:03+2] "The PCC failure is expected, as the facts for `apifeatureusage2001` are out of date. Merging..." [puppet] - 10https://gerrit.wikimedia.org/r/1324732 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [15:16:25] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [15:16:30] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [15:16:30] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:16:30] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors [15:16:33] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors [15:16:53] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [15:17:25] (03CR) 10Eevans: [C:03+2] restbase: Set storage compatibility to UPGRADING [puppet] - 10https://gerrit.wikimedia.org/r/1315185 (https://phabricator.wikimedia.org/T433028) (owner: 10Eevans) [15:18:04] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock1003.eqiad.wmnet with OS trixie [15:18:04] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock1003.eqiad.wmnet [15:18:21] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12208352 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cgoubert@cumin2003 for host rdb-lock1003.eqiad.wmnet w... [15:18:23] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm [15:18:49] (03CR) 10Dreamy Jazz: [C:03+1] WikimediaAntiAbuse: Enable personal info tag display on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324733 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [15:19:25] bking@cumin2003 reimage (PID 1018975) is awaiting input [15:20:21] (03Merged) 10jenkins-bot: EventDetailsParticipantsModule: populate cache with non-local users [extensions/CampaignEvents] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324700 (https://phabricator.wikimedia.org/T434597) (owner: 10Daimona Eaytoy) [15:20:39] !log brennen@deploy1003 Started scap sync-world: Backport for [[gerrit:1324700|EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] [15:20:45] T434597: CampaignEvents: RuntimeException: Cache should be set for valid users - https://phabricator.wikimedia.org/T434597 [15:21:30] brennen: Could you ping me when done? I want to deploy a config patch [15:21:48] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [15:21:52] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Remove records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [15:21:53] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:21:53] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors [15:21:56] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors [15:22:00] (03CR) 10Ssingh: gitlab: point gitlab-replica-a at the CDN (031 comment) [dns] - 10https://gerrit.wikimedia.org/r/1324280 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [15:22:08] !log cgoubert@cumin2003 END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host rdb-lock2003.codfw.wmnet [15:22:12] Dreamy_Jazz: will do. [15:22:19] Thanks [15:22:45] !log brennen@deploy1003 brennen, daimona: Backport for [[gerrit:1324700|EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:22:49] (03CR) 10Ssingh: [C:03+1] gitlab: point gitlab.wikimedia.org at the CDN [dns] - 10https://gerrit.wikimedia.org/r/1324281 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [15:23:05] !log jhancock@cumin1003 START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [15:23:09] (03CR) 10BCornwall: [C:03+1] debian: guard against missing systemd package [debs/pint] - 10https://gerrit.wikimedia.org/r/1319039 (owner: 10Hnowlan) [15:23:14] !log brennen@deploy1003 brennen, daimona: Continuing with deployment [15:23:42] !log eevans@cumin1003 START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — T433028 - eevans@cumin1003 [15:23:46] T433028: Upgrade restbase cluster to Cassandra 5.0.8 - https://phabricator.wikimedia.org/T433028 [15:27:18] !log brennen@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324700|EventDetailsParticipantsModule: populate cache with non-local users (T434597)]] (duration: 06m 38s) [15:27:22] T434597: CampaignEvents: RuntimeException: Cache should be set for valid users - https://phabricator.wikimedia.org/T434597 [15:27:40] Dreamy_Jazz: over to you. [15:27:44] Thanks! [15:27:58] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324733 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [15:27:59] (03CR) 10Clément Goubert: [C:03+1] Point Test Wiki to new docroot (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1315141 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [15:28:55] (03PS1) 10Reedy: InitialiseSettings: Enable 2FA enforcement on more private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324751 (https://phabricator.wikimedia.org/T428103) [15:28:59] (03PS1) 10Reedy: InitialiseSettings: Enable 2FA banners on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324752 (https://phabricator.wikimedia.org/T428103) [15:29:00] (03Merged) 10jenkins-bot: WikimediaAntiAbuse: Enable personal info tag display on enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324733 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [15:29:01] (03PS1) 10Reedy: InitialiseSettings: Enable 2FA enforcement on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324753 (https://phabricator.wikimedia.org/T428103) [15:29:23] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1324733|WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] [15:29:28] T431292: Tag revisions with private change tag for likely oversightable/revision deletable content - https://phabricator.wikimedia.org/T431292 [15:29:37] !log mforns@deploy1003 Started deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] [15:29:54] (03PS1) 10Bking: apifeatureusage: explicitly define java package version [puppet] - 10https://gerrit.wikimedia.org/r/1324754 (https://phabricator.wikimedia.org/T433890) [15:30:09] !log mforns@deploy1003 Finished deploy [analytics/refinery@49c336c] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@49c336cd] (duration: 00m 32s) [15:30:22] !log mforns@deploy1003 Started deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] [15:30:52] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.sanitize-wiki (exit_code=99) Checking sanitization for wikis testwiki in section s3 [15:31:29] !log dreamyjazz@deploy1003 kharlan, dreamyjazz: Backport for [[gerrit:1324733|WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:31:43] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324754 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [15:33:19] !log dreamyjazz@deploy1003 kharlan, dreamyjazz: Continuing with deployment [15:33:41] !log eevans@cumin1003 END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2024,1031]*.wmnet: Set storage compatability to UPGRADING — T433028 - eevans@cumin1003 [15:33:45] T433028: Upgrade restbase cluster to Cassandra 5.0.8 - https://phabricator.wikimedia.org/T433028 [15:34:13] (03PS2) 10Bking: apifeatureusage: explicitly define java package version [puppet] - 10https://gerrit.wikimedia.org/r/1324754 (https://phabricator.wikimedia.org/T433890) [15:34:42] !log mforns@deploy1003 Finished deploy [analytics/refinery@49c336c]: Regular analytics weekly train [analytics/refinery@49c336cd] (duration: 04m 20s) [15:35:28] (03CR) 10Dragoniez: "@fd7ezs8cx@mozmail.com Are you still going to backport this?" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1216721 (https://phabricator.wikimedia.org/T405724) (owner: 10Nvdtn19) [15:35:39] !log mforns@deploy1003 Started deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] [15:35:55] !log jhancock@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [15:36:11] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324754 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [15:37:35] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324733|WikimediaAntiAbuse: Enable personal info tag display on enwiki (T431292)]] (duration: 08m 12s) [15:37:39] !log mforns@deploy1003 Finished deploy [analytics/refinery@49c336c] (thin): Regular analytics weekly train THIN [analytics/refinery@49c336cd] (duration: 01m 59s) [15:37:40] T431292: Tag revisions with private change tag for likely oversightable/revision deletable content - https://phabricator.wikimedia.org/T431292 [15:38:05] (03CR) 10Bking: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1324754 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [15:40:02] (03CR) 10Bking: [C:03+2] "The PCC failure is expected due to stale facts (just reimaged `apifeatureusage2001` to Bookworm). Merging..." [puppet] - 10https://gerrit.wikimedia.org/r/1324754 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [15:40:04] I assume stats.wikimedia.org showing the Apache2 default page is because they’re being re-built or something? Nothing obvious in the SAL. [15:42:31] !log jhancock@cumin1003 START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [15:44:29] (03PS1) 10CDanis: cidergrinder: timer: run half-hr later [puppet] - 10https://gerrit.wikimedia.org/r/1324756 [15:45:52] (03CR) 10Ssingh: [C:03+1] cidergrinder: timer: run half-hr later [puppet] - 10https://gerrit.wikimedia.org/r/1324756 (owner: 10CDanis) [15:46:01] (03CR) 10TChin: [C:03+1] eventgate-analytics-external: Enable testKitchen transform [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324728 (https://phabricator.wikimedia.org/T429898) (owner: 10Phuedx) [15:47:58] (03CR) 10CDanis: [C:03+2] cidergrinder: timer: run half-hr later [puppet] - 10https://gerrit.wikimedia.org/r/1324756 (owner: 10CDanis) [15:52:27] !log jmm@cumin2003 END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host ganeti2046.codfw.wmnet [15:52:27] !log jmm@cumin2003 END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2046.codfw.wmnet [15:56:13] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db1272: New host [15:57:37] !log jhancock@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [16:02:39] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-magru and not (P{cp7001*} or P{cp7009*}) and A:cp - 9.2.15 upgrade (T434620) [16:02:43] T434620: [Update DNS Record Request] - wikimedia.org - Add TXT verification for Mentimeter - https://phabricator.wikimedia.org/T434620 [16:04:09] (03PS1) 10Bking: WIP: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [16:04:26] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) (owner: 10Bking) [16:07:48] !log tchin@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-analytics: apply [16:07:55] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:08:00] !log tchin@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply [16:08:17] !log tchin@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply [16:08:32] btullis@cumin1003 reimage (PID 3067365) is awaiting input [16:08:59] !log tchin@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply [16:10:06] !log tchin@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply [16:10:47] !log tchin@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply [16:12:21] !log eevans@cumin1003 START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — T433028 - eevans@cumin1003 [16:12:25] T433028: Upgrade restbase cluster to Cassandra 5.0.8 - https://phabricator.wikimedia.org/T433028 [16:14:45] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:15:10] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [16:16:02] !log jhancock@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] [16:17:41] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [16:18:24] !log jhancock@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] [16:18:52] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [16:19:00] (03PS1) 10Gehel: WIP: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) (owner: 10Bking) [16:19:00] (03CR) 10Gehel: [C:04-1] "Please confirm what I wrote in the comment, I did not actually test it." [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) (owner: 10Bking) [16:27:24] !log tchin@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply [16:27:33] (03CR) 10Arnaudb: gitlab: point gitlab-replica-a at the CDN (031 comment) [dns] - 10https://gerrit.wikimedia.org/r/1324280 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [16:27:37] !log tchin@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply [16:28:13] 10ops-esams, 06SRE, 06DC-Ops: ganeti3005 shows backplane error after reboot - https://phabricator.wikimedia.org/T434646#12208710 (10RobH) a:03RobH I'll have to detail out a smart hands for this, as it seems we've rebooted and updated the firmware a couple times now (thanks for the task references in the de... [16:30:07] (03PS1) 10Kosta Harlan: WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324765 (https://phabricator.wikimedia.org/T431292) [16:30:32] jouncebot: nowandnext [16:30:32] No deployments scheduled for the next 0 hour(s) and 29 minute(s) [16:30:33] In 0 hour(s) and 29 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1700) [16:31:02] (03CR) 10Dreamy Jazz: [C:03+1] WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324765 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [16:32:13] going to deploy a config patch [16:32:32] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324765 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [16:33:27] (03Merged) 10jenkins-bot: WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324765 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [16:33:45] !log kharlan@deploy1003 Started scap sync-world: Backport for [[gerrit:1324765|WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] [16:33:50] T431292: Tag revisions with private change tag for likely oversightable/revision deletable content - https://phabricator.wikimedia.org/T431292 [16:35:48] !log kharlan@deploy1003 kharlan: Backport for [[gerrit:1324765|WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [16:36:41] !log kharlan@deploy1003 kharlan: Continuing with deployment [16:37:23] !log tchin@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply [16:38:08] !log tchin@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply [16:39:39] !log tchin@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply [16:40:11] !log tchin@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply [16:40:48] !log kharlan@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324765|WikimediaAntiAbuse: Enable PersonalInfoFlagNotifications (T431292)]] (duration: 07m 02s) [16:40:52] T431292: Tag revisions with private change tag for likely oversightable/revision deletable content - https://phabricator.wikimedia.org/T431292 [16:41:27] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1272: New host [16:42:14] (03PS1) 10Btullis: install_server: Generate hadoop worker reuse recipe at install time [puppet] - 10https://gerrit.wikimedia.org/r/1324767 (https://phabricator.wikimedia.org/T434494) [16:42:29] (03PS1) 10CWilliams: mariadb: Productionize db1277 [puppet] - 10https://gerrit.wikimedia.org/r/1324768 (https://phabricator.wikimedia.org/T407942) [16:44:04] 10SRE-Access-Requests, 06Data-Engineering, 06Data-Engineering-Radar, 06Data-Platform-SRE: Create Kerberos identity for Randall Scout - https://phabricator.wikimedia.org/T430598#12208906 (10Milimetric) @Rscout The process to follow here is: https://wikitech.wikimedia.org/wiki/Data_Platform/Data_access#Requ... [16:44:31] !log tchin@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-main: apply [16:44:42] !log tchin@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-main: apply [16:47:55] (03PS2) 10Btullis: install_server: Generate hadoop worker reuse recipe at install time [puppet] - 10https://gerrit.wikimedia.org/r/1324767 (https://phabricator.wikimedia.org/T434494) [16:51:14] (03PS2) 10Gerrit maintenance bot: mariadb: Promote db2213 to s5 master [puppet] - 10https://gerrit.wikimedia.org/r/1324599 (https://phabricator.wikimedia.org/T434635) [16:52:16] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324771 [16:52:33] !log tchin@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-main: apply [16:53:18] !log tchin@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply [16:54:32] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12208979 (10Jdforrester-WMF) 05Open→03In progress [16:55:24] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 27 hosts with reason: Primary switchover s5 T434635 [16:55:28] T434635: Switchover s5 master (db2192 -> db2213) - https://phabricator.wikimedia.org/T434635 [16:55:45] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Set db2213 with weight 0 T434635', diff saved to https://phabricator.wikimedia.org/P96028 and previous config saved to /var/cache/conftool/dbconfig/20260812-165544-cwilliams.json [16:58:29] (03PS2) 10Bking: WIP: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [16:59:45] (03CR) 10CWilliams: [C:03+2] mariadb: Promote db2213 to s5 master [puppet] - 10https://gerrit.wikimedia.org/r/1324599 (https://phabricator.wikimedia.org/T434635) (owner: 10Gerrit maintenance bot) [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1700) [17:01:08] !log Starting s5 codfw failover from db2192 to db2213 - T434635 [17:01:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:01:12] T434635: Switchover s5 master (db2192 -> db2213) - https://phabricator.wikimedia.org/T434635 [17:01:18] (03PS3) 10Bking: WIP: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [17:01:52] (03CR) 10CI reject: [V:04-1] WIP: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) (owner: 10Bking) [17:01:53] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Promote db2213 to s5 primary T434635', diff saved to https://phabricator.wikimedia.org/P96029 and previous config saved to /var/cache/conftool/dbconfig/20260812-170152-cwilliams.json [17:03:30] (03PS4) 10Bking: WIP: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [17:03:39] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depool db2192 T434635', diff saved to https://phabricator.wikimedia.org/P96030 and previous config saved to /var/cache/conftool/dbconfig/20260812-170338-cwilliams.json [17:03:50] !log tchin@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-main: apply [17:04:32] !log tchin@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply [17:05:24] PROBLEM - orchestrator resolve cache non-FQDNs on dborch1002 is CRITICAL: CRITICAL: 2 non-FQDN entries in orchestrator resolve cache: https://wikitech.wikimedia.org/wiki/Orchestrator [17:11:56] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host apifeatureusage2001.codfw.wmnet with OS bookworm [17:13:40] FIRING: SystemdUnitFailed: wmf_auto_restart_rsyslog.service on ml-serve2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:14:45] FIRING: [3x] CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in cloudelastic - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [17:15:58] (03PS5) 10Bking: WIP: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [17:16:36] (03CR) 10Bking: WIP: Point java safepoint logging to a file (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) (owner: 10Bking) [17:18:17] (03PS6) 10Bking: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [17:18:49] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2192.codfw.wmnet with reason: Maintenance [17:20:32] (03PS7) 10Bking: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [17:22:23] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [17:26:47] 10ops-codfw, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install ganeti205[1-6] - https://phabricator.wikimedia.org/T434698 (10RobH) 03NEW [17:27:26] 10ops-codfw, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install ganeti205[1-6] - https://phabricator.wikimedia.org/T434698#12209084 (10RobH) a:03MoritzMuehlenhoff Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/SRE/Dc-operations... [17:27:48] 10ops-codfw, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install ganeti205[1-6] - https://phabricator.wikimedia.org/T434698#12209088 (10RobH) [17:28:02] !log jhancock@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] [17:28:49] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [17:30:44] (03PS8) 10Bking: Point java safepoint logging to a file [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [17:31:14] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-magru and not (P{cp7001*} or P{cp7009*}) and A:cp - 9.2.15 upgrade (T434620) [17:31:18] T434620: [Update DNS Record Request] - wikimedia.org - Add TXT verification for Mentimeter - https://phabricator.wikimedia.org/T434620 [17:32:23] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install puppetserver100[45] and ganeti106[0-4] - https://phabricator.wikimedia.org/T434699 (10RobH) 03NEW [17:33:11] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install puppetserver100[45] and ganeti106[0-4] - https://phabricator.wikimedia.org/T434699#12209125 (10RobH) [17:33:25] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install puppetserver100[45] and ganeti106[0-4] - https://phabricator.wikimedia.org/T434699#12209128 (10RobH) a:03MoritzMuehlenhoff Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wik... [17:35:23] !log jhancock@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] [17:35:56] (03PS1) 10Cparle: Baseline user-facing metrics for thumbnails, with an experiment mechanism [puppet] - 10https://gerrit.wikimedia.org/r/1324778 (https://phabricator.wikimedia.org/T431597) [17:36:16] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [17:36:43] (03PS2) 10Cparle: Baseline user-facing metrics for thumbnails, with an experiment mechanism [puppet] - 10https://gerrit.wikimedia.org/r/1324778 (https://phabricator.wikimedia.org/T431597) [17:41:49] !log jhancock@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] [17:42:01] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [17:47:41] !log jhancock@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['cp5022'] [17:47:57] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [17:53:41] !log jhancock@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022'] [17:55:30] !log sukhe@cumin1003 START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum-eqsin and A:durum [17:55:41] !log jhancock@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022'] [17:56:18] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db2192: Security update [17:59:10] jhancock@cumin1003 upgrade-firmware (PID 3155086) is awaiting input [18:00:04] brennen and jnuche: How many deployers does it take to do MediaWiki train - Utc-7+Utc-0 Version deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T1800). [18:00:09] o/ [18:00:16] still sorting https://phabricator.wikimedia.org/T426102 [18:05:15] !log jhancock@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022'] [18:06:02] brennen: I'm around now. Let me see [18:06:08] yay! [18:06:18] Amir1: thank you and i'm sorry [18:07:27] !log eevans@cumin1003 END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Set storage compatability to UPGRADING — T433028 - eevans@cumin1003 [18:07:32] T433028: Upgrade restbase cluster to Cassandra 5.0.8 - https://phabricator.wikimedia.org/T433028 [18:08:56] !log sukhe@cumin1003 END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum-eqsin and A:durum [18:09:54] okay, I see how Daimona wanted to move the row. It would be better if it was a separate maint script but this could work I think [18:10:00] *rows [18:11:22] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T434668 [18:11:45] (03CR) 10Eevans: [C:03+2] restbase: Set storage compatibility to NONE [puppet] - 10https://gerrit.wikimedia.org/r/1315186 (https://phabricator.wikimedia.org/T433028) (owner: 10Eevans) [18:12:30] (03PS5) 10Eevans: restbase: Set storage compatibility to NONE [puppet] - 10https://gerrit.wikimedia.org/r/1315186 (https://phabricator.wikimedia.org/T433028) [18:14:01] I'm about to migrate testwiki [18:14:14] Woo [18:14:19] (03CR) 10Eevans: [C:03+2] restbase: Set storage compatibility to NONE [puppet] - 10https://gerrit.wikimedia.org/r/1315186 (https://phabricator.wikimedia.org/T433028) (owner: 10Eevans) [18:15:25] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [18:15:51] (03PS3) 10Scott French: ingress: Copy istio 1.2.0 to 1.3.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324342 (https://phabricator.wikimedia.org/T427666) [18:15:52] (03PS5) 10Scott French: ingress: Introduce istio 1.3.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324343 (https://phabricator.wikimedia.org/T427666) [18:15:52] (03PS1) 10Scott French: mw-pretrain: Introduce dedicated Ingress Gateway for jobrunner [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324780 (https://phabricator.wikimedia.org/T427666) [18:16:09] okay, the table doesn't even exist yet. Need to fix that [18:18:21] testwiki migrated, now test2wiki [18:18:28] !log eevans@cumin1003 START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Set storage compatability to NONE — T433028 - eevans@cumin1003 [18:18:34] T433028: Upgrade restbase cluster to Cassandra 5.0.8 - https://phabricator.wikimedia.org/T433028 [18:18:45] !log migrated testwiki entries from ce_worklist_articles to ce_invitation_list_articles (T426102) [18:18:49] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:18:50] T426102: Rename current 'Worklist' texts in invitation lists" - https://phabricator.wikimedia.org/T426102 [18:19:33] test2wiki migrated [18:20:41] officewiki is done too [18:20:53] (03CR) 10Ladsgroup: [C:03+2] Rename ce_worklist_articles table to ce_invitation_list_articles [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321224 (https://phabricator.wikimedia.org/T426102) (owner: 10Daimona Eaytoy) [18:21:05] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T434668 [18:21:39] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T434668 [18:22:34] Amir1 thank you, and sorry everyone for the disruption! Yes I wasn't sure if it was better to do it that way or with an actual rename [18:22:56] That's why I highly discourage table renames :) [18:23:03] thanks both. [18:23:35] I'm going to create the table as empty now on wikishared, backport the patch, then run the migrate script [18:24:03] ack, thanks - i'll await an all clear on the train [18:24:48] (plenty of window here. i'll plan to go to group0 and let it bake in for a few minutes then group1 on schedule.) [18:25:03] !log ce_invitation_list_articles created as empty on wikishared (T426102) [18:25:06] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:25:07] T426102: Rename current 'Worklist' texts in invitation lists" - https://phabricator.wikimedia.org/T426102 [18:25:33] you can go to group0 if you want to [18:27:46] k [18:28:36] (03CR) 10Ssingh: [C:03+1] gitlab: point gitlab-replica-a at the CDN [dns] - 10https://gerrit.wikimedia.org/r/1324280 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [18:30:10] (03CR) 10CI reject: [V:04-1] Rename ce_worklist_articles table to ce_invitation_list_articles [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321224 (https://phabricator.wikimedia.org/T426102) (owner: 10Daimona Eaytoy) [18:30:54] Ah joy, the PHPCS backport [18:31:04] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T434668 [18:31:04] (03CR) 10Ladsgroup: [C:03+2] "try again" [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321224 (https://phabricator.wikimedia.org/T426102) (owner: 10Daimona Eaytoy) [18:31:14] (03PS1) 10Daimona Eaytoy: Bump mediawiki/mediawiki-codesniffer to v52.0.0 [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324782 [18:31:31] It'll need this backport https://gerrit.wikimedia.org/r/c/mediawiki/extensions/CampaignEvents/+/1324782 [18:31:41] (Or force-merging but better not I guess) [18:32:04] (03CR) 10CI reject: [V:04-1] Rename ce_worklist_articles table to ce_invitation_list_articles [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321224 (https://phabricator.wikimedia.org/T426102) (owner: 10Daimona Eaytoy) [18:32:32] (03CR) 10Ladsgroup: [C:03+2] Bump mediawiki/mediawiki-codesniffer to v52.0.0 [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324782 (owner: 10Daimona Eaytoy) [18:36:37] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12209260 (10ARamirez_WMF) Hello, I can't make it past the https://idp.wikimedia.org/login page. I can login into CAS, and have a message that says, "you have successfully... [18:38:52] (03CR) 10Ryan Kemper: install_server: Generate hadoop worker reuse recipe at install time (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324767 (https://phabricator.wikimedia.org/T434494) (owner: 10Btullis) [18:38:57] FIRING: KubernetesCalicoDown: kubestagemaster2005.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=codfw%20prometheus%2Fk8s-staging&var-instance=kubestagemaster2005.codfw.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [18:40:48] (03CR) 10Ladsgroup: [C:03+2] "one more time" [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321224 (https://phabricator.wikimedia.org/T426102) (owner: 10Daimona Eaytoy) [18:41:05] (03Merged) 10jenkins-bot: Bump mediawiki/mediawiki-codesniffer to v52.0.0 [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324782 (owner: 10Daimona Eaytoy) [18:41:29] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2192: Security update [18:49:17] (03Merged) 10jenkins-bot: Rename ce_worklist_articles table to ce_invitation_list_articles [extensions/CampaignEvents] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1321224 (https://phabricator.wikimedia.org/T426102) (owner: 10Daimona Eaytoy) [18:51:06] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1321224|Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] [18:51:12] T426102: Rename current 'Worklist' texts in invitation lists" - https://phabricator.wikimedia.org/T426102 [18:52:19] (03PS5) 10Scott French: api-gateway: Drop support for debug_hosts [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305236 (https://phabricator.wikimedia.org/T433752) [18:52:20] (03PS4) 10Scott French: api-gateway: Remove stale test assertion and noop Lua code [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311874 (https://phabricator.wikimedia.org/T433752) [18:52:21] (03PS4) 10Scott French: api-gateway: Drop support for php_engine_routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311875 (https://phabricator.wikimedia.org/T433752) [18:52:23] (03PS3) 10Scott French: api-gateway: Basic cluster specifier support and Lua plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) [18:52:25] (03PS5) 10Scott French: api-gateway: Support x-wikimedia-debug routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311964 (https://phabricator.wikimedia.org/T433752) [18:52:29] (03PS1) 10Scott French: api-gateway: Support Host-based diversion in the mw-api plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322140 (https://phabricator.wikimedia.org/T433752) [18:53:15] !log ladsgroup@deploy1003 ladsgroup, daimona: Backport for [[gerrit:1321224|Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:53:49] !log ladsgroup@deploy1003 ladsgroup, daimona: Continuing with deployment [18:54:05] FIRING: [6x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:57:53] (03PS2) 10Scott French: api-gateway: Support Host-based diversion in the mw-api plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322140 (https://phabricator.wikimedia.org/T433752) [18:57:57] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1321224|Rename ce_worklist_articles table to ce_invitation_list_articles (T426102)]] (duration: 06m 50s) [18:58:01] T426102: Rename current 'Worklist' texts in invitation lists" - https://phabricator.wikimedia.org/T426102 [19:00:00] !log data migrated on wikishared (T426102) [19:00:04] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:01:05] brennen: you should be good to go [19:02:01] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12209331 (10VRiley-WMF) Hey @fgiunchedi Yes, that would work for me! Would monday next week work for you? [19:02:21] I go afk, will be back in a bit [19:02:34] thanks Amir1, much appreciated. [19:02:54] No worries <3 [19:02:58] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.15 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324785 (https://phabricator.wikimedia.org/T430834) [19:03:00] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by brennen@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324785 (https://phabricator.wikimedia.org/T430834) (owner: 10TrainBranchBot) [19:03:06] Glad I could be useful sometimes [19:03:21] many many times [19:03:24] ++ [19:03:54] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.15 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324785 (https://phabricator.wikimedia.org/T430834) (owner: 10TrainBranchBot) [19:04:57] (03CR) 10Klausman: [C:03+1] "Aside from what Ryan already mentioned: LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1324767 (https://phabricator.wikimedia.org/T434494) (owner: 10Btullis) [19:04:59] 10ops-eqiad, 06DC-Ops: Shipping out RMA for Juniper - https://phabricator.wikimedia.org/T434706 (10VRiley-WMF) 03NEW [19:05:30] 10ops-eqiad, 06DC-Ops: Shipping out RMA for Juniper - https://phabricator.wikimedia.org/T434706#12209353 (10VRiley-WMF) 05Open→03Resolved This has been completed. [19:06:45] <3 [19:06:52] (03PS1) 10CDobbins: installserver: add cp5022 to preseed.yaml regex [puppet] - 10https://gerrit.wikimedia.org/r/1324786 (https://phabricator.wikimedia.org/T414411) [19:09:00] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [19:09:15] (03CR) 10Ssingh: [C:03+1] installserver: add cp5022 to preseed.yaml regex [puppet] - 10https://gerrit.wikimedia.org/r/1324786 (https://phabricator.wikimedia.org/T414411) (owner: 10CDobbins) [19:09:16] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [19:10:04] !log brennen@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs T430834 [19:10:09] T430834: 1.47.0-wmf.15 deployment blockers - https://phabricator.wikimedia.org/T430834 [19:15:49] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.15 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324788 (https://phabricator.wikimedia.org/T430834) [19:15:54] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by brennen@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324788 (https://phabricator.wikimedia.org/T430834) (owner: 10TrainBranchBot) [19:16:54] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.15 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324788 (https://phabricator.wikimedia.org/T430834) (owner: 10TrainBranchBot) [19:17:14] (03CR) 10CDobbins: [C:03+2] installserver: add cp5022 to preseed.yaml regex [puppet] - 10https://gerrit.wikimedia.org/r/1324786 (https://phabricator.wikimedia.org/T414411) (owner: 10CDobbins) [19:19:55] !log eevans@cumin1003 END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Set storage compatability to NONE — T433028 - eevans@cumin1003 [19:20:01] T433028: Upgrade restbase cluster to Cassandra 5.0.8 - https://phabricator.wikimedia.org/T433028 [19:21:11] (03PS3) 10Ryan Kemper: install_server: Generate hadoop worker reuse recipe at install time [puppet] - 10https://gerrit.wikimedia.org/r/1324767 (https://phabricator.wikimedia.org/T434494) (owner: 10Btullis) [19:23:03] !log brennen@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.15 refs T430834 [19:23:08] T430834: 1.47.0-wmf.15 deployment blockers - https://phabricator.wikimedia.org/T430834 [19:26:00] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [19:29:18] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [19:30:16] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [19:32:58] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [19:33:15] (03PS2) 10Scott French: mw-pretrain: Introduce dedicated Ingress Gateway for jobrunner [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324780 (https://phabricator.wikimedia.org/T427666) [19:34:25] (03PS1) 10CDobbins: wikimedia.org: add TXT verification for Mentimeter [dns] - 10https://gerrit.wikimedia.org/r/1324790 (https://phabricator.wikimedia.org/T434620) [19:35:00] (03PS5) 10Scott French: mediawiki: Bump ingress.istio from 1.2.0 to 1.3.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324344 (https://phabricator.wikimedia.org/T427666) [19:35:21] (03CR) 10CI reject: [V:04-1] wikimedia.org: add TXT verification for Mentimeter [dns] - 10https://gerrit.wikimedia.org/r/1324790 (https://phabricator.wikimedia.org/T434620) (owner: 10CDobbins) [19:35:23] !log cdobbins@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie [19:35:36] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12209464 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (**FAIL**) - Removed from... [19:35:45] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie [19:35:57] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12209465 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [19:36:13] (03CR) 10Ryan Kemper: [C:03+2] install_server: Generate hadoop worker reuse recipe at install time [puppet] - 10https://gerrit.wikimedia.org/r/1324767 (https://phabricator.wikimedia.org/T434494) (owner: 10Btullis) [19:36:22] (03CR) 10Ssingh: wikimedia.org: add TXT verification for Mentimeter (032 comments) [dns] - 10https://gerrit.wikimedia.org/r/1324790 (https://phabricator.wikimedia.org/T434620) (owner: 10CDobbins) [19:36:24] (03CR) 10Ryan Kemper: [C:03+2] install_server: Generate hadoop worker reuse recipe at install time (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324767 (https://phabricator.wikimedia.org/T434494) (owner: 10Btullis) [19:38:55] (03PS1) 10Scott French: admin_ng: Add mw-pretrain-jobrunner to mw-pretrain tlsExtraSANs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324775 (https://phabricator.wikimedia.org/T427666) [19:38:56] (03PS1) 10Scott French: wmnet: Add mw-pretrain-jobrunner k8s Ingress CNAMEs [dns] - 10https://gerrit.wikimedia.org/r/1324776 (https://phabricator.wikimedia.org/T427666) [19:47:11] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie [19:47:25] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12209524 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS t... [19:47:39] rolling back for T434712 [19:47:39] T434712: Error: Class "Wikibase\Repo\WikibaseRepo" not found - https://phabricator.wikimedia.org/T434712 [19:47:59] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.15 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324793 (https://phabricator.wikimedia.org/T430834) [19:48:04] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by brennen@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324793 (https://phabricator.wikimedia.org/T430834) (owner: 10TrainBranchBot) [19:48:49] (03PS1) 10Kosta Harlan: WikimediaAntiAbuse: Enable logging channel [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324794 (https://phabricator.wikimedia.org/T431292) [19:49:01] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.15 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324793 (https://phabricator.wikimedia.org/T430834) (owner: 10TrainBranchBot) [19:49:47] (03CR) 10CI reject: [V:04-1] WikimediaAntiAbuse: Enable logging channel [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324794 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [19:49:51] (03PS2) 10Kosta Harlan: WikimediaAntiAbuse: Enable logging channel [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324794 (https://phabricator.wikimedia.org/T431292) [19:55:14] !log brennen@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.15 refs T430834 [19:55:18] T430834: 1.47.0-wmf.15 deployment blockers - https://phabricator.wikimedia.org/T430834 [19:57:08] gah, probably jumped the gun on that rollback honestly. [19:57:57] anyway, i'll leave things alone for backport window. [19:58:05] (03CR) 10Dreamy Jazz: [C:03+1] WikimediaAntiAbuse: Enable logging channel [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324794 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [19:58:06] PROBLEM - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:59:28] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324794 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [19:59:52] brennen: Seems like I'm the only one in the window, so could always use it once I'm done? [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: OwO what's this, a deployment window?? UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T2000). nyaa~ [20:00:05] Dreamy_Jazz: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:17] I'll self-deploy [20:00:40] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324794 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [20:01:36] (03Merged) 10jenkins-bot: WikimediaAntiAbuse: Enable logging channel [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324794 (https://phabricator.wikimedia.org/T431292) (owner: 10Kosta Harlan) [20:01:56] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1324794|WikimediaAntiAbuse: Enable logging channel (T431292)]] [20:01:59] T431292: Tag revisions with private change tag for likely oversightable/revision deletable content - https://phabricator.wikimedia.org/T431292 [20:04:16] !log dreamyjazz@deploy1003 kharlan, dreamyjazz: Backport for [[gerrit:1324794|WikimediaAntiAbuse: Enable logging channel (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:04:24] RECOVERY - orchestrator resolve cache non-FQDNs on dborch1002 is OK: OK: all orchestrator resolve cache entries are FQDNs https://wikitech.wikimedia.org/wiki/Orchestrator [20:04:42] !log dreamyjazz@deploy1003 kharlan, dreamyjazz: Continuing with deployment [20:05:12] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage [20:05:35] (03PS9) 10Bking: opensearch: point gc+age tracing to a file, and stop logging safepoint [puppet] - 10https://gerrit.wikimedia.org/r/1324763 (https://phabricator.wikimedia.org/T434685) [20:08:36] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage [20:08:44] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324794|WikimediaAntiAbuse: Enable logging channel (T431292)]] (duration: 06m 48s) [20:08:47] T431292: Tag revisions with private change tag for likely oversightable/revision deletable content - https://phabricator.wikimedia.org/T431292 [20:09:26] !log Evening UTC backport window done [20:09:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:09:37] brennen: If you wanted to do train stuff, go ahead :D [20:09:38] !log eevans@cumin1003 START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Set storage compatability to NONE — T433028 - eevans@cumin1003 [20:09:43] T433028: Upgrade restbase cluster to Cassandra 5.0.8 - https://phabricator.wikimedia.org/T433028 [20:09:44] thx Dreamy_Jazz [20:11:54] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm [20:11:58] !log ryankemper@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1178.eqiad.wmnet with OS bookworm [20:12:08] Dreamy_Jazz: thanks, I've got a late arrival after yours. I believe it's all clear so I'll start shortly [20:12:32] Was brennen going to run the train? [20:12:42] go ahead Krinkle [20:12:48] i'll roll back to group1 after that. [20:13:31] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1178.eqiad.wmnet with OS bookworm [20:14:18] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, August 12 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324277 (https://phabricator.wikimedia.org/T433891) (owner: 10Krinkle) [20:14:45] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:16:25] FIRING: SystemdUnitFailed: prometheus-pg-replication-lag.service on maps1012:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:17:18] (03PS2) 10CDobbins: wikimedia.org: add TXT verification for Mentimeter [dns] - 10https://gerrit.wikimedia.org/r/1324790 (https://phabricator.wikimedia.org/T434620) [20:19:23] (or maybe i won't. unclear if this will break REST sandbox on more/bigger wikis, but seems like it might.) [20:19:57] (03CR) 10Ssingh: wikimedia.org: add TXT verification for Mentimeter (032 comments) [dns] - 10https://gerrit.wikimedia.org/r/1324790 (https://phabricator.wikimedia.org/T434620) (owner: 10CDobbins) [20:21:25] RESOLVED: SystemdUnitFailed: prometheus-pg-replication-lag.service on maps1012:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:22:22] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324277 (https://phabricator.wikimedia.org/T433891) (owner: 10Krinkle) [20:28:03] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage [20:33:20] vriley@cumin1003 reimage (PID 3166146) is awaiting input [20:33:41] (03Merged) 10jenkins-bot: Improve Math preference labels for SVG/MathJax/MathML [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324277 (https://phabricator.wikimedia.org/T433891) (owner: 10Krinkle) [20:33:46] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie [20:33:58] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12209754 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS trixi... [20:34:02] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1324277|Improve Math preference labels for SVG/MathJax/MathML (T433891)]] [20:34:07] T433891: Improve Math preference labels - https://phabricator.wikimedia.org/T433891 [20:34:20] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1178.eqiad.wmnet with reason: host reimage [20:34:26] (03PS3) 10CDobbins: wikimedia.org: add TXT verification for Mentimeter [dns] - 10https://gerrit.wikimedia.org/r/1324790 (https://phabricator.wikimedia.org/T434620) [20:36:23] !log cdobbins@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" [20:37:07] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [20:38:22] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - cdobbins@cumin1003" [20:38:23] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie [20:38:34] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12209786 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie completed: - cp5022 (**PASS**) - Removed from Puppet and... [20:38:51] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12209788 (10jcrespo) 05Resolved→03Open I can see that you haven't been added to the wmf group. I belive that's needed to access superset at all. In any case, what you suff... [20:38:54] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [20:39:13] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie [20:39:29] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12209791 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS t... [20:39:46] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12209795 (10jcrespo) a:05Arnoldokoth→03jcrespo [20:48:06] RECOVERY - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [20:51:38] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1324277|Improve Math preference labels for SVG/MathJax/MathML (T433891)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:51:43] T433891: Improve Math preference labels - https://phabricator.wikimedia.org/T433891 [20:54:01] !log krinkle@deploy1003 krinkle: Continuing with deployment [20:55:12] (03PS1) 10Bking: apifeatureusage: update config for OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1324805 (https://phabricator.wikimedia.org/T433890) [20:58:08] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1178.eqiad.wmnet with OS bookworm [21:00:04] Deploy window Wikifunctions Services UTC Late (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T2100) [21:04:42] (03PS2) 10Bking: apifeatureusage: update config for OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1324805 (https://phabricator.wikimedia.org/T433890) [21:04:52] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324805 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [21:05:45] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324277|Improve Math preference labels for SVG/MathJax/MathML (T433891)]] (duration: 31m 42s) [21:05:49] T433891: Improve Math preference labels - https://phabricator.wikimedia.org/T433891 [21:07:27] brennen: all yours [21:09:39] !log vriley@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host zuul1006 [21:10:37] !log vriley@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006 [21:10:57] !log eevans@cumin1003 END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Set storage compatability to NONE — T433028 - eevans@cumin1003 [21:11:01] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [21:11:02] T433028: Upgrade restbase cluster to Cassandra 5.0.8 - https://phabricator.wikimedia.org/T433028 [21:11:33] (03PS1) 10Lerickson: Enable egress via urldownloader for WDQS V2. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324809 (https://phabricator.wikimedia.org/T434321) [21:11:47] (03PS1) 10Ahmon Dancy: deployment-info.php: Report dbname and branch for the requested wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324810 (https://phabricator.wikimedia.org/T434726) [21:13:25] RESOLVED: SystemdUnitFailed: wmf_auto_restart_rsyslog.service on ml-serve2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:14:12] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [21:14:36] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:14:44] FIRING: [3x] CirrusSearchSaneitizerFixRateTooHigh: MediaWiki CirrusSearch Saneitizer is fixing an abnormally high number of documents in cloudelastic - https://wikitech.wikimedia.org/wiki/Search/CirrusStreamingUpdater#San(e)itizing - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchSaneitizerFixRateTooHigh [21:15:15] (03CR) 10Bking: [C:03+2] apifeatureusage: update config for OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1324805 (https://phabricator.wikimedia.org/T433890) (owner: 10Bking) [21:17:46] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:25:31] (03PS1) 10Eevans: aqs: canary Cassandra 5.0.8 upgrade [puppet] - 10https://gerrit.wikimedia.org/r/1324812 (https://phabricator.wikimedia.org/T433026) [21:25:33] (03PS1) 10Eevans: aqs: Cassandra 5.0.8 upgrade [puppet] - 10https://gerrit.wikimedia.org/r/1324813 (https://phabricator.wikimedia.org/T433026) [21:25:35] (03PS1) 10Eevans: aqs: upgrade to Java 17 [puppet] - 10https://gerrit.wikimedia.org/r/1324814 (https://phabricator.wikimedia.org/T433026) [21:25:38] (03PS1) 10Eevans: aqs: set storage compatability mode to `UPGRADING` [puppet] - 10https://gerrit.wikimedia.org/r/1324815 (https://phabricator.wikimedia.org/T433026) [21:25:40] (03PS1) 10Eevans: aqs: set storage compatability mode to `NONE` [puppet] - 10https://gerrit.wikimedia.org/r/1324816 (https://phabricator.wikimedia.org/T433026) [21:26:58] (03CR) 10Eevans: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324812 (https://phabricator.wikimedia.org/T433026) (owner: 10Eevans) [21:29:47] jouncebot: nowandnext [21:29:48] For the next 0 hour(s) and 30 minute(s): Wikifunctions Services UTC Late (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T2100) [21:29:48] In 0 hour(s) and 30 minute(s): Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T2200) [21:30:09] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:30:26] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thanks!!" [debs/pint] - 10https://gerrit.wikimedia.org/r/1319039 (owner: 10Hnowlan) [21:30:40] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:30:58] (03PS1) 10Dreamy Jazz: Create a maintenance script to backfill Special:AbuseReview [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324817 (https://phabricator.wikimedia.org/T434688) [21:31:08] (03PS1) 10Dreamy Jazz: Create a maintenance script to backfill Special:AbuseReview [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324818 (https://phabricator.wikimedia.org/T434688) [21:31:13] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:31:40] (03CR) 10Subramanya Sastry: [C:03+1] prv: Enable parsoid rendering for 5 wikisource wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324731 (owner: 10Jgiannelos) [21:32:26] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [21:32:45] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324817 (https://phabricator.wikimedia.org/T434688) (owner: 10Dreamy Jazz) [21:32:45] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324818 (https://phabricator.wikimedia.org/T434688) (owner: 10Dreamy Jazz) [21:35:34] (03Merged) 10jenkins-bot: Create a maintenance script to backfill Special:AbuseReview [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1324817 (https://phabricator.wikimedia.org/T434688) (owner: 10Dreamy Jazz) [21:36:01] (03Merged) 10jenkins-bot: Create a maintenance script to backfill Special:AbuseReview [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.14) - 10https://gerrit.wikimedia.org/r/1324818 (https://phabricator.wikimedia.org/T434688) (owner: 10Dreamy Jazz) [21:36:27] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1324817|Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818|Create a maintenance script to backfill Special:AbuseReview (T434688)]] [21:36:32] T434688: Create maintenance script to evaluate revisions performed between timestamps against content policies - https://phabricator.wikimedia.org/T434688 [21:39:40] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie [21:39:53] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12210023 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS t... [21:40:29] !log dreamyjazz@deploy1003 dreamyjazz: Backport for [[gerrit:1324817|Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818|Create a maintenance script to backfill Special:AbuseReview (T434688)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:43:11] (03PS1) 10Milazg: Add configurable RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) [21:50:33] FIRING: KubernetesAPINotScrapable: k8s-dse@codfw is failing to scrape the k8s api - https://phabricator.wikimedia.org/T343529 - TODO - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPINotScrapable [21:54:34] 10ops-codfw, 06SRE, 06DC-Ops: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12210044 (10Jhancock.wm) checking [21:55:33] RESOLVED: KubernetesAPINotScrapable: k8s-dse@codfw is failing to scrape the k8s api - https://phabricator.wikimedia.org/T343529 - TODO - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPINotScrapable [21:59:28] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie [21:59:42] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12210066 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS trixi... [22:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260812T2200) [22:15:25] FIRING: [2x] GanetiBGPDown: BGP session down between ganeti3005 and asw1-by27-esams - group Ganeti4 - https://wikitech.wikimedia.org/wiki/Ganeti#GanetiBGPDown - https://alerts.wikimedia.org/?q=alertname%3DGanetiBGPDown [22:17:49] !log dreamyjazz@deploy1003 dreamyjazz: Continuing with deployment [22:21:56] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324817|Create a maintenance script to backfill Special:AbuseReview (T434688)]], [[gerrit:1324818|Create a maintenance script to backfill Special:AbuseReview (T434688)]] (duration: 45m 29s) [22:22:01] T434688: Create maintenance script to evaluate revisions performed between timestamps against content policies - https://phabricator.wikimedia.org/T434688 [22:35:26] !log Running `mwscript WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260101000000" --end-timestamp="20260102000000" --sleep=10` [22:35:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:36:48] (03PS1) 10Scott French: kubernetes: Enable and configure mw-pretrain scap deployments [puppet] - 10https://gerrit.wikimedia.org/r/1324826 (https://phabricator.wikimedia.org/T428972) [22:39:12] FIRING: KubernetesCalicoDown: kubestagemaster2005.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=codfw%20prometheus%2Fk8s-staging&var-instance=kubestagemaster2005.codfw.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [22:40:55] !log Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200101010101" --end-timestamp="20260816010101" --sleep=60` [22:40:57] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:45:41] !log Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=testwiki --start-timestamp="20200801010101" --end-timestamp="20260816010101" --sleep=15 --batch-size=5` [22:45:44] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:49:05] !log Running `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260801000000" --end-timestamp="20260802000000" --sleep=2 --batch-size=10` [22:49:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:54:06] FIRING: [6x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [22:59:51] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie [23:00:04] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12210192 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS trixi... [23:00:05] 10ops-codfw, 06SRE, 06DC-Ops: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12210194 (10Jhancock.wm) server is definitely hanging there. i tried two remote reboots (one warm, one cold) and it's still not progressing. went on site and pulled the power cables for a few m... [23:09:32] jouncebot: nowandnext [23:09:32] No deployments scheduled for the next 6 hour(s) and 50 minute(s) [23:09:32] In 6 hour(s) and 50 minute(s): MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260813T0600) [23:09:32] In 6 hour(s) and 50 minute(s): Primary database switchover (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260813T0600) [23:10:01] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324409 (https://phabricator.wikimedia.org/T119117) (owner: 10Ladsgroup) [23:11:07] (03Merged) 10jenkins-bot: Remove $wmg = $wg hacks in Collection [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324409 (https://phabricator.wikimedia.org/T119117) (owner: 10Ladsgroup) [23:11:24] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1324409|Remove $wmg = $wg hacks in Collection (T119117)]] [23:11:28] T119117: Get rid of $wg = $wmg hack - https://phabricator.wikimedia.org/T119117 [23:12:44] 10ops-codfw, 06SRE, 06DC-Ops: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12210206 (10Jhancock.wm) case number: SM2608129193 [23:13:33] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1324409|Remove $wmg = $wg hacks in Collection (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [23:14:02] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [23:14:57] (03CR) 10Scott French: "I think we've all separately agreed that we're ready to take this step :)" [puppet] - 10https://gerrit.wikimedia.org/r/1324826 (https://phabricator.wikimedia.org/T428972) (owner: 10Scott French) [23:18:07] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324409|Remove $wmg = $wg hacks in Collection (T119117)]] (duration: 06m 43s) [23:18:12] T119117: Get rid of $wg = $wmg hack - https://phabricator.wikimedia.org/T119117 [23:23:04] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes, 13Patch-For-Review: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12210211 (10Jhancock.wm) case number: SM2608129194 [23:27:06] (03CR) 10Scott French: "This looks reasonable [0] to me! I *might* be able to deploy this during tomorrow's UTC-late infra window [1], depending on how many other" [puppet] - 10https://gerrit.wikimedia.org/r/1315141 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [23:29:10] (03CR) 10Nvdtn19: "Sorry, I missed email notifications. I still want to backport this." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1216721 (https://phabricator.wikimedia.org/T405724) (owner: 10Nvdtn19) [23:32:11] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 17 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1216721 (https://phabricator.wikimedia.org/T405724) (owner: 10Nvdtn19) [23:33:52] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 17 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1216721 (https://phabricator.wikimedia.org/T405724) (owner: 10Nvdtn19) [23:39:37] (03PS1) 10Ladsgroup: Migrate $wgFlaggedRevsTags from flaggedrevs.php to ext-FlaggedRevs.php [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324832 [23:42:20] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1324833 [23:42:20] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1324833 (owner: 10TrainBranchBot) [23:42:36] (03CR) 10Nvdtn19: "Done" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1216721 (https://phabricator.wikimedia.org/T405724) (owner: 10Nvdtn19) [23:53:29] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1324833 (owner: 10TrainBranchBot)