[00:00:29] !log ryankemper@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" [00:00:34] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1224 - ryankemper@cumin2003" [00:00:34] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [00:00:34] !log ryankemper@cumin2003 START - Cookbook sre.dns.wipe-cache an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [00:00:37] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1224.eqiad.wmnet 21.36.64.10.in-addr.arpa 1.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [00:00:38] !log ryankemper@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1224 [00:02:15] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1224 [00:02:15] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1224 [00:02:58] !log ryankemper@cumin2003 START - Cookbook sre.hosts.move-vlan for host an-worker1225 [00:03:15] !log ryankemper@cumin2003 START - Cookbook sre.dns.netbox [00:04:07] (03CR) 10Cwhite: [C:03+2] profile: add and use placeholder file in security_plugin [puppet] - 10https://gerrit.wikimedia.org/r/1325568 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [00:10:03] !log ryankemper@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" [00:10:07] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1225 - ryankemper@cumin2003" [00:10:08] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [00:10:08] !log ryankemper@cumin2003 START - Cookbook sre.dns.wipe-cache an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [00:10:11] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1225.eqiad.wmnet 22.36.64.10.in-addr.arpa 2.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [00:10:12] !log ryankemper@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1225 [00:11:01] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1225 [00:11:01] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1225 [00:12:03] !log ryankemper@cumin2003 START - Cookbook sre.hosts.move-vlan for host an-worker1226 [00:12:11] !log ryankemper@cumin2003 START - Cookbook sre.dns.netbox [00:13:04] (03CR) 10Ryan Kemper: [C:03+1] dse-k8s-codfw: remove entry for unused host [puppet] - 10https://gerrit.wikimedia.org/r/1325566 (https://phabricator.wikimedia.org/T434793) (owner: 10Bking) [00:16:10] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage [00:17:14] !log ryankemper@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" [00:17:18] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1226 - ryankemper@cumin2003" [00:17:19] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [00:17:19] !log ryankemper@cumin2003 START - Cookbook sre.dns.wipe-cache an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [00:17:22] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1226.eqiad.wmnet 23.36.64.10.in-addr.arpa 3.2.0.0.6.3.0.0.4.6.0.0.0.1.0.0.6.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [00:17:22] !log ryankemper@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1226 [00:17:54] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1226 [00:17:54] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1226 [00:19:35] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1224.eqiad.wmnet with reason: host reimage [00:25:01] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage [00:28:07] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1225.eqiad.wmnet with reason: host reimage [00:31:47] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage [00:38:55] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1226.eqiad.wmnet with reason: host reimage [00:41:50] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1224.eqiad.wmnet with OS bookworm [00:47:51] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1225.eqiad.wmnet with OS bookworm [00:55:02] (03PS1) 10Reedy: Periodic Jobs: Add purge_expired_recovery_codes for SUL and non SUL wikis [puppet] - 10https://gerrit.wikimedia.org/r/1325573 (https://phabricator.wikimedia.org/T422922) [00:55:35] (03CR) 10CI reject: [V:04-1] Periodic Jobs: Add purge_expired_recovery_codes for SUL and non SUL wikis [puppet] - 10https://gerrit.wikimedia.org/r/1325573 (https://phabricator.wikimedia.org/T422922) (owner: 10Reedy) [00:56:13] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [00:57:34] (03PS2) 10Reedy: Periodic Jobs: Add purge_expired_recovery_codes for SUL and non SUL wikis [puppet] - 10https://gerrit.wikimedia.org/r/1325573 (https://phabricator.wikimedia.org/T422922) [00:59:28] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1226.eqiad.wmnet with OS bookworm [01:02:23] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [01:06:02] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [01:07:08] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [01:08:01] (03CR) 10Tim Starling: Enable Produnto on Beta (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324964 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [01:10:32] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [01:11:20] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1325575 [01:11:20] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1325575 (owner: 10TrainBranchBot) [01:12:55] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:14:45] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:17:09] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [01:17:09] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [01:17:09] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [01:21:25] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1325575 (owner: 10TrainBranchBot) [01:21:39] FIRING: [2x] TransitBGPDown: Transit BGP session down between cr2-eqsin and Tata (180.87.164.61) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [01:26:39] RESOLVED: [2x] TransitBGPDown: Transit BGP session down between cr2-eqsin and Tata (180.87.164.61) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [01:38:59] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [01:38:59] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [01:38:59] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.136 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [01:47:55] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [01:53:46] (03CR) 10Subramanya Sastry: [C:03+1] Enable Produnto on Beta (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324964 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [01:56:13] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [01:56:29] PROBLEM - statsv Varnishkafka log producer on cp3070 is CRITICAL: PROCS CRITICAL: 2 processes with args /usr/bin/varnishkafka -S /etc/varnishkafka/statsv.conf https://wikitech.wikimedia.org/wiki/Analytics/Systems/Varnishkafka [01:57:29] RECOVERY - statsv Varnishkafka log producer on cp3070 is OK: PROCS OK: 1 process with args /usr/bin/varnishkafka -S /etc/varnishkafka/statsv.conf https://wikitech.wikimedia.org/wiki/Analytics/Systems/Varnishkafka [02:00:49] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:07:53] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 03s) [02:21:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1147:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1147 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [02:22:52] FIRING: [2x] GitLabRestoreStale: GitLab - replica gitlab1003:0 has not completed a restore in 30h - https://wikitech.wikimedia.org/wiki/GitLab/Backup_and_Restore - https://grafana.wikimedia.org/d/R_1IvBZnz/gitlab-omnibus-overview - https://alerts.wikimedia.org/?q=alertname%3DGitLabRestoreStale [02:31:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1147:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1147 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [03:21:07] FIRING: [2x] GitLabReplicaDataStale: GitLab - replica gitlab1003:0 serves data older than 30h - https://wikitech.wikimedia.org/wiki/GitLab/Backup_and_Restore - https://grafana.wikimedia.org/d/R_1IvBZnz/gitlab-omnibus-overview - https://alerts.wikimedia.org/?q=alertname%3DGitLabReplicaDataStale [03:28:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps2011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:38:29] (03CR) 10Samwilson: [C:03+1] "Looks good to me!" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270141 (https://phabricator.wikimedia.org/T424495) (owner: 10TheDJ) [03:44:06] FIRING: [2x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [04:23:03] (03CR) 10Clare Ming: tk_constructive_edits: use foreachwikiindblist instead and ignore errors (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1325555 (https://phabricator.wikimedia.org/T431493) (owner: 10DLynch) [05:09:07] !log bking@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host apifeatureusage1001.eqiad.wmnet with OS bookworm [05:12:46] (03CR) 10DLynch: tk_constructive_edits: use foreachwikiindblist instead and ignore errors (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1325555 (https://phabricator.wikimedia.org/T431493) (owner: 10DLynch) [05:14:58] (03PS2) 10DLynch: tk_constructive_edits: use foreachwikiindblist instead and ignore errors [puppet] - 10https://gerrit.wikimedia.org/r/1325555 (https://phabricator.wikimedia.org/T431493) [05:47:05] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1175.eqiad.wmnet with OS bookworm [05:47:35] !log ryankemper@cumin2003 START - Cookbook sre.hosts.move-vlan for host an-worker1175 [05:47:38] !log ryankemper@cumin2003 START - Cookbook sre.dns.netbox [05:48:12] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [05:48:38] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1209.eqiad.wmnet with OS bookworm [05:49:01] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1227.eqiad.wmnet with OS bookworm [05:49:46] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1228.eqiad.wmnet with OS bookworm [05:52:36] !log ryankemper@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" [05:52:39] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1175 - ryankemper@cumin2003" [05:52:40] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [05:52:40] !log ryankemper@cumin2003 START - Cookbook sre.dns.wipe-cache an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [05:52:43] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1175.eqiad.wmnet 17.53.64.10.in-addr.arpa 7.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [05:52:43] !log ryankemper@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1175 [05:54:18] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1175 [05:54:18] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1175 [05:54:37] !log ryankemper@cumin2003 START - Cookbook sre.hosts.move-vlan for host an-worker1209 [05:54:40] !log ryankemper@cumin2003 START - Cookbook sre.dns.netbox [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260814T0600) [06:00:08] !log ryankemper@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" [06:00:13] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1209 - ryankemper@cumin2003" [06:00:13] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [06:00:13] !log ryankemper@cumin2003 START - Cookbook sre.dns.wipe-cache an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [06:00:16] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1209.eqiad.wmnet 15.53.64.10.in-addr.arpa 5.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [06:00:17] !log ryankemper@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1209 [06:00:37] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1209 [06:00:38] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1209 [06:00:39] !log ryankemper@cumin2003 START - Cookbook sre.hosts.move-vlan for host an-worker1227 [06:00:43] !log ryankemper@cumin2003 START - Cookbook sre.dns.netbox [06:06:13] !log ryankemper@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" [06:06:17] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1227 - ryankemper@cumin2003" [06:06:17] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [06:06:17] !log ryankemper@cumin2003 START - Cookbook sre.dns.wipe-cache an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [06:06:20] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1227.eqiad.wmnet 19.53.64.10.in-addr.arpa 9.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [06:06:21] !log ryankemper@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1227 [06:07:26] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1227 [06:07:27] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1227 [06:07:47] !log ryankemper@cumin2003 START - Cookbook sre.hosts.move-vlan for host an-worker1228 [06:07:50] !log ryankemper@cumin2003 START - Cookbook sre.dns.netbox [06:08:45] (03CR) 10Ryan Kemper: [C:03+1] hieradata: enable dumps-nfs.w.o usage in production [puppet] - 10https://gerrit.wikimedia.org/r/1324597 (https://phabricator.wikimedia.org/T432212) (owner: 10Filippo Giunchedi) [06:09:58] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage [06:12:42] !log ryankemper@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" [06:12:46] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1228 - ryankemper@cumin2003" [06:12:46] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [06:12:46] !log ryankemper@cumin2003 START - Cookbook sre.dns.wipe-cache an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [06:12:49] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1228.eqiad.wmnet 20.53.64.10.in-addr.arpa 0.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [06:12:50] !log ryankemper@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1228 [06:13:22] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1175.eqiad.wmnet with reason: host reimage [06:14:09] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1228 [06:14:09] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1228 [06:14:25] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage [06:18:08] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1209.eqiad.wmnet with reason: host reimage [06:21:39] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage [06:23:26] FIRING: [2x] GitLabRestoreStale: GitLab - replica gitlab1003:0 has not completed a restore in 30h - https://wikitech.wikimedia.org/wiki/GitLab/Backup_and_Restore - https://grafana.wikimedia.org/d/R_1IvBZnz/gitlab-omnibus-overview - https://alerts.wikimedia.org/?q=alertname%3DGitLabRestoreStale [06:25:04] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1227.eqiad.wmnet with reason: host reimage [06:27:55] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage [06:31:12] PROBLEM - Host an-worker1228 is DOWN: PING CRITICAL - Packet loss = 100% [06:33:11] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1228.eqiad.wmnet with reason: host reimage [06:36:14] RECOVERY - Host an-worker1228 is UP: PING OK - Packet loss = 0%, RTA = 0.28 ms [06:36:56] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1175.eqiad.wmnet with OS bookworm [06:38:10] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1209.eqiad.wmnet with OS bookworm [06:41:32] (03CR) 10Fabfur: "I don't really think using mtail for this kind of analysis would be efficient, especially on CDN cache hosts, I'd prefer eventually to run" [puppet] - 10https://gerrit.wikimedia.org/r/1324778 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [06:42:22] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12214971 (10MoritzMuehlenhoff) >>! In T434681#12212493, @Jhancock.wm wrote: > direct quote from support *sigh*, it can't be some kind of OS-level issue. But anyway, the host i... [06:43:33] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12214972 (10MoritzMuehlenhoff) p:05High→03Medium And since the prio was still high, which I missed. this isn't urgent at all anymore, any time next week would be good [06:44:48] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1227.eqiad.wmnet with OS bookworm [06:53:23] 10ops-codfw, 06SRE, 10Data-Persistence-Backup, 10database-backups, and 3 others: db2201 memory failure (was: both mysql instances crashed in the last few days) - https://phabricator.wikimedia.org/T434532#12214982 (10jcrespo) @Jhancock.wm No worries. Host has notifications disabled, service has been migrate... [06:53:48] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1228.eqiad.wmnet with OS bookworm [06:57:23] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12214985 (10jcrespo) There is no rush on it being stable- I will be on vacations for a month and it is currently not being in use. But it would be nice to try to save it back up, as backup se... [07:00:05] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260814T0700) [07:04:30] !log jayme@cumin1003 START - Cookbook sre.hosts.reimage for host kubestagemaster2005.codfw.wmnet with OS trixie [07:22:37] (03PS1) 10Arnaudb: gitlab: publish the restore version-mismatch state [puppet] - 10https://gerrit.wikimedia.org/r/1325773 (https://phabricator.wikimedia.org/T425441) [07:22:40] (03CR) 10Arnaudb: [C:03+2] gitlab: publish the restore version-mismatch state [puppet] - 10https://gerrit.wikimedia.org/r/1325773 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [07:23:08] !log jayme@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage [07:23:22] FIRING: [2x] GitLabReplicaDataStale: GitLab - replica gitlab1003:0 serves data older than 30h - https://wikitech.wikimedia.org/wiki/GitLab/Backup_and_Restore - https://grafana.wikimedia.org/d/R_1IvBZnz/gitlab-omnibus-overview - https://alerts.wikimedia.org/?q=alertname%3DGitLabReplicaDataStale [07:24:01] (03CR) 10JMeybohm: [C:03+1] CI: fix comment for _helmfile_build [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320166 (https://phabricator.wikimedia.org/T388390) (owner: 10Kamila Součková) [07:24:30] (03PS1) 10Arnaudb: gitlab: distinguish blocked restores from stale restores [alerts] - 10https://gerrit.wikimedia.org/r/1325774 (https://phabricator.wikimedia.org/T425441) [07:26:51] (03Merged) 10jenkins-bot: gitlab: distinguish blocked restores from stale restores [alerts] - 10https://gerrit.wikimedia.org/r/1325774 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [07:27:03] PROBLEM - Host kubestagemaster2005 is DOWN: PING CRITICAL - Packet loss = 100% [07:29:21] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on kubestagemaster2005.codfw.wmnet with reason: host reimage [07:29:48] (03PS1) 10Marostegui: instances.yaml: Remove db1152 [puppet] - 10https://gerrit.wikimedia.org/r/1325775 (https://phabricator.wikimedia.org/T434480) [07:32:05] RECOVERY - Host kubestagemaster2005 is UP: PING OK - Packet loss = 0%, RTA = 30.57 ms [07:32:55] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps2011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:38:37] (03CR) 10Marostegui: [C:03+2] instances.yaml: Remove db1152 [puppet] - 10https://gerrit.wikimedia.org/r/1325775 (https://phabricator.wikimedia.org/T434480) (owner: 10Marostegui) [07:39:42] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove db1152 from dbctl T434480', diff saved to https://phabricator.wikimedia.org/P96090 and previous config saved to /var/cache/conftool/dbconfig/20260814-073941-marostegui.json [07:39:47] T434480: decommission db1152.eqiad.wmnet - https://phabricator.wikimedia.org/T434480 [07:41:29] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1230.eqiad.wmnet with OS bookworm [07:41:50] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1229.eqiad.wmnet with OS bookworm [07:41:56] !log btullis@cumin1003 START - Cookbook sre.hosts.move-vlan for host an-worker1230 [07:44:58] btullis@cumin1003 reimage (PID 3909179) is awaiting input [07:45:59] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215058 (10BTullis) [07:46:57] RECOVERY - Dell PowerEdge or Supermicro Broadcom RAID Controller on an-worker1231 is OK: communication: 0 OK : controller: 0 OK : physical_disk: 0 OK : virtual_disk: 0 OK : bbu: 0 OK : enclosure: 0 OK https://wikitech.wikimedia.org/wiki/PERCCli%23Monitoring [07:48:05] !log btullis@cumin1003 START - Cookbook sre.dns.netbox [07:48:08] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215059 (10BTullis) Fixed up an-worker1231 ` Physical Drives = 14 PD LIST : ======= ------------------------------------------------------------------------... [07:48:20] FIRING: [2x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:51:03] !log jayme@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host kubestagemaster2005.codfw.wmnet with OS trixie [07:51:55] (03CR) 10Klausman: [C:03+1] dse-k8s-codfw: remove entry for unused host [puppet] - 10https://gerrit.wikimedia.org/r/1325566 (https://phabricator.wikimedia.org/T434793) (owner: 10Bking) [07:54:20] btullis@cumin1003 reimage (PID 3909179) is awaiting input [07:55:32] !log btullis@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" [07:55:36] !log btullis@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1230 - btullis@cumin1003" [07:55:37] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [07:55:37] !log btullis@cumin1003 START - Cookbook sre.dns.wipe-cache an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [07:55:41] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1230.eqiad.wmnet 23.53.64.10.in-addr.arpa 3.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [07:55:42] !log btullis@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1230 [07:58:29] !log btullis@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1230 [07:58:29] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1230 [07:59:40] !log btullis@cumin1003 START - Cookbook sre.hosts.move-vlan for host an-worker1229 [08:00:28] !log tappof@cumin1003 START - Cookbook sre.hosts.reboot-single for host titan2002.codfw.wmnet [08:02:45] (03PS6) 10Dreamy Jazz: Register the mediawiki.wikimedia_antiabuse.content_policy_score stream [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321579 (https://phabricator.wikimedia.org/T432848) (owner: 10Mpostoronca) [08:02:47] btullis@cumin1003 reimage (PID 3909199) is awaiting input [08:03:31] (03PS1) 10Marostegui: db2209: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1325841 (https://phabricator.wikimedia.org/T431952) [08:05:00] (03CR) 10Marostegui: [C:03+2] db2209: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1325841 (https://phabricator.wikimedia.org/T431952) (owner: 10Marostegui) [08:05:15] !log btullis@cumin1003 START - Cookbook sre.dns.netbox [08:05:27] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db2209: Pool back [08:06:09] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, 13Patch-For-Review: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12215087 (10Marostegui) 05Open→03Resolved Thanks @Jhancock.wm - I've repooled the host. [08:07:58] !log tappof@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2002.codfw.wmnet [08:08:12] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:08:20] FIRING: [4x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:10:01] !log btullis@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1201.eqiad.wmnet with reason: Fixing a disk [08:11:05] !log btullis@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" [08:11:10] !log btullis@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1229 - btullis@cumin1003" [08:11:10] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [08:11:11] !log btullis@cumin1003 START - Cookbook sre.dns.wipe-cache an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [08:11:14] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1229.eqiad.wmnet 22.53.64.10.in-addr.arpa 2.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [08:11:15] !log btullis@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1229 [08:11:43] !log btullis@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1229 [08:11:43] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1229 [08:12:34] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage [08:19:12] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1230.eqiad.wmnet with reason: host reimage [08:21:09] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet [08:24:57] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet [08:25:40] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage [08:25:53] !log tappof@cumin1003 START - Cookbook sre.hosts.reboot-single for host titan1002.eqiad.wmnet [08:28:23] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1229.eqiad.wmnet with reason: host reimage [08:31:41] !log tappof@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1002.eqiad.wmnet [08:32:17] 06SRE, 06Infrastructure-Foundations, 10netops: Upgrade JunOS on pfw1-eqiad and pfw1-codfw - https://phabricator.wikimedia.org/T434865 (10cmooney) 03NEW p:05Triage→03Low [08:33:20] FIRING: [4x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:34:07] !log fceratto@cumin1003 START - Cookbook sre.mysql.decommission [08:34:28] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db1151.eqiad.wmnet [08:34:46] 06SRE, 06Infrastructure-Foundations, 10netops: Upgrade JunOS on pfw1-eqiad and pfw1-codfw - https://phabricator.wikimedia.org/T434865#12215165 (10cmooney) [08:35:01] 06SRE, 06Infrastructure-Foundations, 10netops: Upgrade JunOS on pfw1-eqiad and pfw1-codfw - https://phabricator.wikimedia.org/T434865#12215168 (10cmooney) [08:35:48] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215172 (10BTullis) [08:38:00] (03PS1) 10Arnaudb: gitlab: discard firewall throttling on the primary [puppet] - 10https://gerrit.wikimedia.org/r/1325844 (https://phabricator.wikimedia.org/T425441) [08:39:39] FIRING: TransitBGPDown: Transit BGP session down between cr2-esams and Hurricane Electric (2001:7f8:13::a500:6939:1) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=esams&var-device=cr2-esams:9804&var-bgp_group=Transit6&var-bgp_neighbor=Hurricane+Electric - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPD [08:40:28] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [08:40:52] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host an-worker1230.eqiad.wmnet with OS bookworm [08:43:07] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet [08:44:39] FIRING: [2x] TransitBGPDown: Transit BGP session down between cr2-esams and Hurricane Electric (193.239.116.14) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [08:46:32] fceratto@cumin1003 decommission (PID 3920087) is awaiting input [08:46:56] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet [08:47:45] !log tappof@cumin1003 START - Cookbook sre.hosts.reboot-single for host titan2001.codfw.wmnet [08:48:12] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:48:28] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866 (10MoritzMuehlenhoff) 03NEW [08:48:46] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1229.eqiad.wmnet with OS bookworm [08:49:39] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215210 (10BTullis) Fixed up an-worker1201. Booted into emergency mode. ` You are in emergency mode. AfterGive root password for maintenance (or press Control... [08:50:32] !log btullis@cumin1003 START - Cookbook sre.hosts.remove-downtime for an-worker1201.eqiad.wmnet [08:50:33] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1201.eqiad.wmnet [08:50:35] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2209: Pool back [08:51:04] btullis@cumin1003: Failed to log message to wiki. Somebody should check the error logs. [08:52:43] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12215227 (10MoritzMuehlenhoff) >>! In T434268#12205495, @VRiley-WMF wrote: > @andrea.denisse is there an estimated day th... [08:53:12] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:53:35] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1152.eqiad.wmnet with OS bookworm [08:53:41] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1231.eqiad.wmnet with OS bookworm [08:53:44] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1232.eqiad.wmnet with OS bookworm [08:54:02] !log btullis@cumin1003 START - Cookbook sre.hosts.move-vlan for host an-worker1152 [08:54:10] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [08:55:23] !log btullis@cumin1003 START - Cookbook sre.dns.netbox [08:55:44] !log tappof@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan2001.codfw.wmnet [08:55:54] PROBLEM - ganeti-noded running on ganeti1034 is CRITICAL: PROCS CRITICAL: 3 processes with UID = 0 (root), command name ganeti-noded https://wikitech.wikimedia.org/wiki/Ganeti [08:56:54] RECOVERY - ganeti-noded running on ganeti1034 is OK: PROCS OK: 2 processes with UID = 0 (root), command name ganeti-noded https://wikitech.wikimedia.org/wiki/Ganeti [08:57:14] fceratto@cumin1003 decommission (PID 3920087) is awaiting input [08:58:12] FIRING: [5x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:58:20] FIRING: [4x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:00:29] (03PS1) 10Arnaudb: sre.gitlab: silence restore staleness alerts during upgrades [cookbooks] - 10https://gerrit.wikimedia.org/r/1325846 (https://phabricator.wikimedia.org/T425441) [09:00:42] (03CR) 10Arnaudb: [C:03+2] sre.gitlab: silence restore staleness alerts during upgrades [cookbooks] - 10https://gerrit.wikimedia.org/r/1325846 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [09:01:24] btullis@cumin1003 reimage (PID 3936526) is awaiting input [09:02:04] (03PS1) 10Arnaudb: Revert "sre.gitlab: silence restore staleness alerts during upgrades" [cookbooks] - 10https://gerrit.wikimedia.org/r/1325847 [09:02:17] (03CR) 10Arnaudb: [V:03+2 C:03+2] Revert "sre.gitlab: silence restore staleness alerts during upgrades" [cookbooks] - 10https://gerrit.wikimedia.org/r/1325847 (owner: 10Arnaudb) [09:03:09] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1151.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [09:03:09] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:03:10] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1151.eqiad.wmnet [09:04:02] (03PS1) 10Arnaudb: Revert^2 "sre.gitlab: silence restore staleness alerts during upgrades" [cookbooks] - 10https://gerrit.wikimedia.org/r/1325849 [09:04:10] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215249 (10BTullis) [09:04:57] 06SRE, 06DC-Ops, 06ServiceOps: Restore kubestagemaster2005 to service - https://phabricator.wikimedia.org/T434844#12215250 (10JMeybohm) 05Open→03Resolved a:03JMeybohm I did the obvious thing of running reimage again and it worked. The etcd cluster came up fine, I updated BGP state in netbox, ran ho... [09:05:27] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215254 (10BTullis) Fixed up an-worker1205. ` Physical Drives = 14 PD LIST : ======= -----------------------------------------------------------------------... [09:06:10] fceratto@cumin1003 decommission (PID 3920087) is awaiting input [09:06:40] !log btullis@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" [09:06:44] !log btullis@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1152 - btullis@cumin1003" [09:06:45] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:06:45] !log btullis@cumin1003 START - Cookbook sre.dns.wipe-cache an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:06:49] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1152.eqiad.wmnet 16.53.64.10.in-addr.arpa 6.1.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:06:49] !log btullis@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1152 [09:06:56] RECOVERY - Dell PowerEdge or Supermicro Broadcom RAID Controller on an-worker1205 is OK: communication: 0 OK : controller: 0 OK : physical_disk: 0 OK : virtual_disk: 0 OK : bbu: 0 OK : enclosure: 0 OK https://wikitech.wikimedia.org/wiki/PERCCli%23Monitoring [09:07:19] (03CR) 10Arnaudb: [C:03+2] Revert^2 "sre.gitlab: silence restore staleness alerts during upgrades" [cookbooks] - 10https://gerrit.wikimedia.org/r/1325849 (owner: 10Arnaudb) [09:08:23] !log btullis@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1152 [09:08:23] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1152 [09:08:27] !log btullis@cumin1003 START - Cookbook sre.hosts.move-vlan for host an-worker1231 [09:08:53] (03PS1) 10Federico Ceratto: site.pp, db1151.yaml: decommission db1151 [puppet] - 10https://gerrit.wikimedia.org/r/1325850 (https://phabricator.wikimedia.org/T434538) [09:09:22] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215260 (10BTullis) Fixed up an-worker1207. ` Physical Drives = 14 PD LIST : ======= -----------------------------------------------------------------------... [09:09:33] !log btullis@cumin1003 START - Cookbook sre.dns.netbox [09:09:35] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215262 (10BTullis) [09:10:11] (03CR) 10Filippo Giunchedi: [C:03+1] O:dumps::distribution::server: Remove duplicate profiles [puppet] - 10https://gerrit.wikimedia.org/r/1325488 (owner: 10Majavah) [09:10:26] (03Merged) 10jenkins-bot: Revert^2 "sre.gitlab: silence restore staleness alerts during upgrades" [cookbooks] - 10https://gerrit.wikimedia.org/r/1325849 (owner: 10Arnaudb) [09:10:59] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org [09:11:35] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12215282 (10MoritzMuehlenhoff) p:05Triage→03Medium [09:12:43] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215283 (10BTullis) Fixed up an-worker1208 ` btullis@an-worker1208:~$ sudo blkid|grep -c hadoop- 11 Physical Drives = 14 PD LIST : ======= ----------------... [09:12:59] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215284 (10BTullis) [09:14:19] !log btullis@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" [09:14:24] !log btullis@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1231 - btullis@cumin1003" [09:14:24] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:14:25] !log btullis@cumin1003 START - Cookbook sre.dns.wipe-cache an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:14:25] !log tappof@cumin1003 START - Cookbook sre.hosts.reboot-single for host titan1001.eqiad.wmnet [09:14:28] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1231.eqiad.wmnet 24.53.64.10.in-addr.arpa 4.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:14:29] !log btullis@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1231 [09:16:43] !log btullis@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1231 [09:16:43] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1231 [09:17:15] !log btullis@cumin1003 START - Cookbook sre.hosts.move-vlan for host an-worker1232 [09:17:28] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org [09:17:30] (03CR) 10Filippo Giunchedi: "Very nice, thank you" [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [09:17:44] !log btullis@cumin1003 START - Cookbook sre.dns.netbox [09:18:33] (03PS4) 10Filippo Giunchedi: openstack: export compute service info to prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1313958 (https://phabricator.wikimedia.org/T284747) [09:22:05] !log btullis@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" [09:22:10] !log btullis@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host an-worker1232 - btullis@cumin1003" [09:22:10] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:22:11] !log btullis@cumin1003 START - Cookbook sre.dns.wipe-cache an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:22:12] !log tappof@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host titan1001.eqiad.wmnet [09:22:14] !log btullis@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) an-worker1232.eqiad.wmnet 25.53.64.10.in-addr.arpa 5.2.0.0.3.5.0.0.4.6.0.0.0.1.0.0.8.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [09:22:15] !log btullis@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host an-worker1232 [09:22:38] !log btullis@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host an-worker1232 [09:22:38] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host an-worker1232 [09:23:12] FIRING: [4x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:23:20] FIRING: [4x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:24:03] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org [09:25:12] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage [09:25:33] !log tappof@cumin1003 START - Cookbook sre.hosts.reboot-single for host centrallog1002.eqiad.wmnet [09:28:00] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1152.eqiad.wmnet with reason: host reimage [09:29:10] FIRING: BFDdown: BFD session down between cr2-eqiad and 10.64.16.86 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:30:21] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage [09:30:44] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org [09:32:30] !log tappof@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog1002.eqiad.wmnet [09:33:19] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org [09:33:20] FIRING: [6x] ProbeDown: Service centrallog1002:6514 has failed probes (tcp_rsyslog_receiver_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:33:54] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1231.eqiad.wmnet with reason: host reimage [09:34:10] RESOLVED: [2x] BFDdown: BFD session down between cr1-eqiad and 10.64.16.86 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:34:14] (03PS1) 10Arnaudb: tunnelencabulator: gitlab moved behind the text-lb CDN [debs/wmf-laptop] - 10https://gerrit.wikimedia.org/r/1325854 (https://phabricator.wikimedia.org/T425441) [09:34:42] !log jayme@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'sync'. [09:36:46] !log jayme@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. [09:36:55] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage [09:39:56] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org [09:41:32] !log tappof@cumin1003 START - Cookbook sre.hosts.reboot-single for host centrallog2002.codfw.wmnet [09:41:48] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org [09:42:43] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1232.eqiad.wmnet with reason: host reimage [09:45:55] FIRING: [4x] BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:46:00] (03CR) 10Blake: [C:03+2] kube-state-metrics: update to upstream 7.3.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [09:46:40] RESOLVED: BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:48:21] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org [09:48:23] !log tappof@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host centrallog2002.codfw.wmnet [09:49:14] !log `mwscript-k8s WikimediaAntiAbuse:BackfillAbuseReview.php --wiki=enwiki --start-timestamp="20260808000000" --end-timestamp="20260812120000" --sleep="5" --batch-size="50"` for T434688 [09:49:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:49:18] T434688: Create maintenance script to evaluate revisions performed between timestamps against content policies - https://phabricator.wikimedia.org/T434688 [09:50:09] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1152.eqiad.wmnet with OS bookworm [09:52:41] (03CR) 10Marostegui: [C:03+1] site.pp, db1151.yaml: decommission db1151 [puppet] - 10https://gerrit.wikimedia.org/r/1325850 (https://phabricator.wikimedia.org/T434538) (owner: 10Federico Ceratto) [09:53:20] FIRING: [4x] ProbeDown: Service centrallog2002:6514 has failed probes (tcp_rsyslog_receiver_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:53:22] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org [09:53:54] !log btullis@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" [09:55:56] (03Merged) 10jenkins-bot: kube-state-metrics: update to upstream 7.3.0. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1319069 (https://phabricator.wikimedia.org/T427405) (owner: 10Blake) [09:56:36] (03CR) 10Majavah: [V:03+1 C:03+2] O:dumps::distribution::server: Remove duplicate profiles [puppet] - 10https://gerrit.wikimedia.org/r/1325488 (owner: 10Majavah) [09:56:37] !log btullis@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003" [09:56:38] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1231.eqiad.wmnet with OS bookworm [09:56:44] (03CR) 10Federico Ceratto: [C:03+2] site.pp, db1151.yaml: decommission db1151 [puppet] - 10https://gerrit.wikimedia.org/r/1325850 (https://phabricator.wikimedia.org/T434538) (owner: 10Federico Ceratto) [09:57:23] (03CR) 10Arnaudb: "I forgot to mention, I tested tunnelencabulator locally. It is documented at https://wikitech.wikimedia.org/wiki/GitLab#Emergency_access" [debs/wmf-laptop] - 10https://gerrit.wikimedia.org/r/1325854 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [09:58:14] (03PS1) 10Majavah: cloudnfs: Remove scratch from PAWS [puppet] - 10https://gerrit.wikimedia.org/r/1325860 (https://phabricator.wikimedia.org/T415819) [09:58:38] !log fceratto@cumin1003 Removing db1151 from zarcillo T434538 [09:58:42] T434538: decommission db1151.eqiad.wmnet - https://phabricator.wikimedia.org/T434538 [09:58:43] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) [09:58:51] 10ops-eqiad, 06DC-Ops: decommission db1151.eqiad.wmnet - https://phabricator.wikimedia.org/T434538#12215362 (10ops-monitoring-bot) db1151 has been deleted from zarcillo [09:58:53] 10ops-eqiad, 06DC-Ops: decommission db1151.eqiad.wmnet - https://phabricator.wikimedia.org/T434538#12215364 (10ops-monitoring-bot) db1151 has been decommissioned by Data Persistence [09:58:55] 10ops-eqiad, 06DC-Ops: decommission db1151.eqiad.wmnet - https://phabricator.wikimedia.org/T434538#12215365 (10ops-monitoring-bot) a:05FCeratto-WMF→03None This host is ready for DC-Ops to decommission [10:00:01] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org [10:00:45] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org [10:02:20] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1232.eqiad.wmnet with OS bookworm [10:03:13] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Add timeouts and limits to HTTP frontend [puppet] - 10https://gerrit.wikimedia.org/r/1325427 (https://phabricator.wikimedia.org/T429930) (owner: 10Majavah) [10:03:32] !log cgoubert@cumin2003 START - Cookbook sre.ganeti.makevm for new host rdb-lock2003.codfw.wmnet [10:03:34] !log cgoubert@cumin2003 START - Cookbook sre.dns.netbox [10:07:20] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org [10:07:21] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1153.eqiad.wmnet with OS bookworm [10:07:24] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm [10:07:27] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1187.eqiad.wmnet with OS bookworm [10:07:33] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1235.eqiad.wmnet with OS bookworm [10:08:38] 10ops-eqiad, 06DBA, 06DC-Ops: decommission db1151.eqiad.wmnet - https://phabricator.wikimedia.org/T434538#12215392 (10FCeratto-WMF) [10:09:27] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org [10:09:48] cgoubert@cumin2003 makevm (PID 1614955) is awaiting input [10:12:01] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [10:12:06] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [10:12:06] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:12:06] !log cgoubert@cumin2003 START - Cookbook sre.dns.wipe-cache rdb-lock2003.codfw.wmnet on all recursors [10:12:09] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) rdb-lock2003.codfw.wmnet on all recursors [10:12:40] !log cgoubert@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [10:12:44] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM rdb-lock2003.codfw.wmnet - cgoubert@cumin2003" [10:15:45] cgoubert@cumin2003 makevm (PID 1614955) is awaiting input [10:16:02] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org [10:16:53] !log cgoubert@cumin2003 START - Cookbook sre.hosts.reimage for host rdb-lock2003.codfw.wmnet with OS trixie [10:17:06] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12215405 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cgoubert@cumin2003 for host rdb-lock2003.codfw.wmn... [10:21:14] (03CR) 10Arnaudb: [C:03+2] ci: enhance ci-build-images script [puppet] - 10https://gerrit.wikimedia.org/r/1268594 (https://phabricator.wikimedia.org/T422488) (owner: 10Hashar) [10:21:53] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage [10:23:47] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage [10:24:08] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage [10:27:21] (03PS29) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) [10:29:05] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12215441 (10BTullis) Hi, sorry for the delay in getting back to you about this. I can't get access to the management interface, either. I tried... [10:29:28] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1187.eqiad.wmnet with reason: host reimage [10:30:24] (03CR) 10CI reject: [V:04-1] cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [10:32:55] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1235.eqiad.wmnet with reason: host reimage [10:35:36] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1153.eqiad.wmnet with reason: host reimage [10:38:19] !log cgoubert@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage [10:44:01] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on rdb-lock2003.codfw.wmnet with reason: host reimage [10:44:37] !log cgoubert@cumin2003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-codfw [10:44:41] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2001.codfw.wmnet [10:45:07] (03CR) 10Urbanecm: [C:03+1] "LGTM" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325527 (https://phabricator.wikimedia.org/T433783) (owner: 10Michael Große) [10:45:15] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2001.codfw.wmnet [10:49:25] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1187.eqiad.wmnet with OS bookworm [10:50:30] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T434668 [10:51:38] (03CR) 10Clément Goubert: "The script actually being used by `mw-script-k8s` is [here](https://gitlab.wikimedia.org/repos/releng/release/-/blob/74090d776f3047782e6d6" [puppet] - 10https://gerrit.wikimedia.org/r/1324348 (https://phabricator.wikimedia.org/T434567) (owner: 10Reedy) [10:51:38] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2001.codfw.wmnet [10:51:39] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2001.codfw.wmnet [10:51:44] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet [10:54:33] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1153.eqiad.wmnet with OS bookworm [10:56:46] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet [10:56:59] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12215510 (10BTullis) [10:57:09] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host rdb-lock2003.codfw.wmnet with OS trixie [10:57:09] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host rdb-lock2003.codfw.wmnet [10:57:24] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12215511 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cgoubert@cumin2003 for host rdb-lock2003.codfw.wmnet w... [10:58:12] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1235.eqiad.wmnet with OS bookworm [10:58:19] PROBLEM - PyBal backends health check on lvs2013 is CRITICAL: PYBAL CRITICAL - CRITICAL - inference-staging_30443: Servers ml-staging2001.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [10:59:19] RECOVERY - PyBal backends health check on lvs2013 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [11:00:04] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260814T0700) [11:00:05] jelto, arnoldokoth, mutante, and arnaudb: Your horoscope predicts another GitLab version upgrades deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260814T1100). [11:01:56] (03PS5) 10Clément Goubert: rdb-lock: Create puppet role redis::lock::instance [puppet] - 10https://gerrit.wikimedia.org/r/1324693 (https://phabricator.wikimedia.org/T434188) [11:01:56] (03PS5) 10Clément Goubert: site.pp: Configure rdb-lock instances [puppet] - 10https://gerrit.wikimedia.org/r/1324694 (https://phabricator.wikimedia.org/T434188) [11:02:27] (03CR) 10Clément Goubert: rdb-lock: Create puppet role redis::lock::instance (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324693 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [11:02:53] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet [11:02:54] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet [11:03:00] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2003.codfw.wmnet [11:03:15] (03CR) 10JMeybohm: [C:03+1] rdb-lock: Create puppet role redis::lock::instance [puppet] - 10https://gerrit.wikimedia.org/r/1324693 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [11:03:31] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2003.codfw.wmnet [11:03:45] (03CR) 10JMeybohm: [C:03+1] site.pp: Configure rdb-lock instances [puppet] - 10https://gerrit.wikimedia.org/r/1324694 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [11:03:52] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1181.eqiad.wmnet with OS bookworm [11:04:13] (03PS6) 10Clément Goubert: rdb-lock: Create puppet role redis::lock::instance [puppet] - 10https://gerrit.wikimedia.org/r/1324693 (https://phabricator.wikimedia.org/T434188) [11:04:16] (03PS6) 10Clément Goubert: site.pp: Configure rdb-lock instances [puppet] - 10https://gerrit.wikimedia.org/r/1324694 (https://phabricator.wikimedia.org/T434188) [11:05:24] (03PS1) 10Urbanecm: Growth: Remove now removed config variable [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325868 [11:05:25] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12215528 (10Clement_Goubert) [11:05:34] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 10ServiceOps-Datastores, and 3 others: eqiad/codfw: 3 trixie VMs for rdb-lock - https://phabricator.wikimedia.org/T434188#12215530 (10Clement_Goubert) 05In progress→03Resolved Resolving, VMs are now created, configuration followup in parent task [11:07:07] aokoth@cumin1003 aokoth: The backup on gitlab1004 is complete, ready to proceed with upgrade. [11:10:00] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2003.codfw.wmnet [11:10:01] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2003.codfw.wmnet [11:10:07] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2004.codfw.wmnet [11:10:07] aokoth@cumin1003 upgrade (PID 4003522) is awaiting input [11:10:41] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2004.codfw.wmnet [11:17:10] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2004.codfw.wmnet [11:17:11] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2004.codfw.wmnet [11:17:11] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-codfw [11:19:41] (03PS1) 10Clément Goubert: mw-on-k8s: Add egress to new rdb-lock hosts [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325872 (https://phabricator.wikimedia.org/T427999) [11:20:08] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T434668 [11:21:00] !log aokoth@cumin1003 START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org [11:21:14] (03CR) 10Ladsgroup: [C:03+2] "I'll deploy it early next week" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270141 (https://phabricator.wikimedia.org/T424495) (owner: 10TheDJ) [11:21:29] FIRING: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:21:38] (03CR) 10Clément Goubert: [C:03+2] rdb-lock: Create puppet role redis::lock::instance [puppet] - 10https://gerrit.wikimedia.org/r/1324693 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [11:22:34] (03PS1) 10Marostegui: instances.yaml: Remove db1153 [puppet] - 10https://gerrit.wikimedia.org/r/1325873 (https://phabricator.wikimedia.org/T434638) [11:22:49] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324694 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [11:23:29] PROBLEM - Host gitlab.wikimedia.org is DOWN: PING CRITICAL - Packet loss = 100% [11:23:32] (03CR) 10Marostegui: [C:03+2] instances.yaml: Remove db1153 [puppet] - 10https://gerrit.wikimedia.org/r/1325873 (https://phabricator.wikimedia.org/T434638) (owner: 10Marostegui) [11:24:21] RECOVERY - Host gitlab.wikimedia.org is UP: PING OK - Packet loss = 0%, RTA = 0.29 ms [11:24:50] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove db1153 from dbctl T434638', diff saved to https://phabricator.wikimedia.org/P96097 and previous config saved to /var/cache/conftool/dbconfig/20260814-112449-marostegui.json [11:24:55] T434638: decommission db1153.eqiad.wmnet - https://phabricator.wikimedia.org/T434638 [11:26:25] RESOLVED: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:27:37] !log aokoth@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org [11:28:12] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:28:18] (03Merged) 10jenkins-bot: Implement remaining, more rare, orientations [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270141 (https://phabricator.wikimedia.org/T424495) (owner: 10TheDJ) [11:30:16] 06SRE, 10SRE-Access-Requests: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877 (10MKrolik-WMF) 03NEW [11:30:22] (03PS1) 10Clément Goubert: external_services: Just in case, add redis-lock [puppet] - 10https://gerrit.wikimedia.org/r/1325875 (https://phabricator.wikimedia.org/T434188) [11:33:10] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps2011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:33:55] (03PS3) 10Ladsgroup: Upgrade to trixie [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1325499 (https://phabricator.wikimedia.org/T419815) [11:34:11] (03CR) 10Ladsgroup: Upgrade to trixie (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1325499 (https://phabricator.wikimedia.org/T419815) (owner: 10Ladsgroup) [11:38:49] (03CR) 10Filippo Giunchedi: [C:03+1] cloudnfs: Remove scratch from PAWS [puppet] - 10https://gerrit.wikimedia.org/r/1325860 (https://phabricator.wikimedia.org/T415819) (owner: 10Majavah) [11:39:03] (03PS1) 10Clément Goubert: ratelimit-media: Switch to per-minute fixed window [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325876 (https://phabricator.wikimedia.org/T414445) [11:41:55] (03CR) 10Majavah: [C:03+2] cloudnfs: Remove scratch from PAWS [puppet] - 10https://gerrit.wikimedia.org/r/1325860 (https://phabricator.wikimedia.org/T415819) (owner: 10Majavah) [11:42:51] (03CR) 10Ladsgroup: Upgrade to trixie (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1325499 (https://phabricator.wikimedia.org/T419815) (owner: 10Ladsgroup) [11:43:29] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324694 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [11:52:43] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12215639 (10klausman) I think the BMC has simply lost it's LAN config (from `impi-config --checkout`): ` Section Lan_Conf ## Possible va... [11:52:51] (03CR) 10Ladsgroup: "Moritz says upgrade to imagemagick 7 is quite a nice improvement." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1325499 (https://phabricator.wikimedia.org/T419815) (owner: 10Ladsgroup) [11:55:01] (03PS30) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) [11:56:25] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet [12:00:17] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet [12:00:54] (03PS1) 10GergesShamon: [arwiki] Enable restricted user page editing and grant edit permissions [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325878 (https://phabricator.wikimedia.org/T434878) [12:01:52] (03CR) 10Reedy: "I’m happy to make the change there too…" [puppet] - 10https://gerrit.wikimedia.org/r/1324348 (https://phabricator.wikimedia.org/T434567) (owner: 10Reedy) [12:02:25] (03PS31) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) [12:02:49] 06SRE, 06Infrastructure-Foundations, 07LDAP: Migrate the r/w LDAP servers to Trixie and MDB storage - https://phabricator.wikimedia.org/T331699#12215662 (10MoritzMuehlenhoff) [12:03:23] (03PS1) 10Federico Ceratto: cookbooks.sre.mysql: fix cookbook import in test [cookbooks] - 10https://gerrit.wikimedia.org/r/1325879 (https://phabricator.wikimedia.org/T426613) [12:03:23] (03CR) 10Federico Ceratto: "A small fix" [cookbooks] - 10https://gerrit.wikimedia.org/r/1325879 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [12:03:41] !log jmm@cumin2003 START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie [12:03:53] 06SRE, 06Infrastructure-Foundations, 07LDAP: Migrate the r/w LDAP servers to Trixie and MDB storage - https://phabricator.wikimedia.org/T331699#12215664 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin2003 for host ldap-rw1001.wikimedia.org with OS trixie [12:04:05] (03CR) 10Federico Ceratto: "This should be ready for review. I used it for https://phabricator.wikimedia.org/T434538" [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [12:07:23] (03CR) 10Clément Goubert: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1324694 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [12:08:12] 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: BFD session fails from Anycast hosts over IPv6 on boot - https://phabricator.wikimedia.org/T434806#12215670 (10cmooney) >>! In T434806#12213146, @BBlack wrote: > 1. That we should switch our dependency above from `network.target` to `network-online.... [12:09:00] (03CR) 10Clément Goubert: [C:03+1] "Yeah I was just pointing out that there's two different places where this needs fixing" [puppet] - 10https://gerrit.wikimedia.org/r/1324348 (https://phabricator.wikimedia.org/T434567) (owner: 10Reedy) [12:14:44] 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: BFD session fails from Anycast hosts over IPv6 on boot - https://phabricator.wikimedia.org/T434806#12215678 (10cmooney) Regarding the patch I submitted and the change from `Requires=` to `Wants=` this seemed sensible, though I am not 100% sure. Fro... [12:14:55] !log cgoubert@cumin2003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-worker-eqiad [12:14:59] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1003.eqiad.wmnet [12:15:09] (03PS1) 10Cathal Mooney: Anycast-healthchecker: wait until network-online target is reached [puppet] - 10https://gerrit.wikimedia.org/r/1325881 (https://phabricator.wikimedia.org/T434806) [12:16:38] (03CR) 10JMeybohm: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1325875 (https://phabricator.wikimedia.org/T434188) (owner: 10Clément Goubert) [12:17:56] !log jmm@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage [12:20:06] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1003.eqiad.wmnet [12:22:01] 06SRE, 06Infrastructure-Foundations, 10netops, 13Patch-For-Review: Problems using Homer with SSH Agent on Debian 13 - https://phabricator.wikimedia.org/T434773#12215699 (10cmooney) [12:22:02] (03CR) 10JMeybohm: ratelimit-media: Switch to per-minute fixed window (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325876 (https://phabricator.wikimedia.org/T414445) (owner: 10Clément Goubert) [12:23:32] (03CR) 10Marostegui: cookbooks/sre/mysql/decommission: add cookbook (033 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [12:24:10] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage [12:26:12] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1003.eqiad.wmnet [12:26:13] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1003.eqiad.wmnet [12:26:19] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1004.eqiad.wmnet [12:27:10] (03CR) 10Marostegui: [C:03+1] cookbooks.sre.mysql: fix cookbook import in test [cookbooks] - 10https://gerrit.wikimedia.org/r/1325879 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [12:27:36] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1004.eqiad.wmnet [12:27:40] (03CR) 10Clément Goubert: ratelimit-media: Switch to per-minute fixed window (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325876 (https://phabricator.wikimedia.org/T414445) (owner: 10Clément Goubert) [12:28:16] (03PS2) 10Clément Goubert: ratelimit-media: Switch to per-minute fixed window [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325876 (https://phabricator.wikimedia.org/T414445) [12:28:40] (03CR) 10JMeybohm: [C:03+1] ratelimit-media: Switch to per-minute fixed window [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325876 (https://phabricator.wikimedia.org/T414445) (owner: 10Clément Goubert) [12:30:35] (03PS3) 10Blake: mediawiki: Refactor lamp.deployment into component containers. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1313111 (https://phabricator.wikimedia.org/T417800) [12:31:45] (03CR) 10Clément Goubert: [C:03+2] ratelimit-media: Switch to per-minute fixed window [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325876 (https://phabricator.wikimedia.org/T414445) (owner: 10Clément Goubert) [12:33:52] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1004.eqiad.wmnet [12:33:53] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1004.eqiad.wmnet [12:33:59] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1005.eqiad.wmnet [12:34:28] (03Merged) 10jenkins-bot: ratelimit-media: Switch to per-minute fixed window [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325876 (https://phabricator.wikimedia.org/T414445) (owner: 10Clément Goubert) [12:35:23] !log cgoubert@deploy1003 helmfile [eqiad] START helmfile.d/services/ratelimit: apply [12:35:45] !log cgoubert@deploy1003 helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply [12:35:51] !log cgoubert@deploy1003 helmfile [codfw] START helmfile.d/services/ratelimit: apply [12:36:12] !log cgoubert@deploy1003 helmfile [codfw] DONE helmfile.d/services/ratelimit: apply [12:39:06] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1005.eqiad.wmnet [12:45:31] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1005.eqiad.wmnet [12:45:33] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1005.eqiad.wmnet [12:45:38] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestage1006.eqiad.wmnet [12:45:56] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie [12:46:07] 06SRE, 06Infrastructure-Foundations, 07LDAP: Migrate the r/w LDAP servers to Trixie and MDB storage - https://phabricator.wikimedia.org/T331699#12215839 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin2003 for host ldap-rw1001.wikimedia.org with OS trixie completed: - ldap-rw1... [12:46:20] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage1006.eqiad.wmnet [12:52:52] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestage1006.eqiad.wmnet [12:52:53] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage1006.eqiad.wmnet [12:52:53] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-worker-eqiad [12:53:20] (03CR) 10Cmelo: [C:03+1] tables-catalog: Drop ce_worklist_articles [puppet] - 10https://gerrit.wikimedia.org/r/1324325 (https://phabricator.wikimedia.org/T426102) (owner: 10Daimona Eaytoy) [12:53:43] (03CR) 10JMeybohm: [C:03+1] "We have https://wikitech.wikimedia.org/wiki/Kubernetes/Deployment_Charts as well as README.md and helmfile.d/services/README.MD to learn a" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325451 (https://phabricator.wikimedia.org/T348856) (owner: 10Ladsgroup) [12:54:10] !log cgoubert@cumin2003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-codfw [12:54:13] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2003.codfw.wmnet [12:54:14] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2003.codfw.wmnet [12:59:19] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2003.codfw.wmnet [12:59:20] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2003.codfw.wmnet [12:59:26] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2004.codfw.wmnet [12:59:26] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2004.codfw.wmnet [13:01:48] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet [13:02:55] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps2011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:04:30] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2004.codfw.wmnet [13:04:32] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2004.codfw.wmnet [13:04:37] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster2005.codfw.wmnet [13:04:37] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster2005.codfw.wmnet [13:05:37] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet [13:09:45] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster2005.codfw.wmnet [13:09:47] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster2005.codfw.wmnet [13:09:47] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-codfw [13:11:49] (03CR) 10Aklapper: Add two new locale files (032 comments) [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1325535 (https://phabricator.wikimedia.org/T412651) (owner: 10Pppery) [13:13:43] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12216006 (10JMeybohm) I'm tentatively closing this. If you still experience issues accessing superset, please feel free to reopen. Thanks! [13:13:55] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-private-users for aramirezwmf - https://phabricator.wikimedia.org/T433999#12216007 (10JMeybohm) 05Open→03Resolved a:03JMeybohm [13:20:11] (03CR) 10Aklapper: [V:03+2 C:03+2] "Thanks!" [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1325534 (owner: 10Pppery) [13:21:52] 06SRE, 10SRE-Access-Requests, 10SRE-tools, 10Cumin, and 2 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12216045 (10JMeybohm) >>! In T433409#12214073, @Jhancock.wm wrote: > Sorry to reopen but i don't seem to be able to run downtime cookbooks. it pro... [13:28:17] (03CR) 10CWilliams: "This has a misleading merge message, at least as I read it. There is no mention on it being for the update-replication cookbook and reads " [cookbooks] - 10https://gerrit.wikimedia.org/r/1325879 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [13:29:43] 06SRE, 10SRE-Access-Requests: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12216102 (10JMeybohm) [13:35:14] 06SRE, 10SRE-Access-Requests, 06Data-Engineering: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12216134 (10JMeybohm) @HShaikh please sign off on this request [13:36:13] 06SRE, 10SRE-Access-Requests, 06Data-Engineering: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12216138 (10JMeybohm) [13:36:31] (03CR) 10Ssingh: "Thanks for the patch @cmooney@wikimedia.org! One question inline:" [puppet] - 10https://gerrit.wikimedia.org/r/1325881 (https://phabricator.wikimedia.org/T434806) (owner: 10Cathal Mooney) [13:41:38] 06SRE, 06Infrastructure-Foundations, 06ServiceOps, 06Traffic: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12216219 (10MoritzMuehlenhoff) >>! In T429175#12172713, @ssingh wrote: > `urldownloader[12]00[34].wikimedia.org` are now behind LVS as a l... [13:45:25] (03CR) 10Ssingh: [C:03+1] "Thanks, +1 on second thought." [puppet] - 10https://gerrit.wikimedia.org/r/1325881 (https://phabricator.wikimedia.org/T434806) (owner: 10Cathal Mooney) [13:50:33] (03PS2) 10Pppery: Add two new locale files [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1325535 (https://phabricator.wikimedia.org/T412651) [13:50:41] (03CR) 10Pppery: Add two new locale files (032 comments) [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1325535 (https://phabricator.wikimedia.org/T412651) (owner: 10Pppery) [13:53:35] FIRING: [2x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:53:43] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet [13:54:34] (03CR) 10Bking: [C:03+2] dse-k8s-codfw: remove entry for unused host [puppet] - 10https://gerrit.wikimedia.org/r/1325566 (https://phabricator.wikimedia.org/T434793) (owner: 10Bking) [13:57:42] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet [14:14:24] (03PS1) 10Hnowlan: ci: parallelise tests, run offline tests before publishing [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1325903 [14:14:57] (03CR) 10CI reject: [V:04-1] ci: parallelise tests, run offline tests before publishing [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1325903 (owner: 10Hnowlan) [14:15:17] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12216332 (10ssingh) @CDobbins: I think we can resolve this task; letting you do it since you worked on it, but just a reminder. [14:16:25] (03CR) 10Aklapper: "Oh and I'm afraid this also needs some lines in src/__phutil_library_map__.php via arc liberate? Sorry, I missed that initially :(" [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1325535 (https://phabricator.wikimedia.org/T412651) (owner: 10Pppery) [14:19:42] (03PS3) 10Pppery: Add two new locale files [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1325535 (https://phabricator.wikimedia.org/T412651) [14:20:14] (03CR) 10Pppery: "Yeah. I really should know better than this." [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1325535 (https://phabricator.wikimedia.org/T412651) (owner: 10Pppery) [14:21:15] FIRING: MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-jobrunner - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?panelId=18&fullscreen&orgId=1&var-datasource=codfw%20prometheus/ops - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [14:23:24] (03PS2) 10Federico Ceratto: cookbooks.sre.mysql.update-replication: fix cookbook import [cookbooks] - 10https://gerrit.wikimedia.org/r/1325879 (https://phabricator.wikimedia.org/T426613) [14:24:17] PROBLEM - LDAP -writable server- on ldap-rw1001 is CRITICAL: Could not search/find objectclasses in dc=wikimedia,dc=org https://wikitech.wikimedia.org/wiki/LDAP%23Troubleshooting [14:26:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-jobrunner - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [14:32:04] (03PS1) 10Kamila Součková: php8.5: initial release of 8.5 image stack [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 [14:33:09] (03PS2) 10Kamila Součková: php8.5: initial release of 8.5 image stack [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 [14:34:08] (03CR) 10Federico Ceratto: [C:03+2] cookbooks.sre.mysql.update-replication: fix cookbook import [cookbooks] - 10https://gerrit.wikimedia.org/r/1325879 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [14:35:10] (03PS3) 10Kamila Součková: php8.5: initial release of 8.5 image stack [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 (https://phabricator.wikimedia.org/T432988) [14:35:59] (03CR) 10Federico Ceratto: [V:03+2 C:03+2] cookbooks.sre.mysql.update-replication: fix cookbook import [cookbooks] - 10https://gerrit.wikimedia.org/r/1325879 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [14:36:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-jobrunner - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [14:36:17] (03CR) 10Kamila Součková: [V:03+2] "PS1 is just a verbatim copy (except changelogs), PS2 is where I changed 8.3 -> 8.5 . (PS3 is "I forgot the Bug trailer" :D)" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 (https://phabricator.wikimedia.org/T432988) (owner: 10Kamila Součková) [14:39:17] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host sretest2003.codfw.wmnet [14:41:15] RESOLVED: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-jobrunner - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [14:45:20] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2003.codfw.wmnet [14:50:32] (03CR) 10Scott French: [C:03+1] "Nice! This looks good." [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 (https://phabricator.wikimedia.org/T432988) (owner: 10Kamila Součková) [14:52:58] (03PS32) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) [14:53:29] (03CR) 10Federico Ceratto: cookbooks/sre/mysql/decommission: add cookbook (033 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1291952 (https://phabricator.wikimedia.org/T426613) (owner: 10Federico Ceratto) [14:54:47] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host sretest2006.codfw.wmnet [14:58:26] (03PS1) 10Majavah: cloudnfs: Remove TWL NFS configuration [puppet] - 10https://gerrit.wikimedia.org/r/1325912 (https://phabricator.wikimedia.org/T402054) [14:58:29] (03PS1) 10Majavah: cloudnfs: Remove fastcci NFS configuration [puppet] - 10https://gerrit.wikimedia.org/r/1325913 [15:01:11] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2006.codfw.wmnet [15:10:17] (03CR) 10Tsevener: "Sure that sounds good, thanks @swfrench@wikimedia.org!" [puppet] - 10https://gerrit.wikimedia.org/r/1315141 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [15:15:02] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12216690 (10Jhancock.wm) putting this on my list for today. Should be ready for y'all when you come back from the weekend. [15:18:12] !log dancy@deploy1003 Installing scap version "4.280.2" for 3 host(s) [15:18:36] (03PS1) 10Lerickson: Bump the WDQS chart 0.0.6 -> 0.0.7. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325917 (https://phabricator.wikimedia.org/T434854) [15:19:17] (03CR) 10Ssingh: "There was a test failing but otherwise I think this is ready to go now?" [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [15:20:08] !log dancy@deploy1003 Installation of scap version "4.280.2" completed for 3 hosts [15:20:45] !log dancy@deploy1003 Started scap sync-world: testing [15:22:35] !log cgoubert@cumin2003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-staging-master-eqiad [15:22:39] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1003.eqiad.wmnet [15:22:40] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1003.eqiad.wmnet [15:24:08] !log dancy@deploy1003 Finished scap sync-world: testing (duration: 03m 23s) [15:25:49] (03CR) 10Trueg: [C:03+1] Bump the WDQS chart 0.0.6 -> 0.0.7. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325917 (https://phabricator.wikimedia.org/T434854) (owner: 10Lerickson) [15:27:46] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1003.eqiad.wmnet [15:27:47] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1003.eqiad.wmnet [15:27:57] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1004.eqiad.wmnet [15:27:58] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1004.eqiad.wmnet [15:28:17] (03CR) 10Lerickson: [C:03+2] Bump the WDQS chart 0.0.6 -> 0.0.7. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325917 (https://phabricator.wikimedia.org/T434854) (owner: 10Lerickson) [15:28:27] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:30:18] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host sretest2009.codfw.wmnet [15:30:48] (03Merged) 10jenkins-bot: Bump the WDQS chart 0.0.6 -> 0.0.7. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325917 (https://phabricator.wikimedia.org/T434854) (owner: 10Lerickson) [15:32:53] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1004.eqiad.wmnet [15:32:55] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1004.eqiad.wmnet [15:33:00] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host kubestagemaster1005.eqiad.wmnet [15:33:01] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestagemaster1005.eqiad.wmnet [15:35:37] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest2009.codfw.wmnet [15:38:07] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host kubestagemaster1005.eqiad.wmnet [15:38:09] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestagemaster1005.eqiad.wmnet [15:38:09] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-staging-master-eqiad [15:47:48] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12216805 (10Jhancock.wm) 05Resolved→03Open I got a list of firmware back from the manufacturer after a meeting. [15:49:56] (03PS1) 10Ebernhardson: [DNM] Test unused variable [puppet] - 10https://gerrit.wikimedia.org/r/1325919 [15:50:32] (03CR) 10Ebernhardson: [V:03+1] "PCC SUCCESS (): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9228/console" [puppet] - 10https://gerrit.wikimedia.org/r/1325919 (owner: 10Ebernhardson) [15:52:02] (03CR) 10Ebernhardson: [V:03+1] "PCC SUCCESS (NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9230/console" [puppet] - 10https://gerrit.wikimedia.org/r/1325919 (owner: 10Ebernhardson) [15:53:28] (03PS1) 10Neriah: Remove sending email to legal team about rejected requests [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325920 (https://phabricator.wikimedia.org/T374053) [15:54:23] (03CR) 10Gehel: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1318274 (https://phabricator.wikimedia.org/T384998) (owner: 10Bking) [16:01:08] (03PS1) 10Bking: stat hosts: update partitioning to retain /srv data [puppet] - 10https://gerrit.wikimedia.org/r/1325921 (https://phabricator.wikimedia.org/T434562) [16:05:50] 10SRE-swift-storage, 10Ceph, 06Data-Persistence, 06DBA: Data persistance: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets - https://phabricator.wikimedia.org/T421719#12216921 (10Eevans) Apologies that this hasn't happened sooner, but with T357791 now imminent, these should all be done in... [16:05:59] 06SRE, 06Infrastructure-Foundations: Re-IP hosts running Cassandra to per-rack subnets in codfw row A and B. - https://phabricator.wikimedia.org/T354871#12216925 (10Eevans) Apologies that this hasn't happened sooner, but with T357791 now imminent, these should all be done in the coming weeks. [16:08:12] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:08:53] (03CR) 10Bking: [C:03+2] cirrus: Remove Cirrus frozen write monitors [puppet] - 10https://gerrit.wikimedia.org/r/1318274 (https://phabricator.wikimedia.org/T384998) (owner: 10Bking) [16:16:17] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: decommission db1151.eqiad.wmnet - https://phabricator.wikimedia.org/T434538#12216975 (10VRiley-WMF) a:03VRiley-WMF [16:16:24] (03PS1) 10Eevans: aptrepo: Create cassandra50 bookworm component [puppet] - 10https://gerrit.wikimedia.org/r/1325922 (https://phabricator.wikimedia.org/T357791) [16:16:56] PROBLEM - Host titan1002 is DOWN: PING CRITICAL - Packet loss = 100% [16:18:12] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:18:20] FIRING: [4x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:19:14] RECOVERY - Host titan1002 is UP: PING OK - Packet loss = 0%, RTA = 0.24 ms [16:23:20] FIRING: [4x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:26:43] (03PS1) 10Bking: apifeatureusage: stop specifying document type [puppet] - 10https://gerrit.wikimedia.org/r/1325928 (https://phabricator.wikimedia.org/T434906) [16:27:48] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1325928 (https://phabricator.wikimedia.org/T434906) (owner: 10Bking) [16:28:28] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Traffic: decommission lvs10[13-16] - https://phabricator.wikimedia.org/T433508#12217044 (10VRiley-WMF) [16:30:05] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Traffic: decommission lvs10[13-16] - https://phabricator.wikimedia.org/T433508#12217047 (10VRiley-WMF) I found that lvs1016 has previously been decommed back in July. I will be removing that one from this list. [16:30:18] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Traffic: decommission lvs10[13-16] - https://phabricator.wikimedia.org/T433508#12217049 (10VRiley-WMF) [16:31:47] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Traffic: decommission lvs10[13-16] - https://phabricator.wikimedia.org/T433508#12217055 (10VRiley-WMF) [16:32:08] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Traffic: decommission lvs10[13-16] - https://phabricator.wikimedia.org/T433508#12217058 (10VRiley-WMF) 05Open→03Resolved This has been completed [16:37:51] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: decommission db1151.eqiad.wmnet - https://phabricator.wikimedia.org/T434538#12217070 (10VRiley-WMF) [16:38:14] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: decommission db1151.eqiad.wmnet - https://phabricator.wikimedia.org/T434538#12217072 (10VRiley-WMF) 05In progress→03Resolved This has been completed [16:40:27] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 17 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1325878 (https://phabricator.wikimedia.org/T434878) (owner: 10GergesShamon) [16:45:25] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12217091 (10Jhancock.wm) the full story is that I'm gonna update the firmware on these smaller parts, generate a new tech support report, send it to them, they'll see if it resolves the is... [16:46:12] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12217092 (10VRiley-WMF) 05Open→03In progress Commencing working on this [16:48:26] PROBLEM - Host an-worker1147 is DOWN: PING CRITICAL - Packet loss = 100% [16:54:08] 10ops-eqiad, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T434926 (10phaultfinder) 03NEW [16:59:44] (03Abandoned) 10Ebernhardson: [DNM] Test unused variable [puppet] - 10https://gerrit.wikimedia.org/r/1325919 (owner: 10Ebernhardson) [17:04:02] (03CR) 10Ebernhardson: [C:03+1] "pcc delta looks like what i would expect. A review of the logstash-output-opensearch repo suggests this should do what we need with plugi" [puppet] - 10https://gerrit.wikimedia.org/r/1325928 (https://phabricator.wikimedia.org/T434906) (owner: 10Bking) [17:16:34] (03CR) 10Kamila Součková: [C:03+2] CI: fix comment for _helmfile_build [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320166 (https://phabricator.wikimedia.org/T388390) (owner: 10Kamila Součková) [17:22:27] (03PS4) 10Kamila Součková: php8.5: initial release of 8.5 image stack [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 (https://phabricator.wikimedia.org/T432988) [17:22:38] (03CR) 10Kamila Součková: "Excellent idea, thank you :D I did this, and besides fixing the memcached serializer and the wikidiff2 filename issue (thanks <3), the onl" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 (https://phabricator.wikimedia.org/T432988) (owner: 10Kamila Součková) [17:28:53] (03CR) 10Bking: [C:03+2] apifeatureusage: stop specifying document type [puppet] - 10https://gerrit.wikimedia.org/r/1325928 (https://phabricator.wikimedia.org/T434906) (owner: 10Bking) [17:31:04] RECOVERY - Host an-worker1147 is UP: PING OK - Packet loss = 0%, RTA = 0.33 ms [17:35:12] PROBLEM - Host an-worker1147 is DOWN: PING CRITICAL - Packet loss = 100% [17:38:34] PROBLEM - Router interfaces on mr1-eqsin is CRITICAL: CRITICAL: No response from remote host 103.102.166.128 for 1.3.6.1.2.1.2.2.1.8 with snmp version 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [17:39:36] RECOVERY - Router interfaces on mr1-eqsin is OK: OK: host 103.102.166.128, interfaces up: 38, down: 0, dormant: 0, excluded: 0, unused: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [17:43:34] PROBLEM - Router interfaces on mr1-eqsin is CRITICAL: CRITICAL: No response from remote host 103.102.166.128 for 1.3.6.1.2.1.2.2.1.8 with snmp version 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [17:44:32] RECOVERY - Router interfaces on mr1-eqsin is OK: OK: host 103.102.166.128, interfaces up: 38, down: 0, dormant: 0, excluded: 0, unused: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [17:45:04] RECOVERY - Host an-worker1147 is UP: PING OK - Packet loss = 0%, RTA = 0.26 ms [17:47:17] (03Merged) 10jenkins-bot: CI: fix comment for _helmfile_build [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320166 (https://phabricator.wikimedia.org/T388390) (owner: 10Kamila Součková) [17:56:18] 10ops-codfw, 06SRE, 10Data-Persistence-Backup, 10database-backups, and 3 others: db2201 memory failure (was: both mysql instances crashed in the last few days) - https://phabricator.wikimedia.org/T434532#12217280 (10Jhancock.wm) @jcrespo replacement has been completed. all yours! [17:57:17] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12217282 (10VRiley-WMF) 05In progress→03Open I tried the following. Rebooting the iDRAC - No change Swapped the cable with 2 other cables -... [18:05:37] 06SRE, 10Wikimedia-Mailing-lists: "You are doing that too often. Please try again later." during subscription a mailing list. - https://phabricator.wikimedia.org/T220914#12217301 (10Gaurav) We've been getting reports of this on https://lists.wikimedia.org/postorius/lists/wikimedia-us-nc.lists.wikimedia.org and... [18:08:21] FIRING: CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:13:21] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:17:23] (03PS4) 10Bking: WIP: stat hosts: update partitioning to retain /srv data [puppet] - 10https://gerrit.wikimedia.org/r/1325921 (https://phabricator.wikimedia.org/T434562) [18:17:38] (03PS5) 10Bking: stat hosts: update partitioning to retain /srv data [puppet] - 10https://gerrit.wikimedia.org/r/1325921 (https://phabricator.wikimedia.org/T434562) [18:19:56] (03CR) 10Bking: "Fixed. removing the WIP status" [puppet] - 10https://gerrit.wikimedia.org/r/1325921 (https://phabricator.wikimedia.org/T434562) (owner: 10Bking) [18:20:26] (03PS1) 10Ssingh: admin_ng: allow-urldownloaders: update urldownloader VIPs for egress [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325949 (https://phabricator.wikimedia.org/T429175) [18:23:43] 10ops-eqiad, 06SRE, 06DC-Ops, 10decommission-hardware, 06Traffic: decommission lvs10[13-16] - https://phabricator.wikimedia.org/T433508#12217352 (10ssingh) Thank you for the help and support @VRiley-WMF. [18:43:21] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:53:21] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:58:21] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [18:58:54] (03PS1) 10Bking: elasticsearch: remove non-working freeze/unfreeze code [software/spicerack] - 10https://gerrit.wikimedia.org/r/1325955 (https://phabricator.wikimedia.org/T433306) [18:59:19] (03PS2) 10Bking: elasticsearch: remove non-working freeze/unfreeze code [software/spicerack] - 10https://gerrit.wikimedia.org/r/1325955 (https://phabricator.wikimedia.org/T433306) [19:03:08] (03CR) 10Scott French: [C:03+1] "Thanks, Sukhbir! This looks good. I'm happy to work with you on Monday to deploy this to the respective wikikube clusters (which will also" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325949 (https://phabricator.wikimedia.org/T429175) (owner: 10Ssingh) [19:03:50] (03PS2) 10Bking: elasticsearch: remove non-working freeze/unfreeze code [software/spicerack] - 10https://gerrit.wikimedia.org/r/1325955 (https://phabricator.wikimedia.org/T433306) [19:10:57] (03CR) 10Lerickson: [C:03+2] Update Qlever to an image including http proxy support. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325512 (https://phabricator.wikimedia.org/T434321) (owner: 10Lerickson) [19:12:39] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T434926#12217474 (10VRiley-WMF) a:03VRiley-WMF [19:13:09] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T434926#12217475 (10VRiley-WMF) 05Open→03Resolved Closing this as the ticket for this is T433348 [19:13:19] (03Merged) 10jenkins-bot: Update Qlever to an image including http proxy support. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325512 (https://phabricator.wikimedia.org/T434321) (owner: 10Lerickson) [19:22:51] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12217505 (10Jhancock.wm) @MoritzMuehlenhoff updated the firmware i could [19:34:36] (03CR) 10Jdlrobson: Transition reading list experiment to instrument (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1259251 (https://phabricator.wikimedia.org/T421939) (owner: 10LorenMora) [19:40:10] FIRING: [2x] BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:43:20] FIRING: [4x] ProbeDown: Service centrallog2002:6514 has failed probes (tcp_rsyslog_receiver_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:45:10] RESOLVED: [2x] BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:45:55] FIRING: [3x] BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:48:20] FIRING: [6x] ProbeDown: Service centrallog1002:6514 has failed probes (tcp_rsyslog_receiver_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:50:40] RESOLVED: [4x] BFDdown: BFD session down between cr1-codfw and 10.192.16.35 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:56:44] PROBLEM - Host logging-hd2001 is DOWN: PING CRITICAL - Packet loss = 100% [19:58:12] RECOVERY - Host logging-hd2001 is UP: PING OK - Packet loss = 0%, RTA = 30.27 ms [20:10:56] PROBLEM - Host logging-sd2001 is DOWN: PING CRITICAL - Packet loss = 100% [20:14:12] RECOVERY - Host logging-sd2001 is UP: PING OK - Packet loss = 0%, RTA = 31.71 ms [20:18:27] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:20:42] (03CR) 10Scott French: "Thanks for the additional investigation! +1 to leaving the APCu tuning aside for now in favor of sorting that out on phabricator." [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 (https://phabricator.wikimedia.org/T432988) (owner: 10Kamila Součková) [20:25:44] PROBLEM - Host logging-hd2002 is DOWN: PING CRITICAL - Packet loss = 100% [20:26:00] PROBLEM - Host logging-sd2002 is DOWN: PING CRITICAL - Packet loss = 100% [20:28:12] RECOVERY - Host logging-hd2002 is UP: PING OK - Packet loss = 0%, RTA = 30.43 ms [20:28:12] RECOVERY - Host logging-sd2002 is UP: PING OK - Packet loss = 0%, RTA = 31.66 ms [20:40:58] PROBLEM - Host logging-hd2003 is DOWN: PING CRITICAL - Packet loss = 100% [20:41:08] PROBLEM - Host logging-sd2003 is DOWN: PING CRITICAL - Packet loss = 100% [20:43:12] RECOVERY - Host logging-sd2003 is UP: PING OK - Packet loss = 0%, RTA = 31.73 ms [20:43:12] RECOVERY - Host logging-hd2003 is UP: PING OK - Packet loss = 0%, RTA = 31.71 ms [20:57:12] PROBLEM - Host logging-hd2004 is DOWN: PING CRITICAL - Packet loss = 100% [20:57:12] PROBLEM - Host logging-sd2004 is DOWN: PING CRITICAL - Packet loss = 100% [20:57:13] RECOVERY - Host logging-hd2004 is UP: PING OK - Packet loss = 0%, RTA = 31.73 ms [20:57:15] RECOVERY - Host logging-sd2004 is UP: PING OK - Packet loss = 0%, RTA = 31.69 ms [21:20:44] PROBLEM - Host logging-hd2005 is DOWN: PING CRITICAL - Packet loss = 100% [21:22:12] RECOVERY - Host logging-hd2005 is UP: PING OK - Packet loss = 0%, RTA = 32.96 ms [21:39:23] (03PS1) 10Dduvall: buildx: Move all docker client calls to driver [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) [21:40:10] 06SRE, 06Content-Platform-Team, 10MediaWiki-extensions-PageAssessments: pageassessments-cleanup alerts tagged with defunct team - https://phabricator.wikimedia.org/T434959 (10Reedy) 03NEW [21:44:29] (03CR) 10CI reject: [V:04-1] buildx: Move all docker client calls to driver [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [22:06:04] (03PS2) 10Dduvall: buildx: Move all docker client calls to driver [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) [22:11:06] (03CR) 10CI reject: [V:04-1] buildx: Move all docker client calls to driver [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [22:16:00] (03PS3) 10Dduvall: buildx: Move all docker client calls to driver [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) [22:58:12] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [23:30:24] (03CR) 10Jasmine: [C:03+1] "Thanks Blake! LGTM" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1313111 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [23:37:30] (03PS1) 10BCornwall: varnish: Introduce ESI fragment support for users [puppet] - 10https://gerrit.wikimedia.org/r/1325993 (https://phabricator.wikimedia.org/T432842) [23:41:29] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1325994 [23:41:29] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1325994 (owner: 10TrainBranchBot) [23:46:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [23:48:35] FIRING: [2x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [23:51:31] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1325994 (owner: 10TrainBranchBot) [23:51:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [23:52:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1147:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1147 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [23:57:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1147:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1147 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown