[00:06:57] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet [00:06:58] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet [00:07:04] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet [00:07:39] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet [00:11:10] FIRING: [4x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:14:20] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet [00:14:21] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet [00:14:30] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet [00:35:48] FIRING: PuppetFailure: Puppet has failed on ml-serve1015:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [00:44:33] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet [00:53:15] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet [00:53:16] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet [00:53:21] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet [00:53:56] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet [01:00:21] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet [01:00:22] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet [01:00:28] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet [01:00:42] (03CR) 10Jasmine: "Thanks for flagging @swfrench@wikimedia.org, @ltoscano@wikimedia.org Will make a note to monitor this during switchover Act 2. (cc @slyngs" [puppet] - 10https://gerrit.wikimedia.org/r/1344678 (owner: 10Elukey) [01:08:57] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1344829 [01:08:57] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1344829 (owner: 10TrainBranchBot) [01:14:41] FIRING: ConfdResourceFailed: confd resource _etc_haproxy_conf.d_tls.cfg.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [01:18:28] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1344829 (owner: 10TrainBranchBot) [01:18:43] 10ops-codfw, 06DC-Ops: Power Supply - Status - issue on cloudbackup2003:9290 - https://phabricator.wikimedia.org/T439192 (10phaultfinder) 03NEW [01:18:44] 10ops-codfw, 06DC-Ops: Power Supply - Status - issue on cirrussearch2080:9290 - https://phabricator.wikimedia.org/T439194 (10phaultfinder) 03NEW [01:18:44] 10ops-codfw, 06DC-Ops: Power Supply - Status - issue on cirrussearch2079:9290 - https://phabricator.wikimedia.org/T439195 (10phaultfinder) 03NEW [01:18:46] 10ops-codfw, 06DC-Ops: Power Supply - Status - issue on wikikube-ctrl2001:9290 - https://phabricator.wikimedia.org/T439193 (10phaultfinder) 03NEW [01:19:49] 10ops-codfw, 06DC-Ops: Power Supply - Status - issue on logstash2036:9290 - https://phabricator.wikimedia.org/T439196 (10phaultfinder) 03NEW [01:30:30] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet [01:41:33] (03PS2) 10DLynch: editcheck-headless: Forward gRPC as HTTP/2 through the ingress [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344803 (https://phabricator.wikimedia.org/T436689) [01:41:33] (03PS1) 10DLynch: python-webapp: Update ingress.istio to 1.4.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344831 (https://phabricator.wikimedia.org/T436689) [01:41:35] (03PS1) 10DLynch: python-webapp releases: Remove ingress.staging and ingress.mlstaging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344832 (https://phabricator.wikimedia.org/T436689) [01:41:37] (03PS1) 10DLynch: ingress: Copy istio 1.4.0 to 1.4.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344833 (https://phabricator.wikimedia.org/T436689) [01:41:39] (03PS1) 10DLynch: ingress.istio: Add useClientProtocol [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344834 (https://phabricator.wikimedia.org/T436689) [01:41:41] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet [01:41:42] (03PS1) 10DLynch: python-webapp: Update ingress.istio to 1.4.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344835 (https://phabricator.wikimedia.org/T436689) [01:41:42] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet [01:41:42] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{dse-k8s-worker1*.eqiad.wmnet} and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad) [01:43:33] FIRING: [2x] KubernetesCalicoDown: dse-k8s-worker1028.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [01:45:24] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:49:52] (03Abandoned) 10DLynch: python-webapp: Update ingress.istio to 1.0.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344802 (https://phabricator.wikimedia.org/T436689) (owner: 10DLynch) [01:50:02] (03Abandoned) 10DLynch: ingress: Copy istio 1.0.3 to 1.0.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344800 (https://phabricator.wikimedia.org/T436689) (owner: 10DLynch) [01:50:23] (03Abandoned) 10DLynch: ingress.istio: Add useClientProtocol [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344801 (https://phabricator.wikimedia.org/T436689) (owner: 10DLynch) [02:00:35] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:08:14] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 38s) [02:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:40:23] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b12-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [04:11:26] FIRING: [4x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:35:48] FIRING: PuppetFailure: Puppet has failed on ml-serve1015:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [05:14:41] FIRING: ConfdResourceFailed: confd resource _etc_haproxy_conf.d_tls.cfg.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [05:21:40] !log [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched [05:21:41] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:40:36] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12362587 (10ayounsi) 05Resolved→03Open p:05Medium→03High There is an outstanding change on the switch for that host: ` Change for lsw1-e5-eqiad.mgmt.eqiad.wmnet:... [05:43:33] FIRING: KubernetesCalicoDown: wikikube-worker2280.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=codfw%20prometheus%2Fk8s&var-instance=wikikube-worker2280.codfw.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [05:45:24] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:52:19] (03PS1) 10Ryan Kemper: wdqs: exclude public main from DC switchover [puppet] - 10https://gerrit.wikimedia.org/r/1344974 (https://phabricator.wikimedia.org/T435443) [05:58:04] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [05:58:13] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260925T0600) [06:03:05] (03PS2) 10Ryan Kemper: wdqs: exclude public main from DC switchover [puppet] - 10https://gerrit.wikimedia.org/r/1344974 (https://phabricator.wikimedia.org/T435443) [06:06:10] FIRING: [4x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:07:36] (03PS3) 10Ryan Kemper: wdqs: exclude public main from DC switchover [puppet] - 10https://gerrit.wikimedia.org/r/1344974 (https://phabricator.wikimedia.org/T435443) [06:08:25] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1344974 (https://phabricator.wikimedia.org/T435443) (owner: 10Ryan Kemper) [06:14:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:16:41] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1344974 (https://phabricator.wikimedia.org/T435443) (owner: 10Ryan Kemper) [06:19:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:25:23] FIRING: [5x] GnmiInterfaceCountersDrop: asw1-b12-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [06:30:23] FIRING: [5x] GnmiInterfaceCountersDrop: asw1-b12-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [06:40:18] (03CR) 10Muehlenhoff: [C:03+2] Remove wmflib::wmf_php_version [puppet] - 10https://gerrit.wikimedia.org/r/1342215 (owner: 10Muehlenhoff) [06:52:17] 10ops-codfw, 06DBA, 06DC-Ops: db2226 PSU loss of redundancy - https://phabricator.wikimedia.org/T439205 (10FCeratto-WMF) 03NEW [07:00:04] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260925T0700) [07:25:18] PROBLEM - Host wikikube-worker1016 is DOWN: PING CRITICAL - Packet loss = 66%, RTA = 2455.22 ms [07:25:22] RECOVERY - Host wikikube-worker1016 is UP: PING OK - Packet loss = 0%, RTA = 103.42 ms [07:28:59] (03PS1) 10Muehlenhoff: Remove cumin1003 from RAPI access [puppet] - 10https://gerrit.wikimedia.org/r/1344980 (https://phabricator.wikimedia.org/T427897) [07:30:21] (03CR) 10Arnaudb: "thanks for the merge and for checking on gerrit2003 :-)" [puppet] - 10https://gerrit.wikimedia.org/r/1344598 (https://phabricator.wikimedia.org/T439104) (owner: 10Arnaudb) [07:31:19] (03CR) 10Arnaudb: [C:03+2] gerrit: hold GerritDiskSpaceExhaustionIncoming for 30m [alerts] - 10https://gerrit.wikimedia.org/r/1344595 (https://phabricator.wikimedia.org/T439039) (owner: 10Arnaudb) [07:33:45] (03Merged) 10jenkins-bot: gerrit: hold GerritDiskSpaceExhaustionIncoming for 30m [alerts] - 10https://gerrit.wikimedia.org/r/1344595 (https://phabricator.wikimedia.org/T439039) (owner: 10Arnaudb) [07:35:19] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2226 PSU loss of redundancy - https://phabricator.wikimedia.org/T439205#12362706 (10Marostegui) p:05Triage→03Medium [07:38:08] (03PS1) 10Marostegui: installserver: Do not reimage db1289 [puppet] - 10https://gerrit.wikimedia.org/r/1345038 [07:41:14] (03PS1) 10Muehlenhoff: Remove cumin1003 from SSH rules for Cumin masters [puppet] - 10https://gerrit.wikimedia.org/r/1345039 (https://phabricator.wikimedia.org/T427897) [07:41:41] (03CR) 10Ayounsi: [C:03+1] Remove cumin1003 from RAPI access [puppet] - 10https://gerrit.wikimedia.org/r/1344980 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [07:45:38] (03CR) 10Marostegui: [C:03+2] installserver: Do not reimage db1289 [puppet] - 10https://gerrit.wikimedia.org/r/1345038 (owner: 10Marostegui) [07:46:30] (03CR) 10Elukey: [C:03+2] "Thanks both!" [puppet] - 10https://gerrit.wikimedia.org/r/1344678 (owner: 10Elukey) [07:47:10] (03CR) 10Ayounsi: [C:03+1] Remove cumin1003 from SSH rules for Cumin masters [puppet] - 10https://gerrit.wikimedia.org/r/1345039 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [07:49:29] FIRING: [4x] CertManagerCertNotReady: Certificate citoid/citoid-staging-tls-proxy-certs is not in a ready state (k8s-staging@codfw) - https://wikitech.wikimedia.org/wiki/Kubernetes/cert-manager - https://alerts.wikimedia.org/?q=alertname%3DCertManagerCertNotReady [07:50:46] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin1003 from SSH rules for Cumin masters [puppet] - 10https://gerrit.wikimedia.org/r/1345039 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [07:57:59] (03PS1) 10Ryan Kemper: wdqs: restore internal scholarly transfer peers [puppet] - 10https://gerrit.wikimedia.org/r/1345041 (https://phabricator.wikimedia.org/T420707) [07:58:21] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1345041 (https://phabricator.wikimedia.org/T420707) (owner: 10Ryan Kemper) [08:00:15] FIRING: MediaWikiMemcachedHighErrorRate: MediaWiki memcached error rate is elevated globally - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?var-datasource=codfw%20prometheus/ops&viewPanel=19 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiMemcachedHighErrorRate [08:00:44] (03PS2) 10Ryan Kemper: wdqs: restore internal scholarly transfer peers [puppet] - 10https://gerrit.wikimedia.org/r/1345041 (https://phabricator.wikimedia.org/T420707) [08:01:45] FIRING: [2x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [08:02:29] (03PS3) 10Ryan Kemper: wdqs: restore internal scholarly transfer peers [puppet] - 10https://gerrit.wikimedia.org/r/1345041 (https://phabricator.wikimedia.org/T420707) [08:03:01] (03CR) 10Ryan Kemper: [V:03+2 C:03+2] "pcc looks great, merging" [puppet] - 10https://gerrit.wikimedia.org/r/1345041 (https://phabricator.wikimedia.org/T420707) (owner: 10Ryan Kemper) [08:05:15] RESOLVED: MediaWikiMemcachedHighErrorRate: MediaWiki memcached error rate is elevated globally - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?var-datasource=codfw%20prometheus/ops&viewPanel=19 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiMemcachedHighErrorRate [08:06:00] (03PS1) 10Elukey: profile::k8s::deployment_server::global_config: allow port 443 for pki [puppet] - 10https://gerrit.wikimedia.org/r/1345042 (https://phabricator.wikimedia.org/T436809) [08:06:45] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [08:07:08] (03CR) 10JMeybohm: [C:03+1] profile::k8s::deployment_server::global_config: allow port 443 for pki [puppet] - 10https://gerrit.wikimedia.org/r/1345042 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [08:08:39] (03CR) 10Jelto: [C:03+1] "lgtm" [puppet] - 10https://gerrit.wikimedia.org/r/1345042 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [08:10:38] (03CR) 10Elukey: [C:03+2] profile::k8s::deployment_server::global_config: allow port 443 for pki [puppet] - 10https://gerrit.wikimedia.org/r/1345042 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [08:13:41] (03PS2) 10Arnaudb: gerrit: remove a unit copy left by systemctl edit --full [puppet] - 10https://gerrit.wikimedia.org/r/1345040 (https://phabricator.wikimedia.org/T439209) [08:14:26] (03PS1) 10Ryan Kemper: wdqs: configure eqiad expansion hosts [puppet] - 10https://gerrit.wikimedia.org/r/1345043 (https://phabricator.wikimedia.org/T420707) [08:14:56] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1345043 (https://phabricator.wikimedia.org/T420707) (owner: 10Ryan Kemper) [08:15:44] !log vgutierrez@puppetserver1001 conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.* [08:17:36] (03CR) 10Arnaudb: [C:03+1] "lgtm for these services, thanks @hnowlan@wikimedia.org" [puppet] - 10https://gerrit.wikimedia.org/r/1344226 (https://phabricator.wikimedia.org/T357099) (owner: 10Hnowlan) [08:20:13] !log vgutierrez@puppetserver1001 conftool action : set/weight=1; selector: dc=codfw,name=cp2059.* [08:21:15] FIRING: MediaWikiMemcachedHighErrorRate: MediaWiki memcached error rate is elevated globally - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?var-datasource=eqiad%20prometheus/ops&viewPanel=19 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiMemcachedHighErrorRate [08:21:45] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [08:22:27] (03PS1) 10Ayounsi: Revert "Disable HE paths to eqsin due to significant packet loss" [homer/public] - 10https://gerrit.wikimedia.org/r/1345044 (https://phabricator.wikimedia.org/T438835) [08:22:56] !log elukey@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'sync'. [08:23:17] !log elukey@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. [08:23:46] (03CR) 10Brouberol: [C:03+1] wdqs: configure eqiad expansion hosts [puppet] - 10https://gerrit.wikimedia.org/r/1345043 (https://phabricator.wikimedia.org/T420707) (owner: 10Ryan Kemper) [08:23:53] !log elukey@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'sync'. [08:24:35] !log elukey@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. [08:26:15] RESOLVED: [2x] MediaWikiMemcachedHighErrorRate: MediaWiki memcached error rate is elevated globally - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiMemcachedHighErrorRate [08:26:45] RESOLVED: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [08:29:07] (03CR) 10Elukey: [C:03+1] Revert "Disable HE paths to eqsin due to significant packet loss" [homer/public] - 10https://gerrit.wikimedia.org/r/1345044 (https://phabricator.wikimedia.org/T438835) (owner: 10Ayounsi) [08:29:20] (03CR) 10Ayounsi: [C:03+2] Revert "Disable HE paths to eqsin due to significant packet loss" [homer/public] - 10https://gerrit.wikimedia.org/r/1345044 (https://phabricator.wikimedia.org/T438835) (owner: 10Ayounsi) [08:30:31] (03Merged) 10jenkins-bot: Revert "Disable HE paths to eqsin due to significant packet loss" [homer/public] - 10https://gerrit.wikimedia.org/r/1345044 (https://phabricator.wikimedia.org/T438835) (owner: 10Ayounsi) [08:34:31] (03CR) 10CWilliams: switchdc.databases.prepare: Retry checks for replication threads (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1344031 (https://phabricator.wikimedia.org/T438833) (owner: 10CWilliams) [08:35:48] FIRING: PuppetFailure: Puppet has failed on ml-serve1015:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [08:39:29] FIRING: [4x] CertManagerCertNotReady: Certificate citoid/citoid-staging-tls-proxy-certs is not in a ready state (k8s-staging@codfw) - https://wikitech.wikimedia.org/wiki/Kubernetes/cert-manager - https://alerts.wikimedia.org/?q=alertname%3DCertManagerCertNotReady [08:42:39] (03CR) 10Marostegui: [C:03+1] switchdc.databases.prepare: Retry checks for replication threads (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1344031 (https://phabricator.wikimedia.org/T438833) (owner: 10CWilliams) [08:44:29] FIRING: [4x] CertManagerCertNotReady: Certificate citoid/citoid-staging-tls-proxy-certs is not in a ready state (k8s-staging@codfw) - https://wikitech.wikimedia.org/wiki/Kubernetes/cert-manager - https://alerts.wikimedia.org/?q=alertname%3DCertManagerCertNotReady [08:49:29] RESOLVED: [4x] CertManagerCertNotReady: Certificate citoid/citoid-staging-tls-proxy-certs is not in a ready state (k8s-staging@codfw) - https://wikitech.wikimedia.org/wiki/Kubernetes/cert-manager - https://alerts.wikimedia.org/?q=alertname%3DCertManagerCertNotReady [08:50:19] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1344205 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [08:52:20] !log jmm@cumin2003 START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet [08:57:21] !log jmm@cumin2003 START - Cookbook sre.dns.netbox [08:57:33] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet [09:00:03] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin1003 from RAPI access [puppet] - 10https://gerrit.wikimedia.org/r/1344980 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [09:00:23] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b12-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [09:01:41] !log jmm@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [09:03:02] (03CR) 10Elukey: [C:03+1] rack depool hiera, add policy: manual and remaining server_pool keys (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1344205 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [09:04:45] jmm@cumin2003 decommission (PID 484173) is awaiting input [09:11:43] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet [09:14:41] FIRING: ConfdResourceFailed: confd resource _etc_haproxy_conf.d_tls.cfg.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [09:17:27] (03CR) 10CWilliams: switchdc.databases.prepare: Retry checks for replication threads (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1344031 (https://phabricator.wikimedia.org/T438833) (owner: 10CWilliams) [09:28:41] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [09:28:41] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:28:42] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet [09:28:48] 06SRE, 06Infrastructure-Foundations: Migrate remaining container build/report steps from build2001 to build2004 - https://phabricator.wikimedia.org/T417389#12363093 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by jmm@cumin2003 for hosts: `build2001.codfw.wmnet` - build2001.codfw.wmnet (... [09:35:03] !log cwilliams@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4 [09:35:23] RESOLVED: GnmiInterfaceCountersDrop: ... [09:35:23] asw1-b13-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=asw1-b13-drmrs:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [09:36:01] !log cwilliams@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4 [09:36:16] (03CR) 10JMeybohm: [C:03+1] services_proxy: drop eventgate timeout to 11s [puppet] - 10https://gerrit.wikimedia.org/r/1344739 (https://phabricator.wikimedia.org/T438896) (owner: 10Hnowlan) [09:36:30] !log cwilliams@cumin1004 START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4 [09:36:47] !log cwilliams@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4 [09:37:23] (03PS1) 10Ameisenigel: Enable FlaggedRevs for the Author NS on ruwikisource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345052 (https://phabricator.wikimedia.org/T439159) [09:40:25] (03CR) 10Jelto: [C:03+1] "lgtm, one nit in-line" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341335 (https://phabricator.wikimedia.org/T437635) (owner: 10Dzahn) [09:43:33] FIRING: KubernetesCalicoDown: wikikube-worker2280.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=codfw%20prometheus%2Fk8s&var-instance=wikikube-worker2280.codfw.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [09:43:35] !log reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - T439105 [09:43:38] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:43:38] T439105: Cleanup manually modified attributes on Icinga hosts and services - https://phabricator.wikimedia.org/T439105 [09:44:06] (03PS1) 10Muehlenhoff: Blocklist KCM [puppet] - 10https://gerrit.wikimedia.org/r/1345053 [09:44:36] (03CR) 10Blake: "Hi Amir! Yeah, I'm very curious about what's going on here as well. Unfortunately all we have at the moment is correlation - we cordoned w" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344653 (https://phabricator.wikimedia.org/T420223) (owner: 10Blake) [09:45:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:48:05] !log cwilliams@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4 [09:49:07] !log cwilliams@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4 [09:49:27] !log cwilliams@cumin1004 START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4 [09:49:45] !log cwilliams@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4 [09:52:19] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345053 (owner: 10Muehlenhoff) [09:53:17] (03PS1) 10Dpogorzelski: ferretdb: make the Service name configurable [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) [09:53:20] (03PS1) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [09:53:22] (03PS1) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [09:54:02] (03PS1) 10Brouberol: site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) [09:56:07] (03PS2) 10Dpogorzelski: ferretdb: make the Service name configurable [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) [09:56:08] (03PS2) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [09:56:08] (03PS2) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [09:56:30] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12363195 (10Jclark-ctr) p:05High→03Medium [09:58:45] FIRING: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [10:00:17] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12363228 (10MoritzMuehlenhoff) [10:01:06] (03CR) 10Btullis: site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) (owner: 10Brouberol) [10:01:29] (03CR) 10Brouberol: site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) (owner: 10Brouberol) [10:02:52] (03PS2) 10Brouberol: site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) [10:03:01] (03CR) 10Brouberol: site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) (owner: 10Brouberol) [10:03:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [10:04:17] (03PS3) 10Brouberol: site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) [10:04:56] (03PS4) 10Brouberol: site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) [10:05:19] (03PS3) 10Dpogorzelski: ferretdb: make the Service name configurable [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) [10:05:19] (03PS3) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [10:05:19] (03PS3) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [10:06:27] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:07:37] (03PS4) 10Dpogorzelski: ferretdb: make the Service name configurable [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) [10:07:37] (03PS4) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [10:07:37] (03PS4) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [10:11:27] (03PS5) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [10:11:27] (03PS5) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [10:12:54] (03CR) 10Brouberol: ferretdb: make the Service name configurable (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [10:13:58] (03PS6) 10CWilliams: switchdc.databases.prepare: Retry checks for replication threads [cookbooks] - 10https://gerrit.wikimedia.org/r/1344031 (https://phabricator.wikimedia.org/T438833) [10:15:41] (03CR) 10Jelto: [C:03+1] mesh: document the idle timeout key the templates actually read [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338960 (https://phabricator.wikimedia.org/T436657) (owner: 10Arnaudb) [10:15:44] (03PS5) 10Dpogorzelski: ferretdb: make the Service name configurable [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) [10:15:44] (03PS6) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [10:15:44] (03PS6) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [10:15:55] RECOVERY - Host wikikube-worker2280 is UP: PING OK - Packet loss = 0%, RTA = 30.33 ms [10:17:32] (03CR) 10Tacsipacsi: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345052 (https://phabricator.wikimedia.org/T439159) (owner: 10Ameisenigel) [10:17:55] (03PS6) 10Dpogorzelski: ferretdb: make the Service name configurable [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) [10:17:55] (03PS7) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [10:17:56] (03PS7) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [10:18:05] (03CR) 10Brouberol: [C:03+1] "LG! Thanks!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [10:18:34] (03CR) 10Jelto: [C:03+1] deployment_server/k8s: set kubeconfig files for aphlict [puppet] - 10https://gerrit.wikimedia.org/r/1341711 (https://phabricator.wikimedia.org/T436657) (owner: 10Arnaudb) [10:18:48] 10ops-codfw, 06DC-Ops, 10Prod-Kubernetes, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: wikikube-worker2280 unreachable - https://phabricator.wikimedia.org/T423395#12363308 (10JMeybohm) 05Resolved→03Open It got stuck again: ` -> start /system1/sol1 /system1/sol1 press , , and then to... [10:19:11] (03CR) 10Dpogorzelski: ferretdb: make the Service name configurable (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [10:19:26] (03CR) 10Jelto: [C:03+1] "lgtm, be sure to merge and deploy I839ff2752070bb93e36b125490cebfc9345f0421 first, see https://wikitech.wikimedia.org/wiki/Kubernetes/Add_" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341855 (https://phabricator.wikimedia.org/T436657) (owner: 10Arnaudb) [10:20:24] 06SRE, 10observability, 06Traffic: HAProxy metrics go down on config reload - https://phabricator.wikimedia.org/T343000#12363311 (10dkertesz) [10:20:40] RESOLVED: KubernetesCalicoDown: wikikube-worker2280.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=codfw%20prometheus%2Fk8s&var-instance=wikikube-worker2280.codfw.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [10:22:11] (03CR) 10Ayounsi: [C:03+1] Blocklist KCM [puppet] - 10https://gerrit.wikimedia.org/r/1345053 (owner: 10Muehlenhoff) [10:27:57] (03CR) 10Brouberol: ceph::osd: disable the OSD service at deletion (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1344626 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [10:28:38] (03CR) 10Btullis: [C:03+1] site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) (owner: 10Brouberol) [10:30:29] (03CR) 10CWilliams: [C:03+2] switchdc.databases.prepare: Retry checks for replication threads [cookbooks] - 10https://gerrit.wikimedia.org/r/1344031 (https://phabricator.wikimedia.org/T438833) (owner: 10CWilliams) [10:33:32] (03Merged) 10jenkins-bot: switchdc.databases.prepare: Retry checks for replication threads [cookbooks] - 10https://gerrit.wikimedia.org/r/1344031 (https://phabricator.wikimedia.org/T438833) (owner: 10CWilliams) [10:37:45] (03CR) 10Dpogorzelski: [C:03+2] ferretdb: make the Service name configurable [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345056 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [10:41:48] FIRING: PuppetFailure: Puppet has failed on cloudidp2001-dev:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [10:47:17] (03PS1) 10Muehlenhoff: Blocklist phonet [puppet] - 10https://gerrit.wikimedia.org/r/1345066 [10:51:48] RESOLVED: PuppetFailure: Puppet has failed on cloudidp2001-dev:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [10:52:41] !log jelto@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues [10:52:42] (03PS1) 10Daniel Kertesz: cache::haproxy: move the stats file to /run/haproxy [puppet] - 10https://gerrit.wikimedia.org/r/1345067 (https://phabricator.wikimedia.org/T343000) [10:52:53] 10ops-eqiad, 06SRE, 06DC-Ops, 10Prod-Kubernetes, and 3 others: wikikube-worker1152.eqiad.wmnet networking issue - https://phabricator.wikimedia.org/T438819#12363460 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=e3cacf74-92ed-4fb6-bf59-087ad26e7040) set by jelto@cumin1004 for 4 days, 0... [10:53:30] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet [10:54:44] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply [10:54:54] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply [10:55:44] (03PS1) 10Matthias Mullie: [MultimediaViewer] Testwiki use same carousel config as other wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345070 [10:56:46] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 28 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345070 (owner: 10Matthias Mullie) [10:58:30] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet [11:00:05] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260925T0700) [11:00:05] jelto, arnoldokoth, mutante, and arnaudb: That opportune time for a GitLab version upgrades deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260925T1100). [11:01:44] (03PS2) 10Daniel Kertesz: cache::haproxy: move the stats file to /run/haproxy [puppet] - 10https://gerrit.wikimedia.org/r/1345067 (https://phabricator.wikimedia.org/T343000) [11:03:54] (03PS1) 10Kevin Bazira: ml-services: update tts-section-generator service to strip script symbols by block not by name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345073 (https://phabricator.wikimedia.org/T438647) [11:04:01] 10ops-eqiad, 06SRE, 06DC-Ops, 10Prod-Kubernetes, and 3 others: wikikube-worker1152.eqiad.wmnet networking issue - https://phabricator.wikimedia.org/T438819#12363472 (10Jelto) a:05Jelto→03None >>! In T438819#12350122, @Jclark-ctr wrote: > @Jelto It looks like the NICs on the motherboard have failed. We... [11:09:54] (03PS1) 10Atsuko: airflow-test-k8s: upgrade to airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345077 (https://phabricator.wikimedia.org/T433383) [11:12:10] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update tts-section-generator service to strip script symbols by block not by name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345073 (https://phabricator.wikimedia.org/T438647) (owner: 10Kevin Bazira) [11:14:41] (03Merged) 10jenkins-bot: ml-services: update tts-section-generator service to strip script symbols by block not by name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345073 (https://phabricator.wikimedia.org/T438647) (owner: 10Kevin Bazira) [11:14:50] (03CR) 10Brouberol: [C:03+2] site: ganeti-jumbo eqiad -> dse-k8s-workers repurposing [puppet] - 10https://gerrit.wikimedia.org/r/1345059 (https://phabricator.wikimedia.org/T439220) (owner: 10Brouberol) [11:17:32] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [11:19:14] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . [11:20:24] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12363549 (10MoritzMuehlenhoff) [11:20:51] !log kevinbazira@deploy1003 helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [11:25:19] (03CR) 10Btullis: [C:03+1] airflow-test-k8s: upgrade to airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345077 (https://phabricator.wikimedia.org/T433383) (owner: 10Atsuko) [11:36:34] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet [11:40:47] !log urbanecm@deploy1003 mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # T438366 [11:40:50] T438366: Run GrowthExperiments:revalidateLinkRecommendations.php on all Wikimedia wikis - https://phabricator.wikimedia.org/T438366 [11:41:22] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet [11:55:03] (03PS1) 10Hnowlan: mediawiki-cache-warmup: use ThreadPoolExecutor for workers [puppet] - 10https://gerrit.wikimedia.org/r/1345079 [12:00:04] (03PS1) 10Brouberol: network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) [12:01:45] (03PS2) 10Brouberol: network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) [12:02:22] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Q1:rack/setup/install 10 new ceph nodes - https://phabricator.wikimedia.org/T438213#12363670 (10BTullis) [12:02:36] (03PS3) 10Brouberol: network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) [12:05:15] (03CR) 10CI reject: [V:04-1] network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) (owner: 10Brouberol) [12:07:51] (03PS4) 10Brouberol: network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) [12:08:31] 07sre-alert-triage, 06Infrastructure-Foundations, 07Essential-Work: Alert in need of triage: PKICertificateExpiry (instance pki2002:9100) - https://phabricator.wikimedia.org/T438703#12363698 (10BTullis) I think that this host is managed by the #infrastructure-foundations team. [12:09:20] (03CR) 10Brouberol: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9479/co" [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) (owner: 10Brouberol) [12:10:11] (03PS5) 10Brouberol: network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) [12:11:42] (03CR) 10Ayounsi: [C:03+1] network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) (owner: 10Brouberol) [12:13:12] (03PS1) 10Majavah: P:puppetdb: Drop support for Puppet 5 CA compat site [puppet] - 10https://gerrit.wikimedia.org/r/1345086 [12:13:26] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 28 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343611 (https://phabricator.wikimedia.org/T432126) (owner: 10Michael Große) [12:13:43] !log jmm@cumin2003 START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet [12:13:48] (03PS6) 10Brouberol: network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) [12:14:32] (03CR) 10Ayounsi: [C:04-1] "I think we can abandon this :)" [puppet] - 10https://gerrit.wikimedia.org/r/1205135 (https://phabricator.wikimedia.org/T410047) (owner: 10Cathal Mooney) [12:14:49] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1 DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1345086 (owner: 10Majavah) [12:16:11] (03PS1) 10Muehlenhoff: Remove cumin1003 from the alertmanager and tcpircbot configs [puppet] - 10https://gerrit.wikimedia.org/r/1345087 (https://phabricator.wikimedia.org/T427897) [12:16:36] (03CR) 10Btullis: [C:03+1] "Looks good to me." [puppet] - 10https://gerrit.wikimedia.org/r/1344205 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [12:16:36] (03CR) 10Brouberol: [C:03+2] network/data: rename the new dse-kubepods ipv4 subnet [puppet] - 10https://gerrit.wikimedia.org/r/1345081 (https://phabricator.wikimedia.org/T439241) (owner: 10Brouberol) [12:17:42] (03PS1) 10Muehlenhoff: Remove cumin1003 from config for private Homer [puppet] - 10https://gerrit.wikimedia.org/r/1345088 (https://phabricator.wikimedia.org/T427897) [12:18:22] (03PS1) 10Muehlenhoff: Remove cumin1003 from site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1345089 (https://phabricator.wikimedia.org/T427897) [12:18:50] !log jmm@cumin2003 START - Cookbook sre.dns.netbox [12:20:18] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply [12:20:28] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply [12:21:14] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply [12:21:24] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply [12:24:17] jmm@cumin2003 decommission (PID 566314) is awaiting input [12:25:24] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin1003 from site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1345089 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [12:26:00] !log jmm@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [12:29:04] jmm@cumin2003 decommission (PID 566314) is awaiting input [12:32:59] (03CR) 10Vladis13: [C:03+1] Enable FlaggedRevs for the Author NS on ruwikisource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345052 (https://phabricator.wikimedia.org/T439159) (owner: 10Ameisenigel) [12:33:48] !log brouberol@cumin1004 START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039 [12:34:09] !log brouberol@cumin1004 START - Cookbook sre.dns.netbox [12:34:11] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [12:34:11] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:34:12] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet [12:34:22] (03PS1) 10Muehlenhoff: Remove obsolete Hiera config [puppet] - 10https://gerrit.wikimedia.org/r/1345092 [12:34:22] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Upgrade Cumin hosts to Trixie - https://phabricator.wikimedia.org/T427897#12363786 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by jmm@cumin2003 for hosts: `cumin1003.eqiad.wmnet` - cumin1003.eqiad.wmnet (**PASS**) - Downtimed hos... [12:34:23] (03PS3) 10Daniel Kertesz: cache::haproxy: move the stats file to /run/haproxy [puppet] - 10https://gerrit.wikimedia.org/r/1345067 (https://phabricator.wikimedia.org/T343000) [12:35:46] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1344811 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [12:35:48] FIRING: PuppetFailure: Puppet has failed on ml-serve1015:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [12:36:47] (03CR) 10Majavah: [V:03+1 C:03+2] P:wmcs::etcd: Add profile to manage cloudinfra client firewall [puppet] - 10https://gerrit.wikimedia.org/r/1344303 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [12:37:00] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Tsevener - https://phabricator.wikimedia.org/T439138#12363812 (10ayounsi) [12:37:12] (03CR) 10Majavah: [C:03+2] O:wmcs::novaproxy: Add confd profile [puppet] - 10https://gerrit.wikimedia.org/r/1344318 (https://phabricator.wikimedia.org/T438971) (owner: 10Majavah) [12:37:17] (03CR) 10Daniel Kertesz: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345067 (https://phabricator.wikimedia.org/T343000) (owner: 10Daniel Kertesz) [12:37:20] 10ops-eqiad, 06SRE, 06DC-Ops, 10Prod-Kubernetes, and 3 others: wikikube-worker1152.eqiad.wmnet networking issue - https://phabricator.wikimedia.org/T438819#12363816 (10Jclark-ctr) a:03Jelto I have installed a new NIC, and it has link. It is also showing up in iDRAC. However, I do not have root access, an... [12:38:25] !log brouberol@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004" [12:39:43] !log brouberol@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004" [12:39:43] !log brouberol@cumin1004 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:39:43] !log brouberol@cumin1004 START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors [12:39:46] !log brouberol@cumin1004 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors [12:39:47] !log brouberol@cumin1004 START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039 [12:40:07] !log brouberol@cumin1004 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039 [12:40:15] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039 [12:42:00] !log brouberol@cumin1004 START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm [12:43:19] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - thanos-query_443: Servers titan1002.eqiad.wmnet are marked down but pooled: thanos-web_443: Servers titan1001.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [12:43:36] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:43:57] FIRING: ProbeDown: Service thanos-query:443 has failed probes (http_thanos-query_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#thanos-query:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:44:11] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:44:19] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - thanos-query_443: Servers titan1002.eqiad.wmnet are marked down but pooled: thanos-web_443: Servers titan1002.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [12:44:24] !ack [12:44:25] 8600 (ACKED) ProbeDown sre (10.2.2.53 ip4 thanos-query:443 probes/service http_thanos-query_ip4 eqiad) [12:44:29] * volans looking [12:44:44] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [12:45:19] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [12:45:41] FIRING: [2x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:45:42] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [12:45:54] (03CR) 10Elukey: [C:03+1] Remove cumin1003 from config for private Homer [puppet] - 10https://gerrit.wikimedia.org/r/1345088 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [12:46:19] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [12:47:13] !log mvernon@cumin1004 START - Cookbook sre.discovery.service-route check sessionstore: maintenance [12:47:13] !log mvernon@cumin1004 END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance [12:48:14] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Tsevener - https://phabricator.wikimedia.org/T439138#12363874 (10ayounsi) @thcipriani @dancy do you approve this request ? @Tsevener can you sign https://phabricator.wikimedia.org/L3 ? Thanks [12:48:14] (03CR) 10Elukey: "Should we try to follow up on this during these weeks?" [software/spicerack] - 10https://gerrit.wikimedia.org/r/1120500 (owner: 10Volans) [12:48:27] RESOLVED: [3x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:48:57] RESOLVED: ProbeDown: Service thanos-query:443 has failed probes (http_thanos-query_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#thanos-query:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:49:24] (03CR) 10Elukey: [C:03+1] "Good to go for a test!" [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [12:51:27] 06SRE, 10SRE-Access-Requests: Requesting access to Deployment for GreyOlson/GOlson-WMF - https://phabricator.wikimedia.org/T439143#12363905 (10ayounsi) @GOlson-WMF could you use an ed25519 ssh key? Or detail why you can't use one. @thcipriani @dancy do you approve this request ? Thanks [12:51:57] 06SRE, 10SRE-Access-Requests: Requesting access to Deployment for GreyOlson/GOlson-WMF - https://phabricator.wikimedia.org/T439143#12363906 (10ayounsi) [12:52:25] !log brouberol@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage [12:53:04] (03CR) 10Brouberol: airflow-test-k8s: upgrade to airflow 3 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345077 (https://phabricator.wikimedia.org/T433383) (owner: 10Atsuko) [12:55:08] 10ops-eqiad, 06SRE, 06DC-Ops: Possible Netbox incorrection on device type - https://phabricator.wikimedia.org/T438963#12363923 (10Jclark-ctr) 05Open→03Resolved a:03Jclark-ctr Fixed errors in netbox [12:55:50] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage [12:56:14] (03PS1) 10Brouberol: site: add dse-k8s-worker1039 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345097 (https://phabricator.wikimedia.org/T439220) [12:57:10] 06SRE, 10SRE-Access-Requests: Requesting access to Deployment for GreyOlson/GOlson-WMF - https://phabricator.wikimedia.org/T439143#12363934 (10GOlson-WMF) @ayounsi I just had one (ed25519) already for other access, figured it would be a little easier for me to have the two separate style ones! [12:57:40] 07sre-alert-triage, 06Infrastructure-Foundations, 10Kafka-Infrastructure, 07Essential-Work: Alert in need of triage: PKICertificateExpiry (instance pki2002:9100) - https://phabricator.wikimedia.org/T438703#12363936 (10elukey) Definitely yes this is I/F related :) I am also going to add the Kafka Infrastruc... [12:57:44] (03CR) 10Ayounsi: [C:03+1] Remove cumin1003 from config for private Homer [puppet] - 10https://gerrit.wikimedia.org/r/1345088 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [12:58:20] !log jclark@cumin1004 START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [12:59:39] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet [13:00:28] 06SRE, 10SRE-Access-Requests: Requesting access to Deployment for GreyOlson/GOlson-WMF - https://phabricator.wikimedia.org/T439143#12363948 (10Seddon) Request approved! [13:02:44] (03PS1) 10Gmodena: wdqs: widen proxy readiness probe timeout/failureThreshold [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345098 (https://phabricator.wikimedia.org/T439248) [13:02:54] (03CR) 10CI reject: [V:04-1] wdqs: widen proxy readiness probe timeout/failureThreshold [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345098 (https://phabricator.wikimedia.org/T439248) (owner: 10Gmodena) [13:03:00] (03PS1) 10DCausse: opensearch-semantic-search: bump to latest image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345099 [13:03:19] !log jclark@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [13:05:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:06:27] (03PS1) 10Daniel Kertesz: hiera: enable HAProxy's persistent stats on two esams nodes [puppet] - 10https://gerrit.wikimedia.org/r/1345100 (https://phabricator.wikimedia.org/T343000) [13:06:29] (03PS1) 10Daniel Kertesz: hiera: remove individual entries for haproxy's persistent stats [puppet] - 10https://gerrit.wikimedia.org/r/1345101 (https://phabricator.wikimedia.org/T343000) [13:06:31] (03PS1) 10Daniel Kertesz: hiera: enable haproxy's persistent stats in magru [puppet] - 10https://gerrit.wikimedia.org/r/1345102 (https://phabricator.wikimedia.org/T343000) [13:06:33] (03PS1) 10Daniel Kertesz: hiera: enable haproxy's persistent stats in drmrs [puppet] - 10https://gerrit.wikimedia.org/r/1345103 (https://phabricator.wikimedia.org/T343000) [13:06:35] (03PS1) 10Daniel Kertesz: hiera: enable haproxy's persistent stats in eqsin [puppet] - 10https://gerrit.wikimedia.org/r/1345104 (https://phabricator.wikimedia.org/T343000) [13:06:37] (03PS1) 10Daniel Kertesz: hiera: enable haproxy's persistent stats in ulsfo [puppet] - 10https://gerrit.wikimedia.org/r/1345105 (https://phabricator.wikimedia.org/T343000) [13:06:40] (03PS1) 10Daniel Kertesz: hiera: enable haproxy's persistent stats in esams [puppet] - 10https://gerrit.wikimedia.org/r/1345106 (https://phabricator.wikimedia.org/T343000) [13:06:42] (03PS1) 10Daniel Kertesz: hiera: enable haproxy's persistent stats in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1345107 (https://phabricator.wikimedia.org/T343000) [13:06:44] (03PS1) 10Daniel Kertesz: hiera: enable haproxy's persistent stats in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1345108 (https://phabricator.wikimedia.org/T343000) [13:08:55] 06SRE, 10SRE-Access-Requests: Requesting access to Deployment for GreyOlson/GOlson-WMF - https://phabricator.wikimedia.org/T439143#12364029 (10ayounsi) >>! In T439143#12363934, @GOlson-WMF wrote: > @ayounsi I just had one (ed25519) already for other access, figured it would be a little easier for me to have th... [13:09:11] 06SRE, 10SRE-Access-Requests: Requesting access to Deployment for GreyOlson/GOlson-WMF - https://phabricator.wikimedia.org/T439143#12364030 (10ayounsi) [13:11:28] 06SRE, 10SRE-Access-Requests: Requesting access to Deployment for GreyOlson/GOlson-WMF - https://phabricator.wikimedia.org/T439143#12364048 (10GOlson-WMF) Oops sorry about that - here's that key ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIHS1w+9rRFLmEcY5GkLzcgiCXpT93o99KT+PQENDw+Lx deployment [13:14:41] FIRING: ConfdResourceFailed: confd resource _etc_haproxy_conf.d_tls.cfg.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [13:15:26] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Tsevener - https://phabricator.wikimedia.org/T439138#12364087 (10ayounsi) [13:15:54] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm [13:17:28] (03PS2) 10Gmodena: wdqs: widen proxy readiness probe timeout/failureThreshold [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345098 (https://phabricator.wikimedia.org/T439248) [13:19:01] (03CR) 10Atsuko: airflow-test-k8s: upgrade to airflow 3 (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345077 (https://phabricator.wikimedia.org/T433383) (owner: 10Atsuko) [13:19:32] (03CR) 10Zabe: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344216 (https://phabricator.wikimedia.org/T386776) (owner: 10Ameisenigel) [13:20:03] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply [13:20:13] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply [13:20:37] (03CR) 10CI reject: [V:04-1] Enable Extension:Translate and Extension:TranslationNotifications on co.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344216 (https://phabricator.wikimedia.org/T386776) (owner: 10Ameisenigel) [13:20:39] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for Tsevener - https://phabricator.wikimedia.org/T439138#12364092 (10ayounsi) [13:20:43] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 28 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [extensions/WikimediaEvents] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1344594 (https://phabricator.wikimedia.org/T437339) (owner: 10Michael Große) [13:21:58] 06SRE, 10SRE-Access-Requests: Requesting access to Deployment for GreyOlson/GOlson-WMF - https://phabricator.wikimedia.org/T439143#12364095 (10ayounsi) [13:22:08] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply [13:22:18] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply [13:23:43] (03CR) 10Bking: [C:03+2] opensearch-semantic-search: bump to latest image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345099 (owner: 10DCausse) [13:23:44] (03CR) 10Ayounsi: [C:03+1] Blocklist phonet [puppet] - 10https://gerrit.wikimedia.org/r/1345066 (owner: 10Muehlenhoff) [13:24:26] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [13:24:29] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [13:24:29] !log mvernon@cumin1004 START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active [13:24:46] !log repool sessionstore in codfw [13:24:47] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:26:48] (03CR) 10Zabe: "Done." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344216 (https://phabricator.wikimedia.org/T386776) (owner: 10Ameisenigel) [13:26:58] PROBLEM - SSH on cp2059 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [13:26:58] PROBLEM - Ensure traffic_exporter for the backend instance binds on port 9122 and responds to HTTP requests on cp2059 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [13:28:48] RECOVERY - SSH on cp2059 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [13:28:48] RECOVERY - Ensure traffic_exporter for the backend instance binds on port 9122 and responds to HTTP requests on cp2059 is OK: HTTP OK: HTTP/1.0 200 OK - 36008 bytes in 0.111 second response time https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [13:28:58] (03CR) 10Btullis: [C:03+1] site: add dse-k8s-worker1039 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345097 (https://phabricator.wikimedia.org/T439220) (owner: 10Brouberol) [13:28:58] PROBLEM - haproxy process on cp2059 is CRITICAL: PROCS CRITICAL: 0 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [13:29:12] PROBLEM - HAProxy HTTPS wikiworkshop.org ECDSA on cp2059 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [13:29:12] PROBLEM - HAProxy HTTPS wikipedia25.org ECDSA on cp2059 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [13:29:12] PROBLEM - HAProxy HTTPS wikipedia.org ECDSA on cp2059 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [13:29:24] (03CR) 10Brouberol: [C:03+2] site: add dse-k8s-worker1039 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345097 (https://phabricator.wikimedia.org/T439220) (owner: 10Brouberol) [13:29:33] !log mvernon@cumin1004 END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active [13:29:41] RESOLVED: ConfdResourceFailed: confd resource _etc_haproxy_conf.d_tls.cfg.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [13:34:36] !log brouberol@cumin1004 START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040 [13:34:41] FIRING: ConfdResourceFailed: confd resource _etc_haproxy_conf.d_tls.cfg.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [13:34:47] !log brouberol@cumin1004 START - Cookbook sre.dns.netbox [13:34:58] RECOVERY - haproxy process on cp2059 is OK: PROCS OK: 2 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [13:35:12] RECOVERY - HAProxy HTTPS wikiworkshop.org ECDSA on cp2059 is OK: SSL OK - Certificate wikiworkshop.org contains all required SANs:Certificate wikiworkshop.org (ECDSA) valid until 2026-11-09 03:18:39 +0000 (expires in 44 days) https://wikitech.wikimedia.org/wiki/HTTPS [13:35:12] RECOVERY - HAProxy HTTPS wikipedia25.org ECDSA on cp2059 is OK: SSL OK - Certificate wikipedia25.org contains all required SANs:Certificate wikipedia25.org (ECDSA) valid until 2026-12-03 05:05:02 +0000 (expires in 68 days) https://wikitech.wikimedia.org/wiki/HTTPS [13:35:12] RECOVERY - HAProxy HTTPS wikipedia.org ECDSA on cp2059 is OK: SSL OK - Certificate *.wikipedia.org contains all required SANs:Certificate *.wikipedia.org (ECDSA) valid until 2026-11-03 19:15:40 +0000 (expires in 39 days) https://wikitech.wikimedia.org/wiki/HTTPS [13:38:39] !log brouberol@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004" [13:39:06] !log brouberol@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004" [13:39:06] !log brouberol@cumin1004 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:39:06] !log brouberol@cumin1004 START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors [13:39:09] !log brouberol@cumin1004 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors [13:39:10] !log brouberol@cumin1004 START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040 [13:39:29] !log brouberol@cumin1004 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040 [13:39:40] (03PS1) 10Btullis: ceph: Enable STS on the codfw rados gateway [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) [13:39:41] RESOLVED: ConfdResourceFailed: confd resource _etc_haproxy_conf.d_tls.cfg.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [13:40:03] (03PS1) 10Btullis: ceph: Add a dummy STS key for the rados gateway [labs/private] - 10https://gerrit.wikimedia.org/r/1345118 (https://phabricator.wikimedia.org/T435590) [13:40:07] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040 [13:40:40] (03CR) 10Vgutierrez: [C:03+1] "thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1345067 (https://phabricator.wikimedia.org/T343000) (owner: 10Daniel Kertesz) [13:40:44] (03CR) 10Btullis: [V:03+2 C:03+2] ceph: Add a dummy STS key for the rados gateway [labs/private] - 10https://gerrit.wikimedia.org/r/1345118 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [13:42:09] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply [13:42:19] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply [13:42:26] (03PS1) 10Hnowlan: mediawiki-cache-warmup: log to file [puppet] - 10https://gerrit.wikimedia.org/r/1345119 [13:43:58] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344325 (https://phabricator.wikimedia.org/T421911) (owner: 10Andrew Bogott) [13:44:17] !log vgutierrez@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P{lvs1020.*} and A:lvs [13:44:51] !log vgutierrez@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P{lvs1020.*} and A:lvs [13:45:29] !log brouberol@cumin1004 START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm [13:45:37] !log brouberol@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet [13:46:35] !log vgutierrez@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P{lvs1019.*} and A:lvs [13:46:58] !log vgutierrez@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P{lvs1019.*} and A:lvs [13:47:31] !log brouberol@cumin1004 START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041 [13:47:54] !log brouberol@cumin1004 START - Cookbook sre.dns.netbox [13:47:59] Pooling back eventgate-analytics eventgate-analytics-external eventgate-logging-external eventgate-main in codfw [13:48:59] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply [13:49:10] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply [13:50:51] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet [13:51:23] !log clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # T438750 [13:51:24] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:51:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:51:26] T438750: Clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis - https://phabricator.wikimedia.org/T438750 [13:51:53] !log brouberol@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004" [13:52:31] (03CR) 10Muehlenhoff: [C:03+2] Remove obsolete Hiera config [puppet] - 10https://gerrit.wikimedia.org/r/1345092 (owner: 10Muehlenhoff) [13:52:38] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [13:52:42] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [13:52:48] !log brouberol@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004" [13:52:48] !log brouberol@cumin1004 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:52:48] !log brouberol@cumin1004 START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors [13:52:52] !log brouberol@cumin1004 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors [13:52:52] !log brouberol@cumin1004 START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041 [13:53:03] !log brouberol@cumin1004 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041 [13:53:35] (03PS1) 10Tiziano Fogli: hiera/pdu_families: update regexps to match new models [puppet] - 10https://gerrit.wikimedia.org/r/1345124 (https://phabricator.wikimedia.org/T438915) [13:53:39] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041 [13:54:37] !log brouberol@cumin1004 START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm [13:55:22] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [13:55:29] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [13:56:50] (03PS1) 10Elukey: profile::amd_gpu: update firmware-amd-graphics to its latest version [puppet] - 10https://gerrit.wikimedia.org/r/1345125 [13:57:12] (03PS1) 10Brouberol: site: add dse-k8s-worker1040 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345126 (https://phabricator.wikimedia.org/T439241) [13:57:15] (03PS1) 10Brouberol: site: add dse-k8s-worker1041 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345127 (https://phabricator.wikimedia.org/T439241) [13:57:56] (03PS2) 10Btullis: ceph: Enable STS on the codfw rados gateway [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) [13:58:18] (03PS2) 10Tiziano Fogli: hiera/pdu_families: update regexps to match new models [puppet] - 10https://gerrit.wikimedia.org/r/1345124 (https://phabricator.wikimedia.org/T438915) [13:58:26] (03CR) 10Btullis: [C:03+1] site: add dse-k8s-worker1040 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345126 (https://phabricator.wikimedia.org/T439241) (owner: 10Brouberol) [13:58:42] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [13:58:45] (03CR) 10Btullis: [C:03+1] site: add dse-k8s-worker1041 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345127 (https://phabricator.wikimedia.org/T439241) (owner: 10Brouberol) [13:58:47] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [13:59:00] (03PS1) 10Mhorsey: Enable the worklist card view on the Beta Cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345129 (https://phabricator.wikimedia.org/T435507) [13:59:01] (03CR) 10Tiziano Fogli: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345124 (https://phabricator.wikimedia.org/T438915) (owner: 10Tiziano Fogli) [13:59:05] (03PS1) 10Sbisson: Remove unused config: ArticleGuidanceRedirectEntryPointTitles [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345130 (https://phabricator.wikimedia.org/T437643) [13:59:22] (03PS3) 10Andrew Bogott: Revert "Puppetize keystone-uwsgi.ini" [puppet] - 10https://gerrit.wikimedia.org/r/1344333 (https://phabricator.wikimedia.org/T421911) [13:59:22] (03PS20) 10Andrew Bogott: openstack apis: Add additional uwsgi ini file to support logstash logging [puppet] - 10https://gerrit.wikimedia.org/r/1344325 (https://phabricator.wikimedia.org/T421911) [13:59:23] (03PS1) 10Andrew Bogott: Remove unneeded duplicated placement-api init.d file [puppet] - 10https://gerrit.wikimedia.org/r/1345131 [13:59:41] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 28 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345130 (https://phabricator.wikimedia.org/T437643) (owner: 10Sbisson) [13:59:44] !log brouberol@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage [14:00:22] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [14:00:26] (03PS2) 10Mhorsey: Enable the worklist card view on the Beta Cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345129 (https://phabricator.wikimedia.org/T437353) [14:00:35] !log atsuko@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw [14:00:36] !log atsuko@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw [14:00:36] !log atsuko@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw [14:00:37] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [14:00:37] !log atsuko@cumin2003 conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw [14:01:05] !log brouberol@cumin1004 conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet [14:02:16] !log brouberol@cumin1004 conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet [14:03:05] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345131 (owner: 10Andrew Bogott) [14:05:27] !log brouberol@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage [14:06:27] FIRING: [2x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:06:42] (03CR) 10JHathaway: [C:03+1] P:puppetdb: Drop support for Puppet 5 CA compat site [puppet] - 10https://gerrit.wikimedia.org/r/1345086 (owner: 10Majavah) [14:06:46] !log brouberol@cumin1004 END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage [14:06:47] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage [14:06:58] (03PS2) 10Trueg: wdqs: widen proxy readiness probe timeout/failureThreshold [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345098 (https://phabricator.wikimedia.org/T439248) (owner: 10Gmodena) [14:07:00] (03CR) 10Andrew Bogott: [C:03+2] Remove unneeded duplicated placement-api init.d file [puppet] - 10https://gerrit.wikimedia.org/r/1345131 (owner: 10Andrew Bogott) [14:07:10] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344325 (https://phabricator.wikimedia.org/T421911) (owner: 10Andrew Bogott) [14:08:12] I know we're not supposed to be doing any releases today, but would anyone mind just merging a tiny config change to BETA only? https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1345129 [14:12:48] cdobbins@cumin1004 reimage (PID 843201) is awaiting input [14:13:12] (03CR) 10Brouberol: [C:03+1] airflow-test-k8s: upgrade to airflow 3 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345077 (https://phabricator.wikimedia.org/T433383) (owner: 10Atsuko) [14:13:21] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply [14:13:31] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply [14:13:44] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply [14:13:54] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply [14:14:20] !log cdobbins@cumin1004 START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie [14:14:46] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply [14:14:56] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply [14:15:16] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply [14:15:25] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply [14:16:46] (03CR) 10JHathaway: [C:03+2] stdlib: remove has_key usage [puppet] - 10https://gerrit.wikimedia.org/r/1344811 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [14:16:47] (03PS21) 10Andrew Bogott: openstack apis: Add additional uwsgi ini file to support logstash logging [puppet] - 10https://gerrit.wikimedia.org/r/1344325 (https://phabricator.wikimedia.org/T421911) [14:17:19] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply [14:17:30] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply [14:17:38] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply [14:17:48] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply [14:17:57] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply [14:18:08] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply [14:20:06] (03PS22) 10Andrew Bogott: openstack apis: Add additional uwsgi ini file to support logstash logging [puppet] - 10https://gerrit.wikimedia.org/r/1344325 (https://phabricator.wikimedia.org/T421911) [14:20:15] (03CR) 10Hashar: [C:03+2] Enable the worklist card view on the Beta Cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345129 (https://phabricator.wikimedia.org/T437353) (owner: 10Mhorsey) [14:20:15] (03PS4) 10Andrew Bogott: Revert "Puppetize keystone-uwsgi.ini" [puppet] - 10https://gerrit.wikimedia.org/r/1344333 (https://phabricator.wikimedia.org/T421911) [14:20:15] (03PS23) 10Andrew Bogott: openstack apis: Add additional uwsgi ini file to support logstash logging [puppet] - 10https://gerrit.wikimedia.org/r/1344325 (https://phabricator.wikimedia.org/T421911) [14:21:40] (03Merged) 10jenkins-bot: Enable the worklist card view on the Beta Cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345129 (https://phabricator.wikimedia.org/T437353) (owner: 10Mhorsey) [14:22:53] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344325 (https://phabricator.wikimedia.org/T421911) (owner: 10Andrew Bogott) [14:23:00] (03CR) 10Daniel Kertesz: [C:03+2] cache::haproxy: move the stats file to /run/haproxy [puppet] - 10https://gerrit.wikimedia.org/r/1345067 (https://phabricator.wikimedia.org/T343000) (owner: 10Daniel Kertesz) [14:23:44] !log vgutierrez@puppetserver1001 conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.* [14:23:56] (03CR) 10Muehlenhoff: "Nuking the P5 bits seems fine, but the underlying site mechanism seems worth keeping around, it might very well be useful when moving to P" [puppet] - 10https://gerrit.wikimedia.org/r/1345086 (owner: 10Majavah) [14:29:28] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm [14:30:14] !log moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - T343000 [14:30:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:30:17] T343000: HAProxy metrics go down on config reload - https://phabricator.wikimedia.org/T343000 [14:32:44] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm [14:33:44] !log cdobbins@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage [14:35:01] (03CR) 10Brouberol: [C:03+2] site: add dse-k8s-worker1040 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345126 (https://phabricator.wikimedia.org/T439241) (owner: 10Brouberol) [14:35:04] (03CR) 10Brouberol: [C:03+2] site: add dse-k8s-worker1041 to the dse-k8s-eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1345127 (https://phabricator.wikimedia.org/T439241) (owner: 10Brouberol) [14:36:50] (03CR) 10Atsuko: [C:03+1] "lgtm" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [14:38:25] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage [14:40:31] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1345125 (owner: 10Elukey) [14:40:50] (03PS3) 10Ebernhardson: semantic ssd test: Drop to 3 masters, fix affinity [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344715 (https://phabricator.wikimedia.org/T438058) [14:41:10] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [14:41:14] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [14:42:36] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [14:43:54] !log brouberol@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet [14:44:22] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [14:44:31] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [14:45:09] (03CR) 10Elukey: [C:03+2] profile::amd_gpu: update firmware-amd-graphics to its latest version [puppet] - 10https://gerrit.wikimedia.org/r/1345125 (owner: 10Elukey) [14:45:41] (03CR) 10Atsuko: liftwing-studio: back the database with ferretdb (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [14:46:30] !log brouberol@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet [14:46:38] (03CR) 10Klausman: [C:03+1] profile::amd_gpu: update firmware-amd-graphics to its latest version [puppet] - 10https://gerrit.wikimedia.org/r/1345125 (owner: 10Elukey) [14:49:08] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet [14:49:26] !log brouberol@cumin1004 conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet [14:49:31] !log brouberol@cumin1004 conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet [14:51:34] !log brouberol@cumin1004 conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet [14:51:39] !log brouberol@cumin1004 conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet [14:51:44] !log brouberol@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet [14:53:21] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [14:55:06] (03PS7) 10Blake: sidecars: enable a restartPolicy: Always option. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) [14:56:41] !log brouberol@cumin1004 conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet [14:56:46] !log brouberol@cumin1004 conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet [14:57:37] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie [15:00:03] (03CR) 10Blake: "Oh, interesting... Does this addition look somewhat correct? How do I look at the outcome of a particular fixture?" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:01:30] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply [15:01:38] !log dcausse@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply [15:01:59] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply [15:02:03] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply [15:03:05] (03PS8) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [15:03:06] (03PS8) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [15:03:33] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [15:03:42] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [15:05:09] 06SRE, 10observability, 06Traffic, 13Patch-For-Review: HAProxy metrics go down on config reload - https://phabricator.wikimedia.org/T343000#12364518 (10dkertesz) Something weird happened: we decided to move the stats file to `/run` so that it //wouldn't// survive reboots, so I prepared a [[ https://gerrit.... [15:07:32] !log cdobbins@cumin1004 conftool action : set/pooled=yes; selector: name=ncredir2001.* [15:07:55] (03PS1) 10Btullis: ceph: Move the dummy STS key to the correct hiera file [labs/private] - 10https://gerrit.wikimedia.org/r/1345151 (https://phabricator.wikimedia.org/T435590) [15:09:12] (03CR) 10Ladsgroup: "definitely go for it. I can't +1 it since I don't understand helm charts/k8s config well enough to review it properly." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344653 (https://phabricator.wikimedia.org/T420223) (owner: 10Blake) [15:13:35] (03CR) 10Btullis: [V:03+2 C:03+2] ceph: Move the dummy STS key to the correct hiera file [labs/private] - 10https://gerrit.wikimedia.org/r/1345151 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [15:13:54] (03PS9) 10Dpogorzelski: liftwing-studio: back the database with ferretdb [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) [15:13:54] (03PS9) 10Dpogorzelski: liftwing-studio: replace Open WebUI with LibreChat [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345058 (https://phabricator.wikimedia.org/T437706) [15:14:08] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [15:14:43] (03CR) 10Dpogorzelski: liftwing-studio: back the database with ferretdb (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [15:15:23] 06SRE, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T438229#12364623 (10bking) @MoritzMuehlenhoff I don't see any value to tying this service to a specific server, particularly one that seems broken. I gu... [15:18:54] (03CR) 10Atsuko: [C:03+1] liftwing-studio: back the database with ferretdb (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345057 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [15:20:48] RESOLVED: PuppetFailure: Puppet has failed on ml-serve1015:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [15:31:03] (03PS1) 10TChin: [beta eventgate] Migrate urls to deployment-eventgate05 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345154 (https://phabricator.wikimedia.org/T429497) [15:32:46] (03CR) 10TChin: "cc'ing @phuedx@wikimedia.org since it technically touches test kitchen config" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345154 (https://phabricator.wikimedia.org/T429497) (owner: 10TChin) [15:33:11] 06SRE, 10observability, 06Traffic, 13Patch-For-Review: HAProxy metrics go down on config reload - https://phabricator.wikimedia.org/T343000#12364687 (10Vgutierrez) it definitely looks like you're hitting the same issue mentioned on https://github.com/haproxy/haproxy/issues/2483#issuecomment-2579845321, def... [15:34:02] 06SRE, 06ServiceOps, 07Datacenter-Switchover: Updates to warmup script (2020-2021) - https://phabricator.wikimedia.org/T269179#12364691 (10jasmine_) 05Open→03Resolved a:03jasmine_ Closing as this has been resolved and fixes have been merged, thanks! [15:34:11] (03PS1) 10JHathaway: stdlib: upgrade to v10.1.0 [puppet] - 10https://gerrit.wikimedia.org/r/1345157 (https://phabricator.wikimedia.org/T438912) [15:34:35] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345157 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [15:40:30] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345157 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [15:53:55] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2226 PSU loss of redundancy - https://phabricator.wikimedia.org/T439205#12364770 (10Jhancock.wm) a:03Jhancock.wm power supply actually failed. no spares but under warranty. opened with Dell. SR232215513 [15:54:34] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - Status - issue on logstash2036:9290 - https://phabricator.wikimedia.org/T439196#12364774 (10Jhancock.wm) 05Open→03Resolved a:03Jhancock.wm [15:54:40] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - Status - issue on logstash2036:9290 - https://phabricator.wikimedia.org/T439196#12364776 (10Jhancock.wm) 05Resolved→03Open [15:54:47] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - Status - issue on logstash2036:9290 - https://phabricator.wikimedia.org/T439196#12364777 (10Jhancock.wm) 05Open→03Resolved [15:55:05] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - Status - issue on cirrussearch2079:9290 - https://phabricator.wikimedia.org/T439195#12364780 (10Jhancock.wm) 05Open→03Resolved a:03Jhancock.wm [15:55:16] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - Status - issue on cirrussearch2080:9290 - https://phabricator.wikimedia.org/T439194#12364782 (10Jhancock.wm) 05Open→03Resolved a:03Jhancock.wm [15:55:34] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - Status - issue on wikikube-ctrl2001:9290 - https://phabricator.wikimedia.org/T439193#12364786 (10Jhancock.wm) 05Open→03Resolved a:03Jhancock.wm [15:55:51] 10ops-codfw, 06SRE, 06DC-Ops: Power Supply - Status - issue on cloudbackup2003:9290 - https://phabricator.wikimedia.org/T439192#12364790 (10Jhancock.wm) 05Open→03Resolved a:03Jhancock.wm [15:58:17] (03PS1) 10Ebernhardson: dse-k8s: add opensearch-semantic-search-ssd records [dns] - 10https://gerrit.wikimedia.org/r/1345160 (https://phabricator.wikimedia.org/T438058) [16:11:42] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:22:42] (03PS3) 10Dzahn: add attribution static site to helmfile, create values file [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341335 (https://phabricator.wikimedia.org/T437635) [16:23:49] (03PS1) 10Bking: dse-k8s: Add LVS/services proxy config for new service [puppet] - 10https://gerrit.wikimedia.org/r/1345171 (https://phabricator.wikimedia.org/T438058) [16:25:45] (03PS2) 10Bking: dse-k8s: Add LVS/services proxy config for new service [puppet] - 10https://gerrit.wikimedia.org/r/1345171 (https://phabricator.wikimedia.org/T438058) [16:29:34] 10ops-codfw, 06SRE, 06DC-Ops, 10Prod-Kubernetes, and 2 others: wikikube-worker2280 unreachable - https://phabricator.wikimedia.org/T423395#12364965 (10Jhancock.wm) opened a case: SM2609250589 [16:30:58] (03CR) 10CDanis: "I'm game!" [software/spicerack] - 10https://gerrit.wikimedia.org/r/1120500 (owner: 10Volans) [16:31:39] (03PS3) 10Bking: dse-k8s: Add LVS/services proxy config for new service [puppet] - 10https://gerrit.wikimedia.org/r/1345171 (https://phabricator.wikimedia.org/T438058) [16:41:02] (03PS4) 10Bking: dse-k8s: Add LVS/services proxy config for new service [puppet] - 10https://gerrit.wikimedia.org/r/1345171 (https://phabricator.wikimedia.org/T438058) [16:42:51] (03CR) 10Btullis: [C:03+1] dse-k8s: Add LVS/services proxy config for new service [puppet] - 10https://gerrit.wikimedia.org/r/1345171 (https://phabricator.wikimedia.org/T438058) (owner: 10Bking) [16:47:30] (03PS3) 10Bking: ceph: Enable STS on the codfw rados gateway [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [16:47:35] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [16:52:12] (03PS4) 10Bking: ceph: Enable STS on the codfw rados gateway [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [16:52:49] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [16:57:58] (03CR) 10Ebernhardson: [C:03+1] dse-k8s: Add LVS/services proxy config for new service [puppet] - 10https://gerrit.wikimedia.org/r/1345171 (https://phabricator.wikimedia.org/T438058) (owner: 10Bking) [17:51:24] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:51:49] (03PS1) 10Ssingh: Revert "Failover Puppet to eqiad" [dns] - 10https://gerrit.wikimedia.org/r/1345186 [17:52:20] (03PS1) 10CDanis: Revert "Failover Puppet to eqiad" [dns] - 10https://gerrit.wikimedia.org/r/1345188 [17:53:01] (03Abandoned) 10Ssingh: Revert "Failover Puppet to eqiad" [dns] - 10https://gerrit.wikimedia.org/r/1345186 (owner: 10Ssingh) [17:53:04] (03CR) 10Ssingh: [C:03+1] Revert "Failover Puppet to eqiad" [dns] - 10https://gerrit.wikimedia.org/r/1345188 (owner: 10CDanis) [17:55:48] (03CR) 10CDanis: [C:03+2] Revert "Failover Puppet to eqiad" [dns] - 10https://gerrit.wikimedia.org/r/1345188 (owner: 10CDanis) [17:56:00] (03PS2) 10CDanis: Revert "Failover Puppet to eqiad" [dns] - 10https://gerrit.wikimedia.org/r/1345188 [17:57:01] (03CR) 10CDanis: [V:03+2 C:03+2] Revert "Failover Puppet to eqiad" [dns] - 10https://gerrit.wikimedia.org/r/1345188 (owner: 10CDanis) [17:57:15] !log cdanis@dns1004 START - running authdns-update [17:58:34] (03CR) 10Andrew Bogott: [C:03+2] cloud-vps eqiad1: Enable zookeeper logs [puppet] - 10https://gerrit.wikimedia.org/r/1344721 (https://phabricator.wikimedia.org/T435503) (owner: 10Andrew Bogott) [17:59:40] !log cdanis@dns1004 END - running authdns-update [18:06:27] FIRING: [2x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:27:18] !log krinkle@deploy1003 Started deploy [statsv/statsv@df3ebff]: T439183: Accept dot, plus, hyphen in label values [18:27:21] T439183: statsv: Counter and timer lines are discarded when a label value contains ".", "+" or "-" - https://phabricator.wikimedia.org/T439183 [18:27:29] !log krinkle@deploy1003 Finished deploy [statsv/statsv@df3ebff]: T439183: Accept dot, plus, hyphen in label values (duration: 00m 11s) [19:21:26] (03PS1) 10Ahmon Dancy: kubernetes.yaml: Revise pretrain logstash thresholds [puppet] - 10https://gerrit.wikimedia.org/r/1345206 (https://phabricator.wikimedia.org/T435419) [19:21:34] (03CR) 10Bking: [C:03+1] ceph: Enable STS on the codfw rados gateway [puppet] - 10https://gerrit.wikimedia.org/r/1345117 (https://phabricator.wikimedia.org/T435590) (owner: 10Btullis) [19:36:54] (03CR) 10Bking: [C:03+1] dse-k8s: add opensearch-semantic-search-ssd records [dns] - 10https://gerrit.wikimedia.org/r/1345160 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [19:40:32] (03PS1) 10Dzahn: zookeeper: remove need to opt-in to enable logging on trixie and beyond [puppet] - 10https://gerrit.wikimedia.org/r/1345209 (https://phabricator.wikimedia.org/T435503) [19:41:22] FIRING: [2x] GnmiInterfaceCountersDrop: cr1-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [19:46:22] RESOLVED: [2x] GnmiInterfaceCountersDrop: cr1-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [20:09:40] (03CR) 10Ottomata: [beta eventgate] Migrate urls to deployment-eventgate05 (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345154 (https://phabricator.wikimedia.org/T429497) (owner: 10TChin) [20:10:33] PROBLEM - MariaDB Replica Lag: s1 on db1240 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 620.11 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [20:10:59] PROBLEM - MariaDB Replica Lag: s8 on db1285 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 649.04 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [20:11:15] PROBLEM - MariaDB Replica Lag: s8 on db2198 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 657.97 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [20:11:15] PROBLEM - MariaDB Replica Lag: s1 on db2250 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 657.08 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [20:26:45] (03PS1) 10Bking: WIP: Add Matomo frontend helm chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1345213 (https://phabricator.wikimedia.org/T436003) [20:31:05] (03PS1) 10Aqu: sqoop_mediawiki: Add weekly change_tag sqoop [puppet] - 10https://gerrit.wikimedia.org/r/1345214 (https://phabricator.wikimedia.org/T437961) [20:34:07] PROBLEM - snapshot of s6 in codfw on backupmon1001 is CRITICAL: snapshot for s6 at codfw (db2197) taken more than 3 days ago: Most recent backup 2026-09-22 20:31:50 https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [20:37:49] (03PS2) 10TChin: [beta eventgate] Migrate urls to deployment-eventgate05 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345154 (https://phabricator.wikimedia.org/T429497) [20:39:01] (03CR) 10TChin: [beta eventgate] Migrate urls to deployment-eventgate05 (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345154 (https://phabricator.wikimedia.org/T429497) (owner: 10TChin) [20:44:25] (03PS3) 10AKhatun: article-feature-counts: add deployment chart files [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343044 (https://phabricator.wikimedia.org/T437000) [20:45:39] (03CR) 10Ottomata: [C:03+1] [beta eventgate] Migrate urls to deployment-eventgate05 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345154 (https://phabricator.wikimedia.org/T429497) (owner: 10TChin) [20:55:14] (03PS1) 10Krinkle: File: Implement LocalFile::getUrlForPurge workaround for caching [core] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1345216 (https://phabricator.wikimedia.org/T425216) [21:01:51] (03PS1) 10Andrea Denisse: alert: add a dummy Klaxon managers incident URL [labs/private] - 10https://gerrit.wikimedia.org/r/1345218 [21:01:51] (03CR) 10Andrea Denisse: "Noticed this secret was missing, PCC didn't complain because it's optional but it also won't detect any changes regarding it." [labs/private] - 10https://gerrit.wikimedia.org/r/1345218 (owner: 10Andrea Denisse) [21:03:03] (03CR) 10CI reject: [V:04-1] File: Implement LocalFile::getUrlForPurge workaround for caching [core] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1345216 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [21:04:07] RECOVERY - snapshot of s6 in codfw on backupmon1001 is OK: Last snapshot for s6 at codfw (db2197) taken on 2026-09-25 20:32:28 (456 GiB, +0.1 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [21:04:07] PROBLEM - snapshot of s5 in codfw on backupmon1001 is CRITICAL: snapshot for s5 at codfw (db2201) taken more than 3 days ago: Most recent backup 2026-09-22 20:36:00 https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [21:06:47] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 28 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1345052 (https://phabricator.wikimedia.org/T439159) (owner: 10Ameisenigel) [21:07:14] (03PS2) 10Andrea Denisse: kubernetes: add dummy secrets for oncall [labs/private] - 10https://gerrit.wikimedia.org/r/1345219 (https://phabricator.wikimedia.org/T438803) [21:07:35] (03CR) 10Andrea Denisse: "Adding the same values as the prod alert host." [labs/private] - 10https://gerrit.wikimedia.org/r/1345219 (https://phabricator.wikimedia.org/T438803) (owner: 10Andrea Denisse) [21:07:57] (03CR) 10Krinkle: "recheck" [core] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1345216 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [21:08:22] 10ops-eqiad, 06DC-Ops: Unresponsive management for krb1002.mgmt:22 - https://phabricator.wikimedia.org/T439280 (10phaultfinder) 03NEW [21:10:25] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - thanos-query_443: Servers titan1002.eqiad.wmnet are marked down but pooled: thanos-web_443: Servers titan1001.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [21:10:37] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - thanos-query_443: Servers titan1001.eqiad.wmnet are marked down but pooled: thanos-web_443: Servers titan1002.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [21:11:59] RECOVERY - MariaDB Replica Lag: s8 on db1285 is OK: OK slave_sql_lag Replication lag: 0.08 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:12:10] FIRING: BFDdown: BFD session down between cr1-codfw and fe80::b6f9:5d07:1130:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:12:13] PROBLEM - Host titan1002 is DOWN: PING CRITICAL - Packet loss = 100% [21:12:25] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [21:12:37] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [21:12:45] RECOVERY - Host titan1002 is UP: PING OK - Packet loss = 0%, RTA = 0.18 ms [21:13:15] RECOVERY - MariaDB Replica Lag: s8 on db2198 is OK: OK slave_sql_lag Replication lag: 0.30 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:13:27] FIRING: [3x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:14:18] (03PS1) 10Andrea Denisse: oncall: add discovery ingress records [dns] - 10https://gerrit.wikimedia.org/r/1345220 (https://phabricator.wikimedia.org/T439279) [21:15:32] (03CR) 10CDanis: [C:03+1] oncall: Add Klaxon to the aux clusters (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344072 (https://phabricator.wikimedia.org/T438897) (owner: 10Andrea Denisse) [21:15:41] RESOLVED: [3x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:16:53] (03PS1) 10Andrea Denisse: wikimedia.org: add oncall [dns] - 10https://gerrit.wikimedia.org/r/1345221 (https://phabricator.wikimedia.org/T439279) [21:17:10] RESOLVED: BFDdown: BFD session down between cr1-codfw and fe80::b6f9:5d07:1130:c93f - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [21:23:59] (03PS3) 10Ameisenigel: Enable Extension:Translate and Extension:TranslationNotifications on co.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344216 (https://phabricator.wikimedia.org/T386776) [21:24:15] PROBLEM - MariaDB Replica Lag: x3 on db2200 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 623.28 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:27:28] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 28 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344216 (https://phabricator.wikimedia.org/T386776) (owner: 10Ameisenigel) [21:33:07] 06SRE, 06serviceops-deprecated, 07Wikimedia-Performance-recommendation: Evaluate using igbinary for MW php-apcu at WMF - https://phabricator.wikimedia.org/T225074#12365625 (10MGoncalves-WMF) == APCu serializer evaluation: local results == We evaluated PHP's native APCu serializer against igbinary in a local... [21:34:07] RECOVERY - snapshot of s5 in codfw on backupmon1001 is OK: Last snapshot for s5 at codfw (db2201) taken on 2026-09-25 20:36:48 (539 GiB, +0.1 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [21:38:35] PROBLEM - snapshot of s8 in codfw on backupmon1001 is CRITICAL: snapshot for s8 at codfw (db2198) taken more than 3 days ago: Most recent backup 2026-09-22 21:11:18 https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [21:40:15] RECOVERY - MariaDB Replica Lag: x3 on db2200 is OK: OK slave_sql_lag Replication lag: 0.13 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:42:49] PROBLEM - snapshot of x3 in codfw on backupmon1001 is CRITICAL: snapshot for x3 at codfw (db2200) taken more than 3 days ago: Most recent backup 2026-09-22 21:38:37 https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [21:44:31] PROBLEM - MariaDB Replica Lag: x3 on db1216 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 607.94 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:46:52] (03CR) 10Andrea Denisse: [V:03+2 C:03+2] kubernetes: add dummy secrets for oncall [labs/private] - 10https://gerrit.wikimedia.org/r/1345219 (https://phabricator.wikimedia.org/T438803) (owner: 10Andrea Denisse) [21:46:55] (03CR) 10Dzahn: "thoughts or hints who I should ask if any?" [puppet] - 10https://gerrit.wikimedia.org/r/1345209 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [21:49:18] (03PS1) 10Andrea Denisse: Revert "ceph: Move the dummy STS key to the correct hiera file" [labs/private] - 10https://gerrit.wikimedia.org/r/1345227 [21:49:31] (03CR) 10Andrea Denisse: [V:03+2 C:03+2] Revert "ceph: Move the dummy STS key to the correct hiera file" [labs/private] - 10https://gerrit.wikimedia.org/r/1345227 (owner: 10Andrea Denisse) [21:50:15] (03PS1) 10Andrea Denisse: Revert^2 "ceph: Move the dummy STS key to the correct hiera file" [labs/private] - 10https://gerrit.wikimedia.org/r/1345228 [21:51:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:51:33] RECOVERY - MariaDB Replica Lag: s1 on db1240 is OK: OK slave_sql_lag Replication lag: 0.08 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:51:55] (03PS2) 10Andrea Denisse: ceph: Move the dummy STS key to the correct hiera file [labs/private] - 10https://gerrit.wikimedia.org/r/1345228 (https://phabricator.wikimedia.org/T435590) [21:53:15] RECOVERY - MariaDB Replica Lag: s1 on db2250 is OK: OK slave_sql_lag Replication lag: 0.28 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:56:31] RECOVERY - MariaDB Replica Lag: x3 on db1216 is OK: OK slave_sql_lag Replication lag: 0.28 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [21:57:41] PROBLEM - snapshot of s2 in codfw on backupmon1001 is CRITICAL: snapshot for s2 at codfw (db2197) taken more than 3 days ago: Most recent backup 2026-09-22 21:52:21 https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [22:02:49] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:03:05] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:03:09] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:06:27] FIRING: [2x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:06:39] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:06:55] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:06:59] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.137 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:08:35] RECOVERY - snapshot of s8 in codfw on backupmon1001 is OK: Last snapshot for s8 at codfw (db2198) taken on 2026-09-25 21:11:00 (1082 GiB, +0.1 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [22:12:49] RECOVERY - snapshot of x3 in codfw on backupmon1001 is OK: Last snapshot for x3 at codfw (db2200) taken on 2026-09-25 21:39:21 (363 GiB, +0.0 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [22:14:30] (03CR) 10Andrew Bogott: [C:04-1] "I probably don't need this anymore" [puppet] - 10https://gerrit.wikimedia.org/r/1344324 (https://phabricator.wikimedia.org/T421911) (owner: 10Andrew Bogott) [22:16:05] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:16:09] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:16:55] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:16:59] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.135 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:19:37] (03PS1) 10Andrea Denisse: service: move oncall to production [puppet] - 10https://gerrit.wikimedia.org/r/1345232 (https://phabricator.wikimedia.org/T439279) [22:19:37] (03CR) 10Andrea Denisse: [C:04-2] "We must merge this after the DNS patches are applied." [puppet] - 10https://gerrit.wikimedia.org/r/1345232 (https://phabricator.wikimedia.org/T439279) (owner: 10Andrea Denisse) [22:20:05] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:20:09] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:21:49] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:22:57] PROBLEM - MariaDB Replica Lag: s4 on db1265 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 614.42 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [22:23:45] PROBLEM - MariaDB Replica Lag: s4 on db2239 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 637.07 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [22:23:50] (03PS1) 10Andrea Denisse: trafficserver: add mapping for oncall [puppet] - 10https://gerrit.wikimedia.org/r/1345234 (https://phabricator.wikimedia.org/T439279) [22:23:50] (03CR) 10Andrea Denisse: [C:04-2] "This must be merged after the DNS patches are applied and after SSO is enabled." [puppet] - 10https://gerrit.wikimedia.org/r/1345234 (https://phabricator.wikimedia.org/T439279) (owner: 10Andrea Denisse) [22:24:05] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 4.432 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:24:05] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:28:39] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:28:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:29:05] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:29:09] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:32:49] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:33:23] 10ops-eqiad, 06SRE, 06DC-Ops, 06cloud-services-team (Hardware), 13Patch-For-Review: Q3:rack/setup/install cloudcephosd105[3456] - https://phabricator.wikimedia.org/T419892#12365789 (10Jclark-ctr) [22:33:39] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:36:49] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:39:12] !log jclark@cumin1004 START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm [22:39:25] 10ops-eqiad, 06SRE, 06DC-Ops, 06cloud-services-team (Hardware), 13Patch-For-Review: Q3:rack/setup/install cloudcephosd105[3456] - https://phabricator.wikimedia.org/T419892#12365795 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jclark@cumin1004 for host cloudcephosd1055.eqiad.... [22:42:24] 10ops-codfw, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q3:rack/setup/install ms-fe202[5-8] - https://phabricator.wikimedia.org/T439292 (10RobH) 03NEW [22:42:37] 10ops-codfw, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q3:rack/setup/install ms-fe202[5-8] - https://phabricator.wikimedia.org/T439292#12365813 (10RobH) [22:43:15] 10ops-codfw, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q3:rack/setup/install ms-fe202[5-8] - https://phabricator.wikimedia.org/T439292#12365814 (10RobH) a:03MatthewVernon Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/SRE/Dc-op... [22:43:39] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:44:48] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q3:rack/setup/install ms-fe102[5-8] - https://phabricator.wikimedia.org/T439293 (10RobH) 03NEW [22:45:17] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q3:rack/setup/install ms-fe102[5-8] - https://phabricator.wikimedia.org/T439293#12365836 (10RobH) a:03MatthewVernon Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/... [22:45:31] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q3:rack/setup/install ms-fe102[5-8] - https://phabricator.wikimedia.org/T439293#12365844 (10RobH) [22:45:57] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:45:59] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.138 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:46:20] !log ssastry@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [22:46:49] !log ssastry@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [22:46:50] !log ssastry@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [22:47:30] !log ssastry@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [22:51:51] !log ssastry@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [22:51:54] !log ssastry@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [22:51:55] !log ssastry@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [22:51:58] !log ssastry@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [22:57:08] jclark@cumin1004 reimage (PID 912937) is awaiting input [22:57:41] RECOVERY - snapshot of s2 in codfw on backupmon1001 is OK: Last snapshot for s2 at codfw (db2197) taken on 2026-09-25 21:53:33 (823 GiB, +0.2 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [22:58:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [23:04:23] PROBLEM - Host wikikube-worker1016 is DOWN: PING CRITICAL - Packet loss = 50%, RTA = 2640.80 ms [23:04:33] RECOVERY - Host wikikube-worker1016 is UP: PING OK - Packet loss = 0%, RTA = 0.27 ms [23:12:49] PROBLEM - snapshot of s7 in codfw on backupmon1001 is CRITICAL: snapshot for s7 at codfw (db2198) taken more than 3 days ago: Most recent backup 2026-09-22 23:05:33 https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [23:15:34] !log jclark@cumin1004 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm [23:15:39] 10ops-eqiad, 06SRE, 06DC-Ops, 06cloud-services-team (Hardware), 13Patch-For-Review: Q3:rack/setup/install cloudcephosd105[3456] - https://phabricator.wikimedia.org/T419892#12365899 (10Jclark-ctr) @fgiunchedi I have not had any luck so far. The other Dell cloudcephosd105 servers have RAID1 for the two sma... [23:15:47] 10ops-eqiad, 06SRE, 06DC-Ops, 06cloud-services-team (Hardware), 13Patch-For-Review: Q3:rack/setup/install cloudcephosd105[3456] - https://phabricator.wikimedia.org/T419892#12365900 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jclark@cumin1004 for host cloudcephosd1055.eqiad.wmne... [23:34:07] PROBLEM - snapshot of x1 in codfw on backupmon1001 is CRITICAL: snapshot for x1 at codfw (db2197) taken more than 3 days ago: Most recent backup 2026-09-22 23:08:10 https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [23:38:13] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, and 3 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12365931 (10BBlack) >>! In T425216#12360080, @Krinkle wrote: > [,,, re: canonical ...] > Is this an inten... [23:38:56] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1345242 [23:38:56] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1345242 (owner: 10TrainBranchBot) [23:47:27] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1345242 (owner: 10TrainBranchBot) [23:59:45] RECOVERY - MariaDB Replica Lag: s4 on db2239 is OK: OK slave_sql_lag Replication lag: 0.17 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response