[00:07:03] !log jclark@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sessionstore1005.eqiad.wmnet with OS bookworm [00:07:09] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12320557 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jclark@cumin1003 for host sessionstore1005.eqiad.wmnet with OS bookworm executed with errors: -... [00:17:36] !log jclark@cumin1003 START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm [00:17:46] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12320577 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jclark@cumin1003 for host sessionstore1005.eqiad.wmnet with OS bookworm [00:19:45] !log jclark@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage [00:23:37] !log jclark@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage [00:33:24] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 15 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334946 (https://phabricator.wikimedia.org/T434487) (owner: 10Sbisson) [00:43:42] !log jclark@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm [00:43:56] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12320610 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jclark@cumin1003 for host sessionstore1005.eqiad.wmnet with OS bookworm completed: - sessionsto... [00:48:00] (03CR) 10Lerickson: "Thanks! Just a couple suggestions to make the comments clearer IMO." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1336094 (https://phabricator.wikimedia.org/T428631) (owner: 10Gmodena) [00:50:07] FIRING: ProbeDown: Service sessionstore1005-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://wikitech.wikimedia.org/wiki/TLS/Runbook#sessionstore1005-a:9042 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [00:55:07] FIRING: [2x] ProbeDown: Service sessionstore1005-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [00:59:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:04:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.79% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:05:07] FIRING: [2x] ProbeDown: Service sessionstore1005-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [01:06:33] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 53604520 and 3 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [01:07:33] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 8272 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [01:10:10] (03PS1) 10TrainBranchBot: Branch commit for wmf/1.47.0-wmf.20 [core] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341449 (https://phabricator.wikimedia.org/T430839) [01:10:12] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/1.47.0-wmf.20 [core] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341449 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [01:11:08] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1341452 [01:11:08] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1341452 (owner: 10TrainBranchBot) [01:16:12] 06SRE, 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Requesting access to deployment and analytics-privatedata-users for Cklimas - https://phabricator.wikimedia.org/T437653#12320638 (10Dzahn) a:05thcipriani→03None [01:16:18] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for EAlbizzati-WMF - https://phabricator.wikimedia.org/T437741#12320640 (10Dzahn) a:05thcipriani→03None [01:16:40] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12320641 (10Dzahn) a:05thcipriani→03None [01:16:48] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to releasers-mobile for LPetty-WMF - https://phabricator.wikimedia.org/T437662#12320642 (10Dzahn) a:05thcipriani→03None [01:17:10] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to releasers-mobile for LPetty-WMF - https://phabricator.wikimedia.org/T437662#12320644 (10Dzahn) [01:17:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 10.92% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:17:20] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12320645 (10Dzahn) [01:17:29] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for EAlbizzati-WMF - https://phabricator.wikimedia.org/T437741#12320646 (10Dzahn) [01:17:51] 06SRE, 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Requesting access to deployment and analytics-privatedata-users for Cklimas - https://phabricator.wikimedia.org/T437653#12320648 (10Dzahn) [01:18:12] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to releasers-mobile for LPetty-WMF - https://phabricator.wikimedia.org/T437662#12320649 (10Dzahn) [01:18:54] (03Merged) 10jenkins-bot: Branch commit for wmf/1.47.0-wmf.20 [core] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341449 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [01:19:39] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for cooltey - https://phabricator.wikimedia.org/T437658#12320652 (10Dzahn) [01:19:46] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for cooltey - https://phabricator.wikimedia.org/T437658#12320653 (10Dzahn) a:05thcipriani→03None [01:21:12] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1341452 (owner: 10TrainBranchBot) [01:42:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.86% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:00:05] Deploy window Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T0200) [02:00:05] Deploy window Automatic deployment of MediaWiki to pretrain wikis - see mw:Pretrain (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T0200) [02:01:17] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:05:07] RESOLVED: ProbeDown: Service sessionstore1005-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://wikitech.wikimedia.org/wiki/TLS/Runbook#sessionstore1005-a:9042 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [02:08:39] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 22s) [02:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:58] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to analytics-privatedata-users, growthbook-customelevatedaccess for Anil - https://phabricator.wikimedia.org/T437611#12320731 (10AKanji-WMF) Thank you so much @ssingh and @Dzahn - I am able to access Growthbook. [02:35:39] RECOVERY - dump of es7 in codfw on backupmon1001 is OK: Last dump for es7 at codfw (es2040) taken on 2026-09-15 00:00:02 (529 GiB, +4.6 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [02:40:41] RECOVERY - dump of es7 in eqiad on backupmon1001 is OK: Last dump for es7 at eqiad (es1040) taken on 2026-09-15 00:00:02 (529 GiB, +4.6 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [02:54:33] (03PS1) 10JHathaway: cron_splay: allow splaying over a single host [puppet] - 10https://gerrit.wikimedia.org/r/1341514 [02:54:51] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1341514 (owner: 10JHathaway) [02:55:39] RECOVERY - dump of es6 in codfw on backupmon1001 is OK: Last dump for es6 at codfw (es2036) taken on 2026-09-15 00:00:02 (529 GiB, +4.6 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [03:00:05] Deploy window Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T0300) [03:00:39] RECOVERY - dump of es6 in eqiad on backupmon1001 is OK: Last dump for es6 at eqiad (es1036) taken on 2026-09-15 00:00:02 (529 GiB, +4.6 %) https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [03:01:53] (03PS1) 10TrainBranchBot: testwikis to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341519 (https://phabricator.wikimedia.org/T430839) [03:01:56] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by mwpresync@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341519 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [03:02:54] (03Merged) 10jenkins-bot: testwikis to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341519 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [03:03:17] !log mwpresync@deploy1003 Started scap sync-world: testwikis to 1.47.0-wmf.20 refs T430839 [03:03:21] T430839: 1.47.0-wmf.20 deployment blockers - https://phabricator.wikimedia.org/T430839 [03:06:00] !log mwpresync@deploy1003 sync-world failed: Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted [03:06:00] /mediawiki-multiversion --singleversion-image-basename docker-registry.discovery.wmnet/restricted/mediawiki-singleversion --webserver-image-name docker-registry.discovery.wmnet/restricted/mediawiki-webserver --latest-tag latest --label vnd.wikimedia.builder.name=scap --label vnd.wikimedia.builder.version=4.289.0 --label vnd.wikimedia.scap.stage_dir=/srv/mediawiki-staging --label vnd.wikimedia.scap.build_state_dir=/srv/med [03:06:00] iawiki-staging/scap/image-build' returned non-zero exit status 1. (scap version: 4.289.0) (duration: 02m 42s) [03:10:32] (03PS2) 10JHathaway: cron_splay: allow splaying over a single host [puppet] - 10https://gerrit.wikimedia.org/r/1341514 [03:10:37] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1341514 (owner: 10JHathaway) [03:11:22] (03CR) 10JHathaway: "Feel free to merge this in, if it looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1341514 (owner: 10JHathaway) [03:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:46:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [03:49:38] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [03:59:39] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:00:05] Deploy window Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T0400) [04:07:12] !log mwpresync@deploy1003 Pruned MediaWiki: 1.47.0-wmf.17 (duration: 07m 10s) [04:16:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [04:31:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.11% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [04:42:14] (03CR) 10Samwilson: [C:03+2] "Looks good to me. I've been trying to get it running locally but to no avail (for unrelated reasons), so shall trust in the cloud gods (i." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341399 (https://phabricator.wikimedia.org/T433964) (owner: 10Ladsgroup-claude) [04:51:40] (03Merged) 10jenkins-bot: Introduce concept of thumbnails with expiry for recently uploaded files [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341399 (https://phabricator.wikimedia.org/T433964) (owner: 10Ladsgroup-claude) [04:55:50] (03Abandoned) 10Giuseppe Lavagetto: php: remove stale images [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1340948 (owner: 10Giuseppe Lavagetto) [04:57:32] (03CR) 10Giuseppe Lavagetto: [C:03+2] shellbox: add gVisor support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333139 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [05:00:22] (03Merged) 10jenkins-bot: shellbox: add gVisor support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333139 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [05:01:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:04:39] (03PS1) 10Giuseppe Lavagetto: shellbox-timeline: enable gVisor in staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341599 (https://phabricator.wikimedia.org/T436649) [05:04:42] (03PS1) 10Giuseppe Lavagetto: shellbox-timeline: enable gVisor everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341600 (https://phabricator.wikimedia.org/T436649) [05:06:40] (03PS4) 10Hashar: rake_modules: support Debian 13 (Trixie) facts [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) [05:06:47] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-timeline: apply [05:06:51] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply [05:06:52] !log oblivian@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply [05:06:59] !log oblivian@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply [05:07:00] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply [05:07:03] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply [05:07:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.21% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:07:47] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox: apply [05:07:51] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox: apply [05:07:52] !log oblivian@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox: apply [05:07:59] !log oblivian@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox: apply [05:08:00] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox: apply [05:08:03] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox: apply [05:08:12] (03CR) 10Hashar: rake_modules: default to test on Bullseye & Bookworm (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329251 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [05:09:14] (03Abandoned) 10Hashar: rake_modules: default to test on Bullseye & Bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1329251 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [05:09:48] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-constraints: apply [05:09:52] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply [05:09:53] !log oblivian@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply [05:10:00] !log oblivian@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply [05:10:01] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply [05:10:04] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply [05:10:06] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-media: apply [05:10:10] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-media: apply [05:10:11] !log oblivian@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-media: apply [05:10:17] !log oblivian@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply [05:10:19] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-media: apply [05:10:22] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply [05:10:23] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply [05:10:27] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [05:10:28] !log oblivian@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply [05:10:35] !log oblivian@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [05:10:36] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply [05:10:40] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [05:10:41] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-video: apply [05:10:45] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-video: apply [05:10:46] !log oblivian@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-video: apply [05:10:52] !log oblivian@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply [05:10:54] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-video: apply [05:10:57] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply [05:12:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:12:55] (03CR) 10Hashar: "> All the depends on commits have been merged!" [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [05:15:08] (03CR) 10Giuseppe Lavagetto: [C:03+2] shellbox-timeline: enable gVisor in staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341599 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [05:17:33] (03Merged) 10jenkins-bot: shellbox-timeline: enable gVisor in staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341599 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [05:27:58] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-timeline: apply [05:30:49] <_joe_> I'm dumb, I forgot to merge a change, sigh [05:34:26] (03PS2) 10Giuseppe Lavagetto: shellbox-timeline: enable gVisor everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341600 (https://phabricator.wikimedia.org/T436649) [05:34:26] (03PS1) 10Giuseppe Lavagetto: gVisor: enable in staging-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341609 [05:34:40] (03CR) 10Giuseppe Lavagetto: [C:03+2] gVisor: enable in staging-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341609 (owner: 10Giuseppe Lavagetto) [05:37:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:38:05] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply [05:42:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.48% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:42:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.11% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:45:13] (03Merged) 10jenkins-bot: gVisor: enable in staging-eqiad [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341609 (owner: 10Giuseppe Lavagetto) [05:47:34] !log oblivian@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [05:47:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.48% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:48:40] !log oblivian@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [05:49:15] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-timeline: apply [05:59:22] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T0600) [06:00:05] marostegui, cezmunsta, and federico3: I, the Bot under the Fountain, call upon thee, The Deployer, to do Primary database switchover deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T0600). [06:03:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:08:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:29:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.11% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:37:04] (03CR) 10Muehlenhoff: [C:03+1] "Looks good!" [puppet] - 10https://gerrit.wikimedia.org/r/1341249 (https://phabricator.wikimedia.org/T428555) (owner: 10Elukey) [06:37:31] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 15 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [extensions/WikimediaMessages] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1341254 (https://phabricator.wikimedia.org/T437183) (owner: 10STran) [06:40:22] (03PS4) 10Muehlenhoff: Remove python-build Bullseye image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339635 (https://phabricator.wikimedia.org/T416452) [06:43:25] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Remove python-build Bullseye image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339635 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [06:48:59] (03CR) 10Slyngshede: "That's a good point. This was focused on fixing the deployment and testing environments, which doesn't have the mmdb files." [puppet] - 10https://gerrit.wikimedia.org/r/1224897 (https://phabricator.wikimedia.org/T414111) (owner: 10Slyngshede) [06:49:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.42% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:49:37] (03PS1) 10Muehlenhoff: Enable gvisor sync definition also for Bookworm/Trixie [puppet] - 10https://gerrit.wikimedia.org/r/1341645 [06:49:44] (03PS4) 10JMeybohm: k8s: Add a script to sync kubelet labels with the API server [puppet] - 10https://gerrit.wikimedia.org/r/1341181 [06:49:54] (03CR) 10JMeybohm: k8s: Add a script to sync kubelet labels with the API server (036 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1341181 (owner: 10JMeybohm) [06:50:54] (03CR) 10JMeybohm: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1341181 (owner: 10JMeybohm) [06:50:55] (03CR) 10Giuseppe Lavagetto: [C:03+2] Enable gvisor sync definition also for Bookworm/Trixie [puppet] - 10https://gerrit.wikimedia.org/r/1341645 (owner: 10Muehlenhoff) [06:51:02] !log pruned obsolete Bullseye image python3-build-bullseye from the docker registry T416452 [06:51:06] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:51:06] T416452: Migrate Docker images running in Production away from Bullseye - https://phabricator.wikimedia.org/T416452 [06:51:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:51:33] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] "python3-build-bullseye was pruned from the Docker registry" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339635 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [06:56:48] (03PS3) 10Muehlenhoff: Remove python-devel container image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339636 (https://phabricator.wikimedia.org/T416452) [07:00:04] Amir1, urbanecm, and awight: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T0700). [07:00:05] Tran: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:10] o/ [07:00:17] All of my patches go together and I can self-deploy [07:01:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.62% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:01:38] starting [07:02:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1341274 (https://phabricator.wikimedia.org/T437183) (owner: 10STran) [07:02:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341242 (https://phabricator.wikimedia.org/T437183) (owner: 10STran) [07:02:05] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [extensions/WikimediaMessages] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1341254 (https://phabricator.wikimedia.org/T437183) (owner: 10STran) [07:03:21] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] "Given debmonitor shows these as running Trixie, I'll go ahead and merge" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339636 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [07:04:00] (03Merged) 10jenkins-bot: SuggestedInvestigations: Add and enable 'sockpuppets' queue view [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341242 (https://phabricator.wikimedia.org/T437183) (owner: 10STran) [07:06:40] !log pruned obsolete Bullseye image python3-devel from the docker registry T416452 [07:06:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:06:44] T416452: Migrate Docker images running in Production away from Bullseye - https://phabricator.wikimedia.org/T416452 [07:06:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [07:08:41] (03PS1) 10Giuseppe Lavagetto: containerd: fix runsc config path definition [puppet] - 10https://gerrit.wikimedia.org/r/1341663 [07:09:16] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] "python3-devel has been pruned from the Docker registry" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339636 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [07:09:58] (03Merged) 10jenkins-bot: SI: Implement "queue view" functionality [extensions/CheckUser] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1341274 (https://phabricator.wikimedia.org/T437183) (owner: 10STran) [07:10:01] (03Merged) 10jenkins-bot: Add wmf-specific Special:SuggestedInvestigations messages [extensions/WikimediaMessages] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1341254 (https://phabricator.wikimedia.org/T437183) (owner: 10STran) [07:11:09] (03PS1) 10Muehlenhoff: Remove obsolete nutcracker container image based on bullseye [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341665 (https://phabricator.wikimedia.org/T416452) [07:11:12] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9420/co" [puppet] - 10https://gerrit.wikimedia.org/r/1341663 (owner: 10Giuseppe Lavagetto) [07:11:32] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1341274|SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242|SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254|Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] [07:11:36] T437183: Likely socks queue - https://phabricator.wikimedia.org/T437183 [07:11:52] (03CR) 10Giuseppe Lavagetto: [V:03+1 C:03+2] containerd: fix runsc config path definition [puppet] - 10https://gerrit.wikimedia.org/r/1341663 (owner: 10Giuseppe Lavagetto) [07:14:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [07:15:17] Tran: just a heads-up. Image building failed last night during the train presync, which means the image building part of your backpart may take longer than usual [07:15:52] ack, thanks for the heads up [07:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:18:04] !log oblivian@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-timeline: apply [07:18:20] !log oblivian@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply [07:19:32] (03PS3) 10Muehlenhoff: Remove obsolete buildkitd image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339620 (https://phabricator.wikimedia.org/T416452) [07:19:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [07:21:43] 06SRE, 06RoadToWiki, 06Traffic, 10WMF-General-or-Unknown: Domain "wikipedia.ar" redirects to Arabic instead of Spanish - https://phabricator.wikimedia.org/T437799#12320964 (10CoderAnimeshPathak) [07:23:28] (03CR) 10Krinkle: profile::mediawiki::php: Add unserialize_callback_func php.ini var (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1337377 (https://phabricator.wikimedia.org/T389402) (owner: 10Gergő Tisza) [07:27:30] (03Abandoned) 10STran: SI: Implement "queue view" data functionality [extensions/CheckUser] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1341253 (https://phabricator.wikimedia.org/T437183) (owner: 10STran) [07:31:20] !log stran@deploy1003 stran: Backport for [[gerrit:1341274|SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242|SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254|Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:31:24] T437183: Likely socks queue - https://phabricator.wikimedia.org/T437183 [07:31:59] testing now [07:36:09] lgtm, continuing [07:37:45] hm? are servers down? I'm seeing the following error: https://test.wikipedia.org/wiki/Special:SpecialPages (/srv/deployment/httpbb-tests/pretrain/test_testwiki.yaml:62) Status code: expected 200, got 500. [07:39:37] Hm no, it's an actual error but afaict my patch in that repo shouldn't have caused it... [07:40:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:40:53] (03CR) 10Filippo Giunchedi: Export a few stats about the magnum capi worker cluster (037 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1339182 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [07:45:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:49:03] (03CR) 10Jelto: [C:03+1] "lgtm, thank you" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341334 (https://phabricator.wikimedia.org/T437635) (owner: 10Dzahn) [07:49:38] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [07:49:41] jnuche: if you're still around, could I get a second opinion? The error I'm seeing suggests that the .20 branch is going to throw an error from WikimediaCustomizations: TypeError: MediaWiki\Rest\Module\ModuleManager::__construct(): Argument #4 ($jsonLocalizer) must be of type MediaWiki\Rest\JsonLocalizer, MediaWiki\Rest\ResponseFactory given, called in [07:49:41] /srv/mediawiki/php-1.47.0-wmf.20/extensions/WikimediaCustomizations/src/RestSandbox/SpecialRestSandbox.php on line 41 but afaict, it's not my caused by my backport to .19 which I currently am midway through. Should I just continue or should I rollback? [07:50:02] I'm failing the ping to testwiki because it's on the .20 branch [07:54:24] Tran: I'd say finish your backport, I have to do the .20 train right after you and I assume the same problem will show up. I'll handle it then [07:54:37] !log stran@deploy1003 stran: Continuing with deployment [07:54:46] thanks, continuing then [07:58:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.82% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:59:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:00:05] jnuche and dduvall: Your horoscope predicts another MediaWiki train - Utc-0+Utc-7 Version deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T0800). [08:01:05] morning, waiting on backport to finish before moving ahead with the train (and presumably failing due to the above) [08:03:44] this change looks like a cause: https://gerrit.wikimedia.org/r/plugins/gitiles/mediawiki/core/+/13c20f9ca6486be088e902c453c8d07cb1f6682c [08:06:17] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [08:06:59] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341274|SI: Implement "queue view" functionality (T437183)]], [[gerrit:1341242|SuggestedInvestigations: Add and enable 'sockpuppets' queue view (T437183)]], [[gerrit:1341254|Add wmf-specific Special:SuggestedInvestigations messages (T437183)]] (duration: 55m 27s) [08:07:03] T437183: Likely socks queue - https://phabricator.wikimedia.org/T437183 [08:07:12] jnuche: backport done, thanks for your patience [08:07:40] Tran: np, glad you could deploy [08:08:27] testwikis is already on .20. I think there's no point in trying to deploy the train at this point. I'm going to create a blocker and signal the train is stuck [08:10:52] (03CR) 10Dpogorzelski: [C:03+2] dns: add liftwing-studio records [dns] - 10https://gerrit.wikimedia.org/r/1341128 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:13:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.18% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:13:25] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti3005.esams.wmnet [08:14:39] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti3005.esams.wmnet [08:15:23] blocker task: T430839 [08:15:24] T430839: 1.47.0-wmf.20 deployment blockers - https://phabricator.wikimedia.org/T430839 [08:15:34] sry: T437982 [08:15:35] T437982: TypeError: MediaWiki\Rest\Module\ModuleManager::__construct(): Argument #4 ($jsonLocalizer) must be of type MediaWiki\Rest\JsonLocalizer, MediaWiki\Rest\ResponseFactory given, called in /srv/mediawiki/php-1.47.0-wmf.20/extensio - https://phabricator.wikimedia.org/T437982 [08:15:46] !log dpogorzelski@dns1004 START - running authdns-update [08:18:02] !log dpogorzelski@dns1004 END - running authdns-update [08:18:55] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie [08:19:04] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321136 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be1098.eqiad.wmnet wit... [08:20:03] (03CR) 10Brouberol: [C:03+1] liftwing-studio: chart deployment [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341174 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:20:23] (03CR) 10Dpogorzelski: [C:03+2] liftwing-studio: chart deployment [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341174 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:23:05] (03Merged) 10jenkins-bot: liftwing-studio: chart deployment [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341174 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:23:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:25:48] !log asw1-b3-magru - Disable logging and file logging for BRCM_PKT - T437984 [08:25:51] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:25:51] T437984: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984 [08:26:12] !log slyngshede@puppetserver1001 conftool action : set/pooled=no; selector: name=cp7010.magru.wmnet [08:26:17] !log mvernon@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1098.eqiad.wmnet with OS trixie [08:26:25] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321151 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be1098.eqiad.wmnet with OS... [08:26:38] (03CR) 10Slyngshede: [C:03+2] site.pp move cp7010 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1338843 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [08:27:05] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341692 [08:28:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:29:29] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321165 (10MatthewVernon) @VRiley-WMF I tried reimaging it myself, and can confirm that it's not currently able to UEFI HTTP boot int... [08:30:12] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321168 (10MatthewVernon) [08:32:45] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 10Ceph, and 2 others: Q1:rack/setup/install apus-be100[7-9] - https://phabricator.wikimedia.org/T436181#12321197 (10MatthewVernon) [08:33:08] !log slyngshede@cumin1003 START - Cookbook sre.hosts.reimage for host cp7010.magru.wmnet with OS trixie [08:33:42] 10ops-codfw, 06SRE, 10SRE-swift-storage, 10Ceph, and 2 others: Q1:rack/setup/install apus-be200[7-9] - https://phabricator.wikimedia.org/T436180#12321200 (10MatthewVernon) [08:34:57] (03CR) 10Ryan Kemper: kafka: converge topic config from hieradata (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [08:36:40] (03CR) 10Elukey: [C:03+2] docker_registry::instance: proper restart when the config changes [puppet] - 10https://gerrit.wikimedia.org/r/1341249 (https://phabricator.wikimedia.org/T428555) (owner: 10Elukey) [08:37:19] PROBLEM - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [08:42:50] (03PS1) 10Jaime Nuche: RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341698 (https://phabricator.wikimedia.org/T437982) [08:44:34] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jnuche@deploy1003 using scap backport" [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341698 (https://phabricator.wikimedia.org/T437982) (owner: 10Jaime Nuche) [08:47:05] (03Merged) 10jenkins-bot: RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341698 (https://phabricator.wikimedia.org/T437982) (owner: 10Jaime Nuche) [08:47:21] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12321267 (10Marostegui) @Jclark-ctr @VRiley-WMF this host is down again. Also the mgmt isn't accessible so I cannot check logs there either. [08:47:33] !log jnuche@deploy1003 Started scap sync-world: Backport for [[gerrit:1341698|RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] [08:47:36] T437982: TypeError: MediaWiki\Rest\Module\ModuleManager::__construct(): Argument #4 ($jsonLocalizer) must be of type MediaWiki\Rest\JsonLocalizer, MediaWiki\Rest\ResponseFactory given, called in /srv/mediawiki/php-1.47.0-wmf.20/extensio - https://phabricator.wikimedia.org/T437982 [08:49:43] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync [08:49:46] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync [08:52:18] !log jnuche@deploy1003 jnuche: Backport for [[gerrit:1341698|RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [08:53:02] !log jnuche@deploy1003 jnuche: Continuing with deployment [08:55:46] !log jmm@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install3004.wikimedia.org with reason: switch reboot [08:55:52] !log marostegui@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2197.codfw.wmnet with reason: cloning db2201 [08:57:14] (03PS1) 10PipelineBot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341702 [08:58:21] 06SRE, 10Data-Persistence-Backup, 10database-backups: Put db2201 back into backup production as a backup source - https://phabricator.wikimedia.org/T437411#12321320 (10Marostegui) Cloning x1 from db2197 [08:59:23] RESOLVED: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [08:59:36] !log jnuche@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341698|RestSandbox: Pass JsonLocalizer instead of ResponseFactory to ModuleManager (T437982)]] (duration: 12m 03s) [08:59:39] T437982: TypeError: MediaWiki\Rest\Module\ModuleManager::__construct(): Argument #4 ($jsonLocalizer) must be of type MediaWiki\Rest\JsonLocalizer, MediaWiki\Rest\ResponseFactory given, called in /srv/mediawiki/php-1.47.0-wmf.20/extensio - https://phabricator.wikimedia.org/T437982 [08:59:44] (03PS1) 10Marostegui: installserver: db1277 do not format /srv [puppet] - 10https://gerrit.wikimedia.org/r/1341703 [09:00:16] !log slyngshede@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp7010.magru.wmnet with reason: host reimage [09:00:58] !log ayounsi@cumin1003 START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switches reboot, T437984] [09:01:01] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switches reboot, T437984] [09:01:01] T437984: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984 [09:01:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.62% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:04:00] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7010.magru.wmnet with reason: host reimage [09:04:26] !log ayounsi@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance [09:04:26] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341706 (https://phabricator.wikimedia.org/T430839) [09:04:29] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341706 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [09:05:32] !log ayounsi@cumin1003 DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on asw1-by27-esams IPv6,asw1-by27-esams.mgmt,asw1-by-27-esams with reason: Switch maintenance [09:05:44] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341706 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [09:06:06] !log ayounsi@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-by27-esams,asw1-by27-esams IPv6,asw1-by27-esams.mgmt with reason: Switch maintenance [09:06:17] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [09:06:49] (03PS1) 10Elukey: sre.hosts.provision: set HttpDev1TlsMode for iDRAC 9 [cookbooks] - 10https://gerrit.wikimedia.org/r/1341708 (https://phabricator.wikimedia.org/T424895) [09:08:19] (03PS1) 10Muehlenhoff: Disable Homer on cumin1003 [puppet] - 10https://gerrit.wikimedia.org/r/1341710 (https://phabricator.wikimedia.org/T427897) [09:10:11] !log elukey@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [09:11:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.62% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:12:21] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1341330 (https://phabricator.wikimedia.org/T437937) (owner: 10Bking) [09:14:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.11% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:15:15] !log jnuche@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs T430839 [09:15:19] T430839: 1.47.0-wmf.20 deployment blockers - https://phabricator.wikimedia.org/T430839 [09:15:57] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:19:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.62% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:21:04] logs look good, train is done for today [09:21:52] (03CR) 10Ayounsi: [C:03+1] Disable Homer on cumin1003 [puppet] - 10https://gerrit.wikimedia.org/r/1341710 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [09:22:07] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1098.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [09:22:56] (03CR) 10Elukey: [C:03+2] Allow registryctl to use a custom Docker config.json filepath [docker-images/docker-report] - 10https://gerrit.wikimedia.org/r/1339621 (https://phabricator.wikimedia.org/T437297) (owner: 10Elukey) [09:23:53] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie [09:23:57] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:24:11] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 3 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321445 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be1098.eqiad.wmnet w... [09:24:15] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie [09:24:28] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 3 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321446 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be1098.eqiad.wmnet with... [09:24:38] !log ayounsi@cumin1003 START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BY27 [09:24:42] hm [09:25:55] arnaudb: I depooled esams for a network maintenance, is drmrs saturating? [09:26:05] it is a bit laggy according to the latency probes XioNoX [09:26:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:26:20] availability is going down also [09:26:57] !log ayounsi@cumin1003 END (FAIL) - Cookbook sre.network.depool-rack (exit_code=99) with action 'depool' for esams rack BY27 [09:27:01] but it looks like the avail issue is only on eqiad wdqs [09:27:10] s/only/mostly/ [09:27:44] arnaudb: let me know if I can proceed with my maintenance of if I should repool/postpone [09:28:06] is it long to run? [09:28:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:28:44] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7010.magru.wmnet with OS trixie [09:28:59] around 09:08 UTC drmrs started lagging/loosing availability [09:29:03] first part is a switch reboot, so should be quick [09:29:09] ack, proceed then! [09:30:12] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie [09:30:32] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 3 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321470 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by mvernon@cumin2003 for host ms-be1098.eqiad.wmnet w... [09:31:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.62% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:31:24] !log asw1-by27-esams> request system reboot - T437984 [09:31:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:31:28] T437984: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984 [09:34:41] PROBLEM - Router interfaces on mr1-esams is CRITICAL: CRITICAL: host 185.15.59.130, interfaces up: 34, down: 1, dormant: 0, excluded: 0, unused: 0: https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [09:35:57] FIRING: [2x] ProbeDown: Ripe Atlas anchor atlas3001:80 is not returning HTTP 200 OK on port 80 - https://wikitech.wikimedia.org/wiki/RIPE_Atlas#HTTP_checks_failing - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:36:18] (03CR) 10Muehlenhoff: [C:03+2] Disable Homer on cumin1003 [puppet] - 10https://gerrit.wikimedia.org/r/1341710 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [09:36:39] FIRING: [4x] CoreBGPDown: Core BGP session down between cr1-esams and asw1-by27-esams (185.15.59.155) - group Switch - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [09:36:40] arnaudb: switch is back up [09:36:41] RECOVERY - Router interfaces on mr1-esams is OK: OK: host 185.15.59.130, interfaces up: 35, down: 0, dormant: 0, excluded: 0, unused: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [09:37:19] RECOVERY - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [09:37:21] arnaudb: we can repool esams, but there is something worrisome if we can't depool esams without saturation in drmrs [09:38:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:39:34] XioNoX: agreed, it did not looked like a "full saturation" though [09:39:47] !log ayounsi@cumin1003 START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switches reboot, T437984] [09:39:50] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switches reboot, T437984] [09:39:51] T437984: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984 [09:39:54] XioNoX: It may help once I'm move a few of the cp-upload node to text. The first is being moved now in magru. We have seen this previous though, so might not be enough [09:39:57] arnaudb: repooling esams [09:40:31] (03PS1) 10Klausman: profiles/amd_gpu: Add GPUs ettings verification script [puppet] - 10https://gerrit.wikimedia.org/r/1341714 (https://phabricator.wikimedia.org/T431553) [09:40:57] FIRING: [6x] ProbeDown: Ripe Atlas anchor atlas3001:80 is not returning HTTP 200 OK on port 80 - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:41:15] drmrs went from 165MB/s to 580+, that's a nice uptick [09:41:39] RESOLVED: [4x] CoreBGPDown: Core BGP session down between cr1-esams and asw1-by27-esams (185.15.59.155) - group Switch - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [09:43:07] (03CR) 10Hnowlan: [C:03+1] "One nit but lgtm. At some point we should look at having multiple executor pools" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337312 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup-claude) [09:43:08] slyngs XioNoX fyi I manually resolved the p.age ↑ [09:43:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:43:57] RESOLVED: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:44:12] (03CR) 10Ryan Kemper: kafka: converge topic config from hieradata (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [09:45:21] (03PS1) 10Samwilson: ImageMagick: Move -background=none to be before -extent [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341716 [09:45:57] RESOLVED: [6x] ProbeDown: Ripe Atlas anchor atlas3001:80 is not returning HTTP 200 OK on port 80 - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:46:13] (03CR) 10Marostegui: [C:03+2] installserver: db1277 do not format /srv [puppet] - 10https://gerrit.wikimedia.org/r/1341703 (owner: 10Marostegui) [09:47:17] (03PS1) 10Daniel Kertesz: icinga: add Daniel Kertesz to authorized users [puppet] - 10https://gerrit.wikimedia.org/r/1341717 [09:48:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:48:39] (03PS4) 10Ryan Kemper: kafka: converge topic config from hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) [09:49:45] FIRING: WidespreadPuppetFailure: Puppet has failed in esams - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [09:52:34] 06SRE, 10Data-Persistence-Backup, 10database-backups: Put db2201 back into backup production as a backup source - https://phabricator.wikimedia.org/T437411#12321625 (10Marostegui) db2201:x1 is now up [09:53:49] (03CR) 10CI reject: [V:04-1] ImageMagick: Move -background=none to be before -extent [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341716 (owner: 10Samwilson) [09:55:33] (03PS1) 10David Caro: toolforge: add the logs-cli package [puppet] - 10https://gerrit.wikimedia.org/r/1341719 (https://phabricator.wikimedia.org/T432572) [09:57:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:57:28] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip for recommendation-api,sessionstore,tegola,termbox [puppet] - 10https://gerrit.wikimedia.org/r/1341131 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [09:57:45] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip for recommendation-api,sessionstore,tegola,termbox [puppet] - 10https://gerrit.wikimedia.org/r/1341132 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [09:58:15] (03PS6) 10Ladsgroup: Run the Swift result storage read off the event loop [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337312 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup-claude) [09:58:33] (03CR) 10Ladsgroup: Run the Swift result storage read off the event loop (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337312 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup-claude) [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1000) [10:00:27] 10ops-eqiad, 06DC-Ops: Unresponsive management for db1245.mgmt:22 - https://phabricator.wikimedia.org/T437998 (10phaultfinder) 03NEW [10:02:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [10:05:16] !log joal@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply [10:05:47] !log joal@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply [10:07:03] !log jelto@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply [10:07:42] 06SRE, 10Ganeti, 06Infrastructure-Foundations: Raise DRBD replication speed for Ganeti clusters - https://phabricator.wikimedia.org/T428878#12321728 (10MoritzMuehlenhoff) [10:07:45] 10ops-eqiad, 06DC-Ops: Unresponsive management for db1245.mgmt:22 - https://phabricator.wikimedia.org/T437998#12321729 (10Jclark-ctr) a:03Jclark-ctr [10:08:06] !log increased DRBD replication speed in Ganeti/esams T428878 [10:08:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:08:09] T428878: Raise DRBD replication speed for Ganeti clusters - https://phabricator.wikimedia.org/T428878 [10:09:15] !log jelto@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply [10:09:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in esams - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [10:09:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [10:10:43] !log hashar@deploy1003 Started deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies [10:10:57] !log hashar@deploy1003 Finished deploy [integration/docroot@5cf09c8]: build: Updating npm dependencies (duration: 00m 13s) [10:11:35] (03CR) 10Hnowlan: [C:03+1] Run the Swift result storage read off the event loop [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337312 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup-claude) [10:12:45] (03CR) 10Ladsgroup: [C:03+2] Run the Swift result storage read off the event loop [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337312 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup-claude) [10:13:09] (03PS1) 10Cathal Mooney: Re-enable transport path cr1-codfw to cr2-eqsin [homer/public] - 10https://gerrit.wikimedia.org/r/1341723 (https://phabricator.wikimedia.org/T435543) [10:15:18] (03Merged) 10jenkins-bot: Run the Swift result storage read off the event loop [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337312 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup-claude) [10:16:03] (03CR) 10Cathal Mooney: [C:03+2] Re-enable transport path cr1-codfw to cr2-eqsin [homer/public] - 10https://gerrit.wikimedia.org/r/1341723 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [10:18:40] (03Merged) 10jenkins-bot: Re-enable transport path cr1-codfw to cr2-eqsin [homer/public] - 10https://gerrit.wikimedia.org/r/1341723 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [10:20:40] !log increased DRBD replication speed in Ganeti/magru T428878 [10:20:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:20:44] T428878: Raise DRBD replication speed for Ganeti clusters - https://phabricator.wikimedia.org/T428878 [10:21:59] !log failover Ganeti master in magru to ganeti7001 [10:22:00] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:22:36] 06SRE, 10Ganeti, 06Infrastructure-Foundations: Raise DRBD replication speed for Ganeti clusters - https://phabricator.wikimedia.org/T428878#12321817 (10MoritzMuehlenhoff) [10:23:31] mvernon@cumin2003 reimage (PID 2182066) is awaiting input [10:24:27] PROBLEM - ganeti-wconfd running on ganeti7004 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 110 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [10:27:01] 06SRE, 10Wikimedia-Mailing-lists: create a new mailing list for ar-arbcom - https://phabricator.wikimedia.org/T437832#12321852 (10Ladsgroup) 05Open→03Resolved a:03Ladsgroup https://lists.wikimedia.org/postorius/lists/wikipedia-ar-arbcom.lists.wikimedia.org/members/owner/ [10:30:19] (03CR) 10MVernon: [C:03+1] "LGTM, and thanks for the help with this !" [cookbooks] - 10https://gerrit.wikimedia.org/r/1341708 (https://phabricator.wikimedia.org/T424895) (owner: 10Elukey) [10:30:56] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS trixie [10:31:09] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 3 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321861 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by mvernon@cumin2003 for host ms-be1098.eqiad.wmnet with... [10:31:31] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1341234 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [10:33:48] !log jelto@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply [10:34:15] !log jelto@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply [10:34:24] (03PS1) 10Elukey: Release version 0.0.20 [docker-images/docker-report] - 10https://gerrit.wikimedia.org/r/1341832 (https://phabricator.wikimedia.org/T437297) [10:35:14] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 3 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12321872 (10MatthewVernon) @VRiley-WMF with the help of @elukey I've made some progress here. ms-be1098 now boots into the installer O... [10:35:43] (03PS1) 10Btullis: sre.ceph.rotate-osd-keys: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1341834 [10:38:22] (03PS2) 10Samwilson: ImageMagick: Move -background=none to be before -extent [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341716 (https://phabricator.wikimedia.org/T438000) [10:39:04] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 10Ceph, and 2 others: Q1:rack/setup/install apus-be100[7-9] - https://phabricator.wikimedia.org/T436181#12321886 (10MatthewVernon) [10:39:23] 10ops-codfw, 06SRE, 10SRE-swift-storage, 10Ceph, and 2 others: Q1:rack/setup/install apus-be200[7-9] - https://phabricator.wikimedia.org/T436180#12321888 (10MatthewVernon) [10:39:35] (03PS2) 10Btullis: sre.ceph.rotate-osd-keys: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1341834 [10:41:12] (03CR) 10CI reject: [V:04-1] ImageMagick: Move -background=none to be before -extent [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341716 (https://phabricator.wikimedia.org/T438000) (owner: 10Samwilson) [10:43:00] (03Abandoned) 10Hashar: tests: Remove newline after header from MassMessageJob [extensions/MassMessage] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337961 (owner: 10Hashar) [10:48:43] (03CR) 10JavierMonton: [C:03+2] stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333854 (https://phabricator.wikimedia.org/T431555) (owner: 10JavierMonton) [10:51:19] (03Merged) 10jenkins-bot: stream: pageview-trending-relative [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333854 (https://phabricator.wikimedia.org/T431555) (owner: 10JavierMonton) [10:51:46] !log slyngshede@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp7010.magru.wmnet [10:52:42] (03PS1) 10Muehlenhoff: thumbor-plugins: Rebuild against latest package versions in Trixie [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341835 [10:54:09] (03PS3) 10Btullis: sre.ceph.rotate-osd-keys: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1341834 (https://phabricator.wikimedia.org/T428445) [10:54:43] (03PS1) 10Federico Ceratto: new_storage.pp: Monitor versitygw TLS certs [puppet] - 10https://gerrit.wikimedia.org/r/1341836 [10:57:24] 06SRE, 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Requesting access to analytics-privatedata-users level 3 for derenrich - https://phabricator.wikimedia.org/T437668#12322034 (10hnowlan) Pinging @HSwan-WMF for manager approval [10:58:13] (03CR) 10Hnowlan: [C:03+2] admin: add Lamar Petty and add them to releasers-mobile [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) (owner: 10Dzahn) [10:59:54] (03CR) 10Hnowlan: [C:03+2] admin: add cooltey to deployment [puppet] - 10https://gerrit.wikimedia.org/r/1339906 (https://phabricator.wikimedia.org/T437658) (owner: 10Ssingh) [11:00:04] (03CR) 10Hnowlan: [C:03+2] "Approval is in!" [puppet] - 10https://gerrit.wikimedia.org/r/1339906 (https://phabricator.wikimedia.org/T437658) (owner: 10Ssingh) [11:01:20] (03PS2) 10Federico Ceratto: new_storage.pp: Monitor versitygw TLS certs [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) [11:04:01] (03PS2) 10Hnowlan: admin: add wrai to deployment [puppet] - 10https://gerrit.wikimedia.org/r/1339866 (https://phabricator.wikimedia.org/T437652) (owner: 10Ssingh) [11:04:49] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart P{lvs7001.magru.wmnet} and A:liberica [11:05:04] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.admin depooling P{lvs7001.magru.wmnet} and A:liberica [11:05:15] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs7001.magru.wmnet} and A:liberica [11:05:28] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.admin pooling P{lvs7001.magru.wmnet} and A:liberica [11:05:30] (03PS3) 10Samwilson: ImageMagick: Move -background=none to be before -extent [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341716 (https://phabricator.wikimedia.org/T438000) [11:05:50] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs7001.magru.wmnet} and A:liberica [11:05:52] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P{lvs7001.magru.wmnet} and A:liberica [11:06:52] (03CR) 10Hnowlan: [C:03+2] "Group approval is in!" [puppet] - 10https://gerrit.wikimedia.org/r/1339866 (https://phabricator.wikimedia.org/T437652) (owner: 10Ssingh) [11:07:47] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12322086 (10hnowlan) 05Open→03Resolved a:03hnowlan Access granted to `deployment` [11:10:00] (03CR) 10CWilliams: new_storage.pp: Monitor versitygw TLS certs (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:11:12] (03PS2) 10LWatson: Enable ReaderExperiments in eswiki, jawiki, and ptwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339741 (https://phabricator.wikimedia.org/T438009) [11:11:17] 06SRE, 10LDAP-Access-Requests: Grant Access to wmf, for WMF staff/contractors nda group for gsduser - https://phabricator.wikimedia.org/T435852#12322140 (10hnowlan) 05Open→03Stalled [11:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:16:55] (03CR) 10Ladsgroup: [C:03+2] ImageMagick: Move -background=none to be before -extent [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341716 (https://phabricator.wikimedia.org/T438000) (owner: 10Samwilson) [11:17:37] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti7004.magru.wmnet [11:18:44] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti7004.magru.wmnet [11:19:56] (03Merged) 10jenkins-bot: ImageMagick: Move -background=none to be before -extent [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341716 (https://phabricator.wikimedia.org/T438000) (owner: 10Samwilson) [11:20:27] (03CR) 10Federico Ceratto: new_storage.pp: Monitor versitygw TLS certs (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:22:04] (03CR) 10Ladsgroup: "Now that I71f9875d5921a is merged (and I'm about to deploy it), is this still needed?" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341835 (owner: 10Muehlenhoff) [11:24:19] (03CR) 10Ayounsi: [C:03+1] sre.hosts.provision: set HttpDev1TlsMode for iDRAC 9 [cookbooks] - 10https://gerrit.wikimedia.org/r/1341708 (https://phabricator.wikimedia.org/T424895) (owner: 10Elukey) [11:31:37] (03CR) 10Muehlenhoff: "I just grabbed wikimedia/operations-software-thumbor-plugins:2026-09-15-112014-production from the registry and all changes are in. Thumbo" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341835 (owner: 10Muehlenhoff) [11:33:40] !log marostegui@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: cloning db2201 [11:33:57] (03CR) 10Finchgold: [C:03+1] ImageMagick: Move -background=none to be before -extent [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341716 (https://phabricator.wikimedia.org/T438000) (owner: 10Samwilson) [11:35:12] (03PS1) 10Ladsgroup: thumbor: Update to the most recent build [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341842 (https://phabricator.wikimedia.org/T438000) [11:36:57] (03CR) 10Ladsgroup: "I'm about to deploy it (if all goes well) Ia8f8a052346c25" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341835 (owner: 10Muehlenhoff) [11:38:43] (03CR) 10Marostegui: [C:03+1] new_storage.pp: Monitor versitygw TLS certs [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:40:29] (03CR) 10Muehlenhoff: "Perfect, can you please simply abandon this patch when it's deployed to production?" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341835 (owner: 10Muehlenhoff) [11:43:12] (03PS1) 10PipelineBot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341844 [11:45:07] (03PS1) 10Ayounsi: move BGPalerter to netmon hosts [puppet] - 10https://gerrit.wikimedia.org/r/1341845 (https://phabricator.wikimedia.org/T437024) [11:48:55] (03CR) 10Ayounsi: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1341845 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [11:49:32] (03CR) 10Muehlenhoff: [C:03+1] "Nice catch!" [puppet] - 10https://gerrit.wikimedia.org/r/1341514 (owner: 10JHathaway) [11:49:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [11:50:02] (03PS2) 10Ayounsi: move BGPalerter to netmon hosts [puppet] - 10https://gerrit.wikimedia.org/r/1341845 (https://phabricator.wikimedia.org/T437024) [11:51:30] (03CR) 10Ladsgroup: [C:03+2] thumbor: Update to the most recent build [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341842 (https://phabricator.wikimedia.org/T438000) (owner: 10Ladsgroup) [11:51:46] (03CR) 10Ladsgroup: "willdo." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341835 (owner: 10Muehlenhoff) [11:52:30] (03CR) 10JMeybohm: [C:03+1] "I'm not 100% sure the chain of envoys will detect/forward the gRPC chain all the way to the app. But it's probably easier to test/verify t" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [11:54:00] (03Merged) 10jenkins-bot: thumbor: Update to the most recent build [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341842 (https://phabricator.wikimedia.org/T438000) (owner: 10Ladsgroup) [11:54:03] (03CR) 10Elukey: [C:03+2] sre.hosts.provision: set HttpDev1TlsMode for iDRAC 9 [cookbooks] - 10https://gerrit.wikimedia.org/r/1341708 (https://phabricator.wikimedia.org/T424895) (owner: 10Elukey) [11:54:24] (03CR) 10Elukey: [C:03+2] Release version 0.0.20 [docker-images/docker-report] - 10https://gerrit.wikimedia.org/r/1341832 (https://phabricator.wikimedia.org/T437297) (owner: 10Elukey) [11:54:28] (03PS3) 10Ayounsi: move BGPalerter to netmon hosts [puppet] - 10https://gerrit.wikimedia.org/r/1341845 (https://phabricator.wikimedia.org/T437024) [11:54:46] (03CR) 10Ayounsi: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1341845 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [11:56:17] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [11:56:25] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [11:57:41] (03CR) 10Gergő Tisza: CommonSettings: Use a restrictive CSP for auth.wikimedia.org (032 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [11:58:54] (03PS1) 10PipelineBot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341848 [11:59:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1200) [12:00:42] (03CR) 10Ayounsi: "RPKI hosts might need a cleanup after deploy." [puppet] - 10https://gerrit.wikimedia.org/r/1341845 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [12:00:50] (03PS2) 10Arendpieter: CommonSettings: Use a restrictive CSP for auth.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) [12:00:54] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.admin config_reloading P{lvs7001.magru.wmnet} and A:liberica [12:01:13] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P{lvs7001.magru.wmnet} and A:liberica [12:01:39] (03CR) 10Arendpieter: "Thanks! Trimmed the comment in PS2; answered the question inline." (032 comments) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [12:01:44] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [12:02:20] (03CR) 10Dbrant: [C:03+2] wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341848 (owner: 10PipelineBot) [12:03:25] !log push pfw policies - T437627 [12:03:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:03:50] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [12:04:57] (03Merged) 10jenkins-bot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341848 (owner: 10PipelineBot) [12:06:01] !log dbrant@deploy1003 helmfile [staging] START helmfile.d/services/wikifeeds: apply [12:06:22] (03CR) 10Marostegui: new_storage.pp: Monitor versitygw TLS certs [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [12:06:27] (03PS2) 10Federico Ceratto: backups-disk-space.yaml: backups disk space [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) [12:06:49] !log dbrant@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifeeds: apply [12:07:07] !log dbrant@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifeeds: apply [12:07:38] !log dbrant@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply [12:07:47] !log dbrant@deploy1003 helmfile [codfw] START helmfile.d/services/wikifeeds: apply [12:08:03] (03CR) 10Volans: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1341298 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [12:08:12] (03CR) 10CI reject: [V:04-1] backups-disk-space.yaml: backups disk space [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [12:08:15] !log dbrant@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply [12:09:43] !log jmm@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on install7002.wikimedia.org with reason: switch reboot [12:10:13] (03Abandoned) 10Dbrant: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341702 (owner: 10PipelineBot) [12:10:20] (03Abandoned) 10Dbrant: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341331 (owner: 10PipelineBot) [12:11:17] (03CR) 10Majavah: [C:03+2] O:wmcs: Add new cloudinfra_etcd role [puppet] - 10https://gerrit.wikimedia.org/r/1341298 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [12:11:19] !log ayounsi@cumin1003 START - Cookbook sre.dns.admin DNS admin: depool magru [reason: switch reboot, T437984] [12:11:22] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: switch reboot, T437984] [12:11:22] T437984: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984 [12:11:51] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [12:12:36] !log ayounsi@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b4-magru,asw1-b4-magru IPv6,asw1-b4-magru.mgmt with reason: Switch maintenance [12:13:19] !log ayounsi@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 12 hosts with reason: Switch maintenance [12:13:58] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [12:17:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:18:17] (03PS1) 10Arnaudb: deployment_server/k8s: set kubeconfig files for aphlict [puppet] - 10https://gerrit.wikimedia.org/r/1341711 (https://phabricator.wikimedia.org/T436657) [12:19:05] !log slyngshede@puppetserver1001 conftool action : set/weight=1; selector: name=cp7010.magru.wmnet [12:19:36] (03CR) 10Federico Ceratto: [C:03+2] new_storage.pp: Monitor versitygw TLS certs [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [12:22:11] !log installing shadow security updates [12:22:12] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:22:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:23:27] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart P{lvs7001.magru.wmnet} and A:liberica [12:23:40] (03PS1) 10STran: SuggestedInvestigations: Update "sockpuppet" queue view defaults [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341861 (https://phabricator.wikimedia.org/T438018) [12:23:43] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.admin depooling P{lvs7001.magru.wmnet} and A:liberica [12:23:54] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs7001.magru.wmnet} and A:liberica [12:24:07] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.admin pooling P{lvs7001.magru.wmnet} and A:liberica [12:24:30] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs7001.magru.wmnet} and A:liberica [12:24:32] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P{lvs7001.magru.wmnet} and A:liberica [12:26:20] (03Abandoned) 10Ladsgroup: thumbor-plugins: Rebuild against latest package versions in Trixie [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1341835 (owner: 10Muehlenhoff) [12:27:12] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 15 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341861 (https://phabricator.wikimedia.org/T438018) (owner: 10STran) [12:29:28] !log asw1-b4-magru> request system reboot - T437984 [12:29:30] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:29:31] T437984: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984 [12:30:29] (03PS1) 10KartikMistry: Update Apertium to 2026-09-15-084320-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341866 (https://phabricator.wikimedia.org/T437213) [12:30:44] (03CR) 10Marostegui: "@fceratto@wikimedia.org please note that this change has been +2 but not submitted/merged because it depends on another one from you too." [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [12:31:26] (03CR) 10Muehlenhoff: [C:03+1] "Looks good! No need for specific cleanups on rpki* IMO, now that bgpalerter is out of the way, we can reimage them to trixie and clean the" [puppet] - 10https://gerrit.wikimedia.org/r/1341845 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [12:32:45] PROBLEM - Router interfaces on mr1-magru is CRITICAL: CRITICAL: host 195.200.68.132, interfaces up: 34, down: 1, dormant: 0, excluded: 0, unused: 0: https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [12:33:22] FIRING: ProbeDown: Ripe Atlas anchor atlas7001:80 is not returning HTTP 200 OK on port 80 - https://wikitech.wikimedia.org/wiki/RIPE_Atlas#HTTP_checks_failing - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:34:39] FIRING: [4x] CoreBGPDown: Core BGP session down between cr1-magru and asw1-b4-magru (195.200.68.149) - group Switch - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:34:45] RECOVERY - Router interfaces on mr1-magru is OK: OK: host 195.200.68.132, interfaces up: 35, down: 0, dormant: 0, excluded: 0, unused: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [12:34:53] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-magru:et-0/0/2 (Core: asw1-b4-magru:et-0/0/48) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:35:07] (03PS1) 10Mvolz: zotero: update to latest [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341867 (https://phabricator.wikimedia.org/T435179) [12:36:21] switch is back up, we short start seeing recoveries [12:38:22] RESOLVED: [5x] ProbeDown: Ripe Atlas anchor atlas7001:80 is not returning HTTP 200 OK on port 80 - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:39:03] (03CR) 10Marostegui: [C:03+1] "PCC failed but checking its results it looks unrelated to this change. Also checking some random results they look fine." [puppet] - 10https://gerrit.wikimedia.org/r/1341257 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [12:39:39] RESOLVED: [4x] CoreBGPDown: Core BGP session down between cr1-magru and asw1-b4-magru (195.200.68.149) - group Switch - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:39:53] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-magru:et-0/0/2 (Core: asw1-b4-magru:et-0/0/48) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:39:59] jouncebot: nowandnext [12:39:59] For the next 0 hour(s) and 20 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1200) [12:39:59] In 0 hour(s) and 20 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1300) [12:40:20] (03PS5) 10Anzx: lift IP cap for edit-a-thon /workshop [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) [12:40:53] !log ayounsi@cumin1003 START - Cookbook sre.dns.admin DNS admin: pool magru [reason: network maintenance finished, T437984] [12:40:56] T437984: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984 [12:41:19] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 15 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [12:41:26] (03CR) 10Brouberol: [C:03+1] "Nice work" [cookbooks] - 10https://gerrit.wikimedia.org/r/1341834 (https://phabricator.wikimedia.org/T428445) (owner: 10Btullis) [12:42:35] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: network maintenance finished, T437984] [12:42:57] (03CR) 10Marostegui: [C:04-1] "Actually, I am not sure this covers core_test role. Let me check a bit more" [puppet] - 10https://gerrit.wikimedia.org/r/1341257 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [12:43:39] (03PS4) 10Btullis: sre.ceph.rotate-osd-keys: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1341834 (https://phabricator.wikimedia.org/T428445) [12:44:19] (03CR) 10Arnaudb: [C:03+2] "I'll apply that change on the replicas first, with puppet disabled on the primary" [puppet] - 10https://gerrit.wikimedia.org/r/1341184 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [12:45:29] !log btullis@cumin1004 START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P{cephosd2001.codfw.wmnet} and (A:cephosd-codfw or A:cephosd-eqiad) [12:46:41] !log btullis@cumin1004 END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P{cephosd2001.codfw.wmnet} and (A:cephosd-codfw or A:cephosd-eqiad) [12:50:02] (03CR) 10Giuseppe Lavagetto: [C:03+1] k8s: Add a script to sync kubelet labels with the API server (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341181 (owner: 10JMeybohm) [12:52:01] (03CR) 10Marostegui: [C:04-1] "Can you run the PCC against test hosts to see if the change actually fixes anything?" [puppet] - 10https://gerrit.wikimedia.org/r/1341257 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [12:57:09] (03PS1) 10CWilliams: admin: Updated aliases for cwilliams [puppet] - 10https://gerrit.wikimedia.org/r/1341871 [12:57:35] (03PS5) 10Btullis: sre.ceph.rotate-osd-keys: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1341834 (https://phabricator.wikimedia.org/T428445) [12:57:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [12:58:42] (03PS1) 10Giuseppe Lavagetto: gvisor: install in wikikube-codfw [puppet] - 10https://gerrit.wikimedia.org/r/1341873 (https://phabricator.wikimedia.org/T435796) [12:58:45] (03PS1) 10Giuseppe Lavagetto: gvisor: enable in wikikube [puppet] - 10https://gerrit.wikimedia.org/r/1341874 (https://phabricator.wikimedia.org/T435796) [12:59:44] !log btullis@cumin1004 START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P{cephosd2001.codfw.wmnet} and (A:cephosd-codfw or A:cephosd-eqiad) [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: I seem to be stuck in Groundhog week. Sigh. Time for (yet another) UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1300). [13:00:05] stephanebisson, Tran, and anzx: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:08] o/ [13:00:10] o/ [13:00:10] o/ [13:00:55] !log btullis@cumin1004 END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P{cephosd2001.codfw.wmnet} and (A:cephosd-codfw or A:cephosd-eqiad) [13:03:48] !log btullis@cumin1004 START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on P{cephosd2001.codfw.wmnet} and (A:cephosd-codfw or A:cephosd-eqiad) [13:04:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:04:54] (03CR) 10Giuseppe Lavagetto: [C:03+1] icinga: add Daniel Kertesz to authorized users [puppet] - 10https://gerrit.wikimedia.org/r/1341717 (owner: 10Daniel Kertesz) [13:05:15] (03CR) 10Slyngshede: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1341717 (owner: 10Daniel Kertesz) [13:06:15] is everyone self-deploying or do we need a deployer? [13:06:16] (03CR) 10Hnowlan: [C:03+2] "Approval is in, key verified oob." [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) (owner: 10Dzahn) [13:06:39] I need someone to deploy for me [13:07:00] Do we know if the deployment blocker (expired image) is resolved? [13:07:27] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.upgrade restart P{lvs7003.magru.wmnet} and A:liberica [13:07:31] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.admin depooling P{lvs7003.magru.wmnet} and A:liberica [13:07:42] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs7003.magru.wmnet} and A:liberica [13:07:44] (03CR) 10Daniel Kertesz: [C:03+2] icinga: add Daniel Kertesz to authorized users [puppet] - 10https://gerrit.wikimedia.org/r/1341717 (owner: 10Daniel Kertesz) [13:07:55] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.admin pooling P{lvs7003.magru.wmnet} and A:liberica [13:08:05] Anyway, I'll try deploying my patch [13:08:16] was that more recent than the train blocker earlier today? Because that's been resolved afaik [13:08:17] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs7003.magru.wmnet} and A:liberica [13:08:19] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P{lvs7003.magru.wmnet} and A:liberica [13:08:20] (03PS3) 10Federico Ceratto: new_storage.pp: Monitor versitygw TLS certs [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) [13:08:30] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbisson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334946 (https://phabricator.wikimedia.org/T434487) (owner: 10Sbisson) [13:09:54] (03Merged) 10jenkins-bot: ArticleGuidance: Remove the experiment configuration keys [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334946 (https://phabricator.wikimedia.org/T434487) (owner: 10Sbisson) [13:10:13] !log slyngshede@cumin1003 START - Cookbook sre.loadbalancer.admin config_reloading P{lvs7003.magru.wmnet} and A:liberica [13:10:20] !log sbisson@deploy1003 Started scap sync-world: Backport for [[gerrit:1334946|ArticleGuidance: Remove the experiment configuration keys (T434487)]] [13:10:23] T434487: Stop AG experiment [tr, en.simple, fr] and switch to enabled by default for junior editors in tr and en.simple - https://phabricator.wikimedia.org/T434487 [13:10:31] (03CR) 10Gergő Tisza: [C:03+2] "It was rolled back via scap / spiderpig (which I guess undoes the last commit locally and then deploys that?). I probably should have done" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341276 (owner: 10Gergő Tisza) [13:10:32] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P{lvs7003.magru.wmnet} and A:liberica [13:10:57] (03CR) 10Federico Ceratto: "I've noticed, I'm rebasing it to merge it." [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [13:11:00] (03CR) 10Federico Ceratto: [C:03+2] new_storage.pp: Monitor versitygw TLS certs [puppet] - 10https://gerrit.wikimedia.org/r/1341836 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [13:11:31] (03PS1) 10Jelto: etherpad: update image tag to newest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341877 (https://phabricator.wikimedia.org/T435509) [13:14:38] !log sbisson@deploy1003 sbisson: Backport for [[gerrit:1334946|ArticleGuidance: Remove the experiment configuration keys (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:15:09] !log sbisson@deploy1003 sbisson: Continuing with deployment [13:15:11] (03CR) 10Jelto: [C:03+2] etherpad: update image tag to newest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341877 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [13:16:59] !log btullis@cumin1004 END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on P{cephosd2001.codfw.wmnet} and (A:cephosd-codfw or A:cephosd-eqiad) [13:18:00] (03Merged) 10jenkins-bot: etherpad: update image tag to newest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341877 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [13:19:39] !log sbisson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1334946|ArticleGuidance: Remove the experiment configuration keys (T434487)]] (duration: 09m 19s) [13:19:43] T434487: Stop AG experiment [tr, en.simple, fr] and switch to enabled by default for junior editors in tr and en.simple - https://phabricator.wikimedia.org/T434487 [13:20:08] (03CR) 10Gergő Tisza: [C:03+1] CommonSettings: Use a restrictive CSP for auth.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [13:20:35] !log jelto@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/etherpad: apply [13:20:56] o/ [13:21:03] !log jelto@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/etherpad: apply [13:21:04] (03CR) 10Dreamy Jazz: [C:03+1] SuggestedInvestigations: Update "sockpuppet" queue view defaults [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341861 (https://phabricator.wikimedia.org/T438018) (owner: 10STran) [13:21:15] I could deploy if needed (was in a meeting earlier) [13:22:08] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 15 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [13:22:39] I think stephanebisson just finished? I can deploy my own but everything left in the window is a config; should they all go at the same time? [13:23:49] (03CR) 10Btullis: [C:03+2] sre.ceph.rotate-osd-keys: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1341834 (https://phabricator.wikimedia.org/T428445) (owner: 10Btullis) [13:23:55] I'll deploy one more patch at the end (config but risky) [13:24:26] in absence of other feedback, ig I'll just deploy my own for now [13:24:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341861 (https://phabricator.wikimedia.org/T438018) (owner: 10STran) [13:25:48] (03Merged) 10jenkins-bot: SuggestedInvestigations: Update "sockpuppet" queue view defaults [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341861 (https://phabricator.wikimedia.org/T438018) (owner: 10STran) [13:26:02] (03CR) 10Muehlenhoff: [C:03+2] Disable bullseye-security in apt config for Bullseye [puppet] - 10https://gerrit.wikimedia.org/r/1340987 (https://phabricator.wikimedia.org/T437069) (owner: 10Muehlenhoff) [13:26:07] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1341861|SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] [13:26:10] T438018: Uncheck overlapping editing from Likely Socks queue - https://phabricator.wikimedia.org/T438018 [13:26:28] (03PS1) 10Blake: golang: add trixie-based golang-1.26 image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341880 (https://phabricator.wikimedia.org/T423851) [13:27:00] (03Merged) 10jenkins-bot: sre.ceph.rotate-osd-keys: add cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1341834 (https://phabricator.wikimedia.org/T428445) (owner: 10Btullis) [13:27:44] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341883 [13:28:22] (03PS3) 10Giuseppe Lavagetto: shellbox-timeline: enable gVisor everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341600 (https://phabricator.wikimedia.org/T436649) [13:28:22] (03PS1) 10Giuseppe Lavagetto: wikikube: enable gVisor [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341884 (https://phabricator.wikimedia.org/T436649) [13:28:43] (03CR) 10Muehlenhoff: [C:03+1] "Looks good, all approvals are in" [puppet] - 10https://gerrit.wikimedia.org/r/1339879 (https://phabricator.wikimedia.org/T437653) (owner: 10Ssingh) [13:30:23] !log stran@deploy1003 stran: Backport for [[gerrit:1341861|SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:31:51] lgtm, continuing [13:32:02] !log stran@deploy1003 stran: Continuing with deployment [13:36:30] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341861|SuggestedInvestigations: Update "sockpuppet" queue view defaults (T438018)]] (duration: 10m 23s) [13:36:33] T438018: Uncheck overlapping editing from Likely Socks queue - https://phabricator.wikimedia.org/T438018 [13:36:34] done [13:40:20] (03PS2) 10Krinkle: varnish: Refactor 19-normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) [13:42:02] (03CR) 10JMeybohm: [C:04-1] gvisor: install in wikikube-codfw (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341873 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [13:42:33] Lucas_WMDE: If you could deploy i have a patch for deployment [13:43:32] (03CR) 10JMeybohm: [C:03+1] wikikube: enable gVisor [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341884 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [13:43:48] (03CR) 10JMeybohm: [C:03+1] shellbox-timeline: enable gVisor everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341600 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [13:43:55] !log btullis@cumin1004 START - Cookbook sre.ceph.rotate-osd-keys rolling rotate_keys on A:cephosd-codfw [13:50:59] (03PS1) 10Muehlenhoff: Remove obsolete gpu-tester image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341889 (https://phabricator.wikimedia.org/T416452) [13:51:35] (03PS1) 10Krinkle: Restore table borders for client-side MathJax [extensions/Math] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341890 (https://phabricator.wikimedia.org/T435274) [13:53:33] (03CR) 10JMeybohm: [C:03+2] k8s: Add a script to sync kubelet labels with the API server [puppet] - 10https://gerrit.wikimedia.org/r/1341181 (owner: 10JMeybohm) [13:55:40] (03PS1) 10Muehlenhoff: Don't add the envoy-future repo component [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341892 [13:56:01] (03PS1) 10Jelto: trafficserver: add mapping for etherpad-next [puppet] - 10https://gerrit.wikimedia.org/r/1341893 (https://phabricator.wikimedia.org/T435509) [13:56:03] (03PS1) 10Jelto: cache-text: set caching for etherpad-next "websockets" [puppet] - 10https://gerrit.wikimedia.org/r/1341894 (https://phabricator.wikimedia.org/T435509) [13:56:28] (03CR) 10JHathaway: [C:03+2] cron_splay: allow splaying over a single host [puppet] - 10https://gerrit.wikimedia.org/r/1341514 (owner: 10JHathaway) [13:57:15] (03PS1) 10Dreamy Jazz: ReportIncidentController: Instance cache expensive methods [extensions/ReportIncident] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341895 (https://phabricator.wikimedia.org/T437588) [13:57:16] (03PS25) 10Herron: sre.opensearch.roll-restart-reboot: include checklist items [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) [13:58:00] jouncebot: nowandnext [13:58:00] For the next 0 hour(s) and 1 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1300) [13:58:00] In 0 hour(s) and 1 minute(s): Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1400) [13:58:17] (03CR) 10Hashar: zuul: refactor zookeeper inclusion to fix firewalling (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [13:59:14] (03PS2) 10Giuseppe Lavagetto: gvisor: install in wikikube-codfw [puppet] - 10https://gerrit.wikimedia.org/r/1341873 (https://phabricator.wikimedia.org/T435796) [13:59:14] (03PS2) 10Giuseppe Lavagetto: gvisor: enable in wikikube [puppet] - 10https://gerrit.wikimedia.org/r/1341874 (https://phabricator.wikimedia.org/T435796) [13:59:42] anzx: I'm looking at yours [13:59:52] (03CR) 10Krinkle: lift IP cap for edit-a-thon /workshop (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [13:59:52] sorry, I missed the ping [14:00:05] Deploy window Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1400) [14:00:13] (03CR) 10Giuseppe Lavagetto: gvisor: install in wikikube-codfw (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341873 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:00:56] I want to deploy a patch, does anzx have time to get the feedback on their patch addressed before I go? [14:01:02] (03CR) 10JMeybohm: [C:03+1] gvisor: install in wikikube-codfw [puppet] - 10https://gerrit.wikimedia.org/r/1341873 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:01:16] (03CR) 10Hashar: "The `ci-build-images` script is invoked once to create the image. There is no need to have the Puppet agent to run it afterward ;)" [puppet] - 10https://gerrit.wikimedia.org/r/1333796 (https://phabricator.wikimedia.org/T436775) (owner: 10Hashar) [14:01:20] (03CR) 10JMeybohm: [C:03+1] gvisor: enable in wikikube [puppet] - 10https://gerrit.wikimedia.org/r/1341874 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:01:30] (03CR) 10Anzx: lift IP cap for edit-a-thon /workshop (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [14:01:35] tgr_: Are you intending to deploy BTW? [14:01:42] (03PS3) 10Muehlenhoff: profile::dbbackups::transfer: Default db transfers to false [puppet] - 10https://gerrit.wikimedia.org/r/1341121 (https://phabricator.wikimedia.org/T427897) [14:02:21] Dreamy_Jazz: go ahead. [14:02:30] Thanks [14:02:51] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [extensions/ReportIncident] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341895 (https://phabricator.wikimedia.org/T437588) (owner: 10Dreamy Jazz) [14:02:55] 06SRE, 10SRE-Access-Requests: Requesting access to for - https://phabricator.wikimedia.org/T438034 (10Hany.elmokadem) 03NEW [14:03:41] 06SRE, 10SRE-Access-Requests: Requesting access to Superset Dashboard for Hany EL Mokadem - https://phabricator.wikimedia.org/T438034#12322991 (10Hany.elmokadem) [14:04:21] (03CR) 10Giuseppe Lavagetto: [C:03+2] gvisor: install in wikikube-codfw [puppet] - 10https://gerrit.wikimedia.org/r/1341873 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:04:44] (03Merged) 10jenkins-bot: ReportIncidentController: Instance cache expensive methods [extensions/ReportIncident] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341895 (https://phabricator.wikimedia.org/T437588) (owner: 10Dreamy Jazz) [14:05:10] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1341895|ReportIncidentController: Instance cache expensive methods (T437588)]] [14:05:14] T437588: Wikimedia\RequestTimeout\RequestTimeoutException: The maximum execution time of {limit} seconds was exceeded - https://phabricator.wikimedia.org/T437588 [14:05:47] (03CR) 10Krinkle: lift IP cap for edit-a-thon /workshop (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [14:06:41] 06SRE, 06Infrastructure-Foundations, 10netops: Add link from cloudsw1-e4-eqiad to cloudsw1-f4-eiqad - https://phabricator.wikimedia.org/T372061#12323044 (10cmooney) 05Open→03Declined Closing this one, we will review how to provision the bandwidth when the C8/D5 switches are upgraded. [14:06:45] (03CR) 10Arnaudb: [C:03+1] "looks good to me" [puppet] - 10https://gerrit.wikimedia.org/r/1341893 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [14:07:23] (03CR) 10Krinkle: lift IP cap for edit-a-thon /workshop (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [14:07:31] (03PS4) 10CWilliams: mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [14:07:56] Dreamy_Jazz: if no test kitchen deployment is happening then yes [14:07:59] 06SRE, 10SRE-Access-Requests: Requesting access to Superset Dashboard for Hany EL Mokadem - https://phabricator.wikimedia.org/T438034#12323059 (10AndrewTavis_WMDE) I'm not sure if I have approval rights for these sorts of requests, but as @Hany.elmokadem is an Engineering Manager for WMDE, this request should... [14:08:06] (03CR) 10CI reject: [V:04-1] mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [14:09:01] (03CR) 10Arnaudb: "small nit inline, otherwise lgtm!" [puppet] - 10https://gerrit.wikimedia.org/r/1341894 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [14:09:31] !log dreamyjazz@deploy1003 dreamyjazz: Backport for [[gerrit:1341895|ReportIncidentController: Instance cache expensive methods (T437588)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:10:34] (03PS1) 10Elukey: profile::k8s::deployment_server: add python3-docker-report [puppet] - 10https://gerrit.wikimedia.org/r/1341899 (https://phabricator.wikimedia.org/T437297) [14:10:45] (03CR) 10CWilliams: "I have rebased this MR and changed its content. The `HUP` can be seen working on the ticket with my update earlier today." [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [14:11:04] (03CR) 10Marostegui: [C:03+1] admin: Updated aliases for cwilliams [puppet] - 10https://gerrit.wikimedia.org/r/1341871 (owner: 10CWilliams) [14:11:46] Still testing [14:12:39] !log dreamyjazz@deploy1003 dreamyjazz: Continuing with deployment [14:12:43] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:12:56] (03PS5) 10CWilliams: mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [14:12:57] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:13:30] (03PS3) 10Hnowlan: admin: add access for cklimas (deployment, analytics-private-data; krb) [puppet] - 10https://gerrit.wikimedia.org/r/1339879 (https://phabricator.wikimedia.org/T437653) (owner: 10Ssingh) [14:15:37] (03CR) 10Elukey: [C:03+1] "A lot of memories :D" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341889 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [14:16:17] (03PS6) 10CWilliams: mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [14:16:48] !log btullis@cumin1004 END (PASS) - Cookbook sre.ceph.rotate-osd-keys (exit_code=0) rolling rotate_keys on A:cephosd-codfw [14:16:49] (03CR) 10CWilliams: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [14:17:02] (03CR) 10Hnowlan: [C:03+2] admin: add access for cklimas (deployment, analytics-private-data; krb) [puppet] - 10https://gerrit.wikimedia.org/r/1339879 (https://phabricator.wikimedia.org/T437653) (owner: 10Ssingh) [14:17:07] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341895|ReportIncidentController: Instance cache expensive methods (T437588)]] (duration: 11m 56s) [14:17:11] T437588: Wikimedia\RequestTimeout\RequestTimeoutException: The maximum execution time of {limit} seconds was exceeded - https://phabricator.wikimedia.org/T437588 [14:17:14] I'm done with scap [14:18:26] 06SRE, 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Requesting access to deployment and analytics-privatedata-users for Cklimas - https://phabricator.wikimedia.org/T437653#12323130 (10hnowlan) Access granted! [14:18:52] (03CR) 10CWilliams: mediabackups: Update versitygw systemd unit file to SIGHUP on reload (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [14:19:53] (03PS3) 10Krinkle: varnish: Refactor 19-normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) [14:20:26] (03CR) 10Krinkle: "Aye, the HTTP 400 error prevents the testreq appendix from running it seems. This isn't really the point of the test though, so I've chang" [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [14:21:10] (03PS1) 10Dpogorzelski: liftwing-studio: add external-services egress to the chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341903 (https://phabricator.wikimedia.org/T437706) [14:21:12] (03PS1) 10Dpogorzelski: liftwing-studio: OIDC auth against idp.wikimedia.org [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341904 (https://phabricator.wikimedia.org/T437706) [14:21:52] jouncebot: nowandnext [14:21:52] For the next 0 hour(s) and 8 minute(s): Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1400) [14:21:52] In 0 hour(s) and 8 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1430) [14:22:14] looks like the IP exemption patch is still being discussed? [14:22:22] 06SRE, 10vm-requests: eqiad: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438037 (10BTullis) 03NEW p:05Triage→03Medium [14:22:33] I'll deploy the other one then [14:23:16] 06SRE, 10vm-requests: codfw: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438038 (10BTullis) 03NEW [14:24:03] 06SRE, 10vm-requests: codfw: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438038#12323185 (10BTullis) p:05Low→03Medium [14:24:11] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tgr@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [14:26:07] 06SRE, 10vm-requests: eqiad: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438037#12323201 (10BTullis) I have updated https://wikitech.wikimedia.org/wiki/SRE/Infrastructure_naming_conventions#Hostname_prefixes and I plan to use `ceph-admin1001.eqiad.wmnet` for this host. [14:26:15] (03PS6) 10Anzx: lift IP cap for edit-a-thon /workshop [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) [14:26:39] 06SRE, 10vm-requests: codfw: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438038#12323203 (10BTullis) I have updated https://wikitech.wikimedia.org/wiki/SRE/Infrastructure_naming_conventions#Hostname_prefixes and I plan to use ceph-admin2001.eqiad.wmnet for this host. [14:27:03] (03CR) 10Anzx: lift IP cap for edit-a-thon /workshop (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [14:27:09] (03CR) 10CI reject: [V:04-1] CommonSettings: Use a restrictive CSP for auth.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [14:28:28] (03CR) 10Gergő Tisza: [C:03+1] "CI error is T352319 I think." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [14:29:02] (03CR) 10Gergő Tisza: [C:03+2] CommonSettings: Use a restrictive CSP for auth.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1430) [14:30:10] (03Merged) 10jenkins-bot: CommonSettings: Use a restrictive CSP for auth.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [14:30:15] (03CR) 10Marostegui: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [14:30:17] (03CR) 10Krinkle: lift IP cap for edit-a-thon /workshop (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [14:30:52] (03CR) 10Krinkle: [C:03+1] lift IP cap for edit-a-thon /workshop [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [14:32:09] (03PS1) 10Dpogorzelski: liftwing-studio: OIDC auth against idp.wikimedia.org [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341909 (https://phabricator.wikimedia.org/T437706) [14:33:50] (03CR) 10Klausman: [C:03+1] Remove obsolete gpu-tester image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341889 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [14:34:54] !log tgr@deploy1003 Started scap sync-world: Backport for [[gerrit:1341292|CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] [14:34:58] T419684: Add restrictive CSP to auth.wikimedia.org - https://phabricator.wikimedia.org/T419684 [14:35:08] (03CR) 10Krinkle: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [14:35:29] (03PS4) 10Krinkle: varnish: Refactor 19-normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) [14:35:32] (03CR) 10Krinkle: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [14:35:42] (03CR) 10Krinkle: "Passes for me now." [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [14:39:12] !log tgr@deploy1003 tgr, arendpieter: Backport for [[gerrit:1341292|CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:39:43] (03PS2) 10Dpogorzelski: liftwing-studio: OIDC auth against idp.wikimedia.org [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341904 (https://phabricator.wikimedia.org/T437706) [14:42:03] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply [14:42:38] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 15 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339741 (https://phabricator.wikimedia.org/T438009) (owner: 10LWatson) [14:42:47] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply [14:44:20] 06SRE, 10vm-requests: eqiad: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438037#12323344 (10BTullis) Corresponding VM request in codfw is: {T438038} [14:44:23] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply [14:44:41] 06SRE, 10vm-requests: codfw: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438038#12323350 (10BTullis) Corresponding VM request in eqiad is: {T438037} [14:44:58] (03CR) 10Hnowlan: [C:03+2] icinga: remove dead ElasticSearch checks breaking CI [puppet] - 10https://gerrit.wikimedia.org/r/1338157 (owner: 10Hnowlan) [14:45:02] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply [14:45:20] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply [14:45:42] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply [14:46:43] (03PS3) 10Hnowlan: graphite: remove module, references, config [puppet] - 10https://gerrit.wikimedia.org/r/1332767 (https://phabricator.wikimedia.org/T435340) [14:47:24] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply [14:47:28] (03CR) 10Dzahn: [C:03+1] deployment_server/k8s: set kubeconfig files for aphlict [puppet] - 10https://gerrit.wikimedia.org/r/1341711 (https://phabricator.wikimedia.org/T436657) (owner: 10Arnaudb) [14:47:58] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply [14:48:05] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply [14:48:37] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply [14:48:52] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply [14:48:59] 10ops-eqiad, 06SRE, 06DC-Ops, 10Kafka-Infrastructure, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12323368 (10cmooney) >>! In T435775#12312748, @VRiley-WMF wrote: > Row A currently only have a few cabinets that... [14:49:25] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply [14:52:36] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply [14:52:39] (03CR) 10Giuseppe Lavagetto: [C:03+2] gvisor: enable in wikikube [puppet] - 10https://gerrit.wikimedia.org/r/1341874 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:52:44] (03CR) 10Elukey: [C:03+2] admin_ng: move cfssl's issuer endpoint to PKI hosts [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341183 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [14:53:06] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply [14:53:50] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply [14:54:26] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply [14:54:35] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply [14:54:58] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply [14:54:59] (03CR) 10Gergő Tisza: [C:03+2] "Tested login, signup, top-level autologin, subrequest autologin, passkeys. All seem to work." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341292 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [14:55:06] !log tgr@deploy1003 tgr, arendpieter: Continuing with deployment [14:56:01] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12323401 (10Jdforrester-WMF) 05In progress→03Resolved Live at https://grafana.wikimedia.org/d/slot-pilot-slo-detail/sloth-... [14:59:19] (03CR) 10Elukey: [C:03+2] profile::pki::client: move the signer endpoint to pki[12]002 [puppet] - 10https://gerrit.wikimedia.org/r/1341234 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [15:00:04] jelto, arnoldokoth, mutante, and arnaudb: That opportune time for a SRE Collaboration Services office hours deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1500). [15:00:06] !log tgr@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341292|CommonSettings: Use a restrictive CSP for auth.wikimedia.org (T419684)]] (duration: 25m 11s) [15:00:10] T419684: Add restrictive CSP to auth.wikimedia.org - https://phabricator.wikimedia.org/T419684 [15:03:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.41% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:09:04] (03CR) 10RLazarus: [C:04-1] "Thanks for the find. But I think instead of removing the component, it needs s/bullseye/bookworm/ in envoy-future.list. We sometimes have " [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341892 (owner: 10Muehlenhoff) [15:09:21] (03CR) 10Herron: [C:03+1] Puppet 8: Replace unscoped legacy facts in module opensearch_dashboards [puppet] - 10https://gerrit.wikimedia.org/r/1338040 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:09:35] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module opensearch_dashboards [puppet] - 10https://gerrit.wikimedia.org/r/1338040 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:11:44] !log elukey@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'sync'. [15:11:49] !log elukey@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'sync'. [15:12:19] !log elukey@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'sync'. [15:12:22] !log elukey@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'. [15:13:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:15:22] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-presto1007 - https://phabricator.wikimedia.org/T434505#12323524 (10VRiley-WMF) 05Open→03Resolved Checked in today with this unit and the disk has recovered. [15:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:21:03] (03PS3) 10Federico Ceratto: backups-disk-space.yaml: backups disk space [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) [15:22:05] 06SRE, 10vm-requests: eqiad: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438037#12323553 (10MoritzMuehlenhoff) Looks good, please use row/group B. And best to use 20G disks, 15G can be a little tight, kernel images have become bigger the past years. [15:22:30] 06SRE, 10vm-requests: codfw: 1 VM requested for ceph-admin - https://phabricator.wikimedia.org/T438038#12323559 (10MoritzMuehlenhoff) Looks good, please use row/group C. And best to use 20G disks, 15G can be a little tight, kernel images have become bigger the past years. [15:24:35] (03PS2) 10Muehlenhoff: envoy-future: Add the apt source for bookworm [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341892 [15:24:55] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12323575 (10VRiley-WMF) 05Open→03In progress Working on this now. [15:25:07] (03CR) 10Muehlenhoff: "Makes sense, I had though this was just an obsolete kludge for an old version. I've updated the patch accordingly." [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341892 (owner: 10Muehlenhoff) [15:25:35] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Remove obsolete gpu-tester image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341889 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [15:26:27] !log pruned obsolete Bullseye image amd-gpu-tester from the docker registry T416452 [15:26:29] (03CR) 10Thilio: "recheck" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341892 (owner: 10Muehlenhoff) [15:26:29] (03CR) 10RLazarus: [C:03+1] "Thanks!" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341892 (owner: 10Muehlenhoff) [15:26:30] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:26:30] T416452: Migrate Docker images running in Production away from Bullseye - https://phabricator.wikimedia.org/T416452 [15:26:41] !log jelto@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on phab2003.codfw.wmnet,phab[1005-1006].eqiad.wmnet with reason: Phabricator deploy [15:26:52] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] "amd-gpu-tester has been pruned from the Docker registry" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341889 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [15:27:18] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] envoy-future: Add the apt source for bookworm [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341892 (owner: 10Muehlenhoff) [15:30:41] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341890 (https://phabricator.wikimedia.org/T435274) (owner: 10Krinkle) [15:30:41] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [15:31:37] (03Merged) 10jenkins-bot: lift IP cap for edit-a-thon /workshop [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340558 (https://phabricator.wikimedia.org/T437609) (owner: 10Anzx) [15:31:37] !log elukey@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [15:31:43] !log elukey@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [15:31:54] !log brennen@deploy1003 Started deploy [phabricator/deployment@c386249]: deploy phab2003 for T437930 [15:31:58] T437930: Deploy Phab/Phorge 2026-09-15 - https://phabricator.wikimedia.org/T437930 [15:32:19] !log elukey@deploy1003 helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'. [15:32:24] !log elukey@deploy1003 helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'. [15:32:47] !log brennen@deploy1003 Finished deploy [phabricator/deployment@c386249]: deploy phab2003 for T437930 (duration: 00m 52s) [15:32:48] !log elukey@deploy1003 helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'. [15:32:51] !log elukey@deploy1003 helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'. [15:32:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [15:33:09] !log brennen@deploy1003 Started deploy [phabricator/deployment@c386249]: deploy phab1005 for T437930 [15:33:16] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye [15:33:25] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12323647 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host ms-be1098.eqiad.wmnet wi... [15:33:49] !log brennen@deploy1003 Finished deploy [phabricator/deployment@c386249]: deploy phab1005 for T437930 (duration: 00m 39s) [15:34:03] !log elukey@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'. [15:34:06] !log elukey@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'. [15:34:35] 06SRE, 10SRE-Access-Requests: Requesting access to Superset Dashboard for Hany EL Mokadem - https://phabricator.wikimedia.org/T438034#12323650 (10hnowlan) [15:35:45] (03PS5) 10Krinkle: varnish: Refactor normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) [15:36:06] 10ops-eqsin, 06SRE: Inbound errors on interface cr2-eqsin:xe-0/1/1 (Transit: Tata (1028127)) - https://phabricator.wikimedia.org/T437800#12323687 (10RobH) 05Open→03Declined [15:36:13] 10ops-eqsin, 06SRE: Inbound errors on interface cr2-eqsin:et-0/0/2 (Core: asw1-604-eqsin:ethernet-1/55) - https://phabricator.wikimedia.org/T436182#12323688 (10RobH) 05Open→03Declined [15:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:36:36] 10ops-eqsin, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: EQSIN:Switch refresh diagram and wiring - https://phabricator.wikimedia.org/T423724#12323689 (10RobH) 05Open→03Resolved [15:37:12] 10ops-eqsin, 06SRE: Unresponsive management for cp5022.mgmt:22 - https://phabricator.wikimedia.org/T416499#12323695 (10RobH) 05Open→03Declined [15:37:15] 10ops-eqsin, 06SRE: Unresponsive management for ganeti5007.mgmt:22 - https://phabricator.wikimedia.org/T433602#12323697 (10RobH) 05Open→03Declined [15:38:06] (03PS6) 10Krinkle: varnish: Refactor normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) [15:38:25] 06SRE, 10SRE-Access-Requests: Requesting access to Superset Dashboard for Hany EL Mokadem - https://phabricator.wikimedia.org/T438034#12323715 (10hnowlan) Hi @Hany.elmokadem, thanks for the request! I don't see an NDA on file for you - please reach out to rstallman@wikimedia.org or naramayo@wikimedia.org to re... [15:38:59] (03PS4) 10Herron: logstash: add service ordering to unit [puppet] - 10https://gerrit.wikimedia.org/r/1341921 (https://phabricator.wikimedia.org/T435265) [15:39:46] (03Merged) 10jenkins-bot: Restore table borders for client-side MathJax [extensions/Math] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1341890 (https://phabricator.wikimedia.org/T435274) (owner: 10Krinkle) [15:40:08] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1341890|Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558|lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] [15:40:17] T435274: Missing table grid borders in Client side MathJax (check polyfills for SVG support) - https://phabricator.wikimedia.org/T435274 [15:40:18] T437609: Lift IP cap on these dates 2026-09-29; 2026-10-27 and 2026-11-24 for edit-a-thon for eswiki, commons and wikidata - https://phabricator.wikimedia.org/T437609 [15:40:18] T437594: Lift IP cap on these dates 2026-09-28; 2026-10-05; 2026-10-19 and 2026-10-26 for edit-a-thon for eswiki, commons and wikidata - https://phabricator.wikimedia.org/T437594 [15:40:19] T437470: IP whitelist request for editathon at the University of Cambridge - Friday 16th October - https://phabricator.wikimedia.org/T437470 [15:40:24] (03CR) 10Herron: sre.opensearch.roll-restart-reboot: include checklist items (032 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [15:40:49] (03CR) 10Krinkle: "Interesting, the text/thumb version has different expected values. This seems to be because it is performing different normalization, but " [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [15:41:40] (03CR) 10Elukey: "It looks very good, I left some small questions in the Python script to better understand it!" [puppet] - 10https://gerrit.wikimedia.org/r/1341098 (https://phabricator.wikimedia.org/T311005) (owner: 10Ayounsi) [15:44:43] !log krinkle@deploy1003 anzx, krinkle: Backport for [[gerrit:1341890|Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558|lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:44:53] Krinkle: nothing to test on throttle, ok to continue [15:45:16] (03PS7) 10Krinkle: varnish: Refactor normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) [15:47:20] (03CR) 10Dzahn: [V:03+1 C:03+2] zuul: refactor zookeeper inclusion to fix firewalling (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [15:47:38] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052 (10RobH) 03NEW p:05Triage→03Medium [15:48:13] (03CR) 10Dzahn: [V:03+1 C:03+2] "entire error message and explanation was on the original change and/or revert of it" [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [15:49:14] (03CR) 10Dzahn: [C:03+2] ci: logoutput on failure of ci-build-images script [puppet] - 10https://gerrit.wikimedia.org/r/1333795 (https://phabricator.wikimedia.org/T436775) (owner: 10Hashar) [15:50:02] (03PS3) 10Hashar: ci: set ci-build-images image and update creates [puppet] - 10https://gerrit.wikimedia.org/r/1333796 (https://phabricator.wikimedia.org/T436775) [15:51:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:53:37] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12323874 (10RobH) [15:54:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:54:26] 10SRE-SLO, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th): page_change SLO windows - https://phabricator.wikimedia.org/T438054 (10APizzata-WMF) 03NEW [15:58:36] anzx: ack, testing my change as well. [15:58:38] (03CR) 10Dzahn: ci: set ci-build-images image and update creates (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1333796 (https://phabricator.wikimedia.org/T436775) (owner: 10Hashar) [15:59:42] !log krinkle@deploy1003 anzx, krinkle: Continuing with deployment [15:59:55] (03CR) 10Dduvall: [C:03+1] Remove obsolete buildkitd image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339620 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [16:00:05] jhathaway and rzl: #bothumor Q:Why did functions stop calling each other? A:They had arguments. Rise for Puppet request window . (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1600). [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:00:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [16:04:16] 06SRE, 10LDAP-Access-Requests: Grant Access to <500> for - https://phabricator.wikimedia.org/T437488#12323940 (10hnowlan) 05Open→03Stalled [16:04:27] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341890|Restore table borders for client-side MathJax (T435274)]], [[gerrit:1340558|lift IP cap for edit-a-thon /workshop (T437609 T437594 T437470)]] (duration: 24m 19s) [16:04:35] T435274: Missing table grid borders in Client side MathJax (check polyfills for SVG support) - https://phabricator.wikimedia.org/T435274 [16:04:36] T437609: Lift IP cap on these dates 2026-09-29; 2026-10-27 and 2026-11-24 for edit-a-thon for eswiki, commons and wikidata - https://phabricator.wikimedia.org/T437609 [16:04:36] T437594: Lift IP cap on these dates 2026-09-28; 2026-10-05; 2026-10-19 and 2026-10-26 for edit-a-thon for eswiki, commons and wikidata - https://phabricator.wikimedia.org/T437594 [16:04:36] T437470: IP whitelist request for editathon at the University of Cambridge - Friday 16th October - https://phabricator.wikimedia.org/T437470 [16:04:37] (03CR) 10DLynch: "Yeah, we'd mostly like to get this into staging so that we can start actually working out what about the environment will/won't work. I've" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [16:05:27] (03PS1) 10Krinkle: varnish: Fix thumb.wm.o normalization to match upload.wm.o [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) [16:05:32] (03PS1) 10Krinkle: varnish: Move query string strip for upload.wm.o to pre-purge [puppet] - 10https://gerrit.wikimedia.org/r/1341937 (https://phabricator.wikimedia.org/T425216) [16:05:38] (03CR) 10Krinkle: "I ran these two" [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [16:05:56] 06SRE, 10SRE-Access-Requests, 06Data-Engineering: Requesting access to deployment and analytics-privatedata-users for Cklimas - https://phabricator.wikimedia.org/T437653#12323954 (10hnowlan) 05Open→03Resolved a:03hnowlan [16:05:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [16:06:00] (03PS8) 10Krinkle: varnish: Refactor normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) [16:06:00] (03PS2) 10Krinkle: varnish: Fix thumb.wm.o normalization to match upload.wm.o [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) [16:06:00] (03PS2) 10Krinkle: varnish: Move query string strip for upload.wm.o to pre-purge [puppet] - 10https://gerrit.wikimedia.org/r/1341937 (https://phabricator.wikimedia.org/T425216) [16:09:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 19.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:11:42] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:15:16] (03PS3) 10Krinkle: varnish: Fix thumb.wm.o normalization to match upload.wm.o [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) [16:15:16] (03PS3) 10Krinkle: varnish: Move query string strip for upload.wm.o to pre-purge [puppet] - 10https://gerrit.wikimedia.org/r/1341937 (https://phabricator.wikimedia.org/T425216) [16:16:22] (03CR) 10Krinkle: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [16:16:36] PROBLEM - Host arclamp2001 is DOWN: PING CRITICAL - Packet loss = 100% [16:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:19:45] (03CR) 10Dzahn: [V:03+1] "https://puppet-compiler.wmflabs.org/output/1327569/9419/" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [16:21:18] !log temp disabling puppet on C:zookeeper (32 hosts) - safe deploy of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327569 [16:21:19] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:21:31] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [16:21:56] (03CR) 10Dzahn: [V:03+1] "16:21 < mutante> !log temp disabling puppet on C:zookeeper (32 hosts) - safe deploy of https://gerrit.wikimedia.org/r/c/operations/puppet/" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [16:22:00] (03CR) 10Dzahn: [V:03+1 C:03+2] zookeeper: fix log4j addition when tls is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [16:23:06] (03CR) 10Dzahn: [V:03+1 C:03+2] "> I found Zookeeper on `zuul1001.eqiad.wmnet` does not produce any log under `/var/log/zookeeper`. It is an issue in the Bookworm/Trixie " [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [16:27:12] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:33:51] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS bullseye [16:34:04] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12324119 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host ms-be1098.eqiad.wmnet wi... [16:38:32] (03CR) 10JHathaway: [C:03+2] rake_modules: support Debian 13 (Trixie) facts [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [16:40:24] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12324169 (10VRiley-WMF) [16:47:05] (03CR) 10Dzahn: [V:03+1 C:03+2] "I can't confirm this. This changes more than just order. For example on cloudcontrol:" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [16:48:04] (03PS1) 10Bking: apt: Add repo component for OpenSearch 3 [puppet] - 10https://gerrit.wikimedia.org/r/1341948 (https://phabricator.wikimedia.org/T433697) [16:48:59] (03PS2) 10Bking: apt: Add repo component for OpenSearch 3 [puppet] - 10https://gerrit.wikimedia.org/r/1341948 (https://phabricator.wikimedia.org/T433697) [16:49:06] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1341948 (https://phabricator.wikimedia.org/T433697) (owner: 10Bking) [16:49:39] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [16:49:45] 10ops-eqiad, 06DC-Ops: Alert for device ps1-d1-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438067 (10phaultfinder) 03NEW [16:51:50] (03PS1) 10Subramanya Sastry: Parsoid Read Views: Enable on 60 wikiquote wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341949 (https://phabricator.wikimedia.org/T437917) [16:52:45] (03CR) 10Hnowlan: "Based on successful testing, I am going to give a cautious +1. I'd be curious to see if @rkemper@wikimedia.org or @bking@wikimedia.org hav" [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [16:53:53] !log vriley@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003" [16:53:57] !log vriley@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [ms-be1099] - vriley@cumin1003" [16:53:57] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:54:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.8% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:54:18] vriley@cumin1003 reimage (PID 202162) is awaiting input [16:54:38] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1098.eqiad.wmnet with OS bullseye [16:54:54] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12324261 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host ms-be1098.eqiad.wmnet with O... [16:54:59] !log vriley@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099 [16:55:00] !log vriley@cumin1003 END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099 [16:55:22] !log vriley@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099 [16:55:23] !log vriley@cumin1003 END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099 [16:55:51] (03PS9) 10CDobbins: prometheus: fix prometheus-ferm-mss.py [puppet] - 10https://gerrit.wikimedia.org/r/1333244 (https://phabricator.wikimedia.org/T433672) [16:55:55] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [16:55:56] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [16:56:46] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [16:57:04] (03CR) 10CI reject: [V:04-1] prometheus: fix prometheus-ferm-mss.py [puppet] - 10https://gerrit.wikimedia.org/r/1333244 (https://phabricator.wikimedia.org/T433672) (owner: 10CDobbins) [16:58:41] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Eqiad: replace unmanaged msw in racks A4, B7 & C5 - https://phabricator.wikimedia.org/T438071 (10cmooney) 03NEW p:05Triage→03Low [16:58:53] (03PS1) 10Dzahn: zookeeper: add feature flag to enable log4j [puppet] - 10https://gerrit.wikimedia.org/r/1341950 (https://phabricator.wikimedia.org/T435503) [16:59:18] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:59:44] !log vriley@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host ms-be1099 [16:59:45] !log vriley@cumin1003 END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host ms-be1099 [16:59:56] (03CR) 10CI reject: [V:04-1] zookeeper: add feature flag to enable log4j [puppet] - 10https://gerrit.wikimedia.org/r/1341950 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [16:59:58] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:00:00] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:00:04] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1700) [17:01:37] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:01:38] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:02:57] !log brett@puppetserver1001 conftool action : set/pooled=no; selector: name=cp7009.* [17:03:21] 10ops-eqiad, 06DC-Ops: Alert for device ps1-d1-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438067#12324333 (10VRiley-WMF) a:03VRiley-WMF [17:04:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:05:28] (03CR) 10Andrew Bogott: Export a few stats about the magnum capi worker cluster (037 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1339182 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [17:05:52] PROBLEM - Thanos swift https on thanos-fe1007 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Thanos [17:05:52] PROBLEM - Thanos swift https on thanos-fe1006 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Thanos [17:06:02] (03PS5) 10Andrew Bogott: Export a few stats about the magnum capi worker cluster [puppet] - 10https://gerrit.wikimedia.org/r/1339182 (https://phabricator.wikimedia.org/T429557) [17:07:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:07:47] (03PS1) 10BCornwall: site: Move cp7009 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341951 (https://phabricator.wikimedia.org/T436363) [17:08:42] RECOVERY - Thanos swift https on thanos-fe1007 is OK: HTTP OK: HTTP/1.1 200 OK - 279 bytes in 0.079 second response time https://wikitech.wikimedia.org/wiki/Thanos [17:08:42] RECOVERY - Thanos swift https on thanos-fe1006 is OK: HTTP OK: HTTP/1.1 200 OK - 279 bytes in 0.060 second response time https://wikitech.wikimedia.org/wiki/Thanos [17:09:23] (03PS1) 10Ladsgroup: thumbor: Enable cache expiry TTL on two containers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341952 (https://phabricator.wikimedia.org/T433964) [17:12:03] FIRING: ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [17:12:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.04% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [17:13:53] 10ops-eqiad, 06DC-Ops: Alert for device ps1-d1-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438067#12324393 (10VRiley-WMF) 05Open→03Resolved Rebalanced powerr [17:14:19] (03CR) 10Ssingh: [C:03+1] site: Move cp7009 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341951 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [17:14:33] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1339182 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [17:14:44] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12324400 (10VRiley-WMF) [17:15:00] (03CR) 10Ladsgroup: "Reference for config I7b712d3283154" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341952 (https://phabricator.wikimedia.org/T433964) (owner: 10Ladsgroup) [17:15:19] (03PS1) 10BCornwall: site: Move cp404[56] from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341957 (https://phabricator.wikimedia.org/T436363) [17:15:27] (03CR) 10BCornwall: [C:03+2] site: Move cp7009 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341951 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [17:16:01] (03PS2) 10Dzahn: zookeeper: add feature flag to enable log4j [puppet] - 10https://gerrit.wikimedia.org/r/1341950 (https://phabricator.wikimedia.org/T435503) [17:16:11] jhathaway: cool if I merge your change in? [17:17:03] RESOLVED: ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [17:17:15] (03CR) 10CDobbins: [C:03+2] site: Move cp7009 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341951 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [17:18:37] (03PS1) 10Ssingh: site.pp/conftool: move cp1101 to text from upload [puppet] - 10https://gerrit.wikimedia.org/r/1341958 (https://phabricator.wikimedia.org/T436363) [17:18:39] (03PS1) 10Ssingh: site.pp/conftool: move cp1103 to text from upload [puppet] - 10https://gerrit.wikimedia.org/r/1341959 (https://phabricator.wikimedia.org/T436363) [17:19:25] (03CR) 10BCornwall: [C:03+1] site.pp/conftool: move cp1103 to text from upload [puppet] - 10https://gerrit.wikimedia.org/r/1341959 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [17:19:47] (03CR) 10BCornwall: [C:03+1] site.pp/conftool: move cp1101 to text from upload [puppet] - 10https://gerrit.wikimedia.org/r/1341958 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [17:21:00] jhathaway: I'm going to merge in the changes [17:21:02] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [17:21:09] brett: +1 [17:22:38] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:22:39] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:22:59] (03CR) 10Ssingh: [C:03+1] site: Move cp404[56] from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341957 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [17:23:09] !log sukhe@puppetserver1001 conftool action : set/pooled=no; selector: name=cp1101.eqiad.wmnet [17:23:20] (03CR) 10Ssingh: [C:03+2] site.pp/conftool: move cp1101 to text from upload [puppet] - 10https://gerrit.wikimedia.org/r/1341958 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [17:23:59] (03PS3) 10Dzahn: zookeeper: add feature flag to enable log4j [puppet] - 10https://gerrit.wikimedia.org/r/1341950 (https://phabricator.wikimedia.org/T435503) [17:24:04] 10ops-eqsin: Unresponsive management for ganeti5007.mgmt:22 - https://phabricator.wikimedia.org/T438073 (10phaultfinder) 03NEW [17:24:21] (03CR) 10Dzahn: [C:03+2] zookeeper: add feature flag to enable log4j [puppet] - 10https://gerrit.wikimedia.org/r/1341950 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [17:24:51] !log cdobbins@puppetserver1001 conftool action : set/pooled=no; selector: name=cp7009.magru.wmnet [17:25:35] !log sukhe@cumin1004 START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie [17:26:14] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie [17:26:51] (03CR) 10Dzahn: [C:03+2] "doing this because I already have merged the parent and disabled puppet on all zookeeper servers and I want to re-enable it without changi" [puppet] - 10https://gerrit.wikimedia.org/r/1341950 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [17:28:51] !log brett@puppetserver1001 conftool action : set/pooled=no; selector: name=cp4045.* [17:28:54] !log brett@puppetserver1001 conftool action : set/pooled=no; selector: name=cp4046.* [17:29:09] (03CR) 10BCornwall: [C:03+2] site: Move cp404[56] from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341957 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [17:30:16] (03PS1) 10Ssingh: site: Move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341970 (https://phabricator.wikimedia.org/T436363) [17:30:18] (03PS1) 10Ssingh: site: Move cp3075 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341971 (https://phabricator.wikimedia.org/T436363) [17:30:52] (03CR) 10BCornwall: [C:03+1] site: Move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341970 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [17:31:33] (03CR) 10BCornwall: [C:03+1] site: Move cp3075 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341971 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [17:33:38] !log sukhe@puppetserver1001 conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet [17:34:14] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie [17:34:20] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie [17:34:44] (03PS1) 10CDobbins: site: move cp2044 and cp2046 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341973 (https://phabricator.wikimedia.org/T436363) [17:34:54] (03CR) 10Ssingh: [C:03+2] site: Move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341970 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [17:35:36] (03PS1) 10Dzahn: zookeeper: pass through new feature flag to enable log4j from profile [puppet] - 10https://gerrit.wikimedia.org/r/1341974 (https://phabricator.wikimedia.org/T435503) [17:36:49] !log sukhe@cumin1004 START - Cookbook sre.hosts.reimage for host cp3074.esams.wmnet with OS trixie [17:37:29] (03PS2) 10CDobbins: site: move cp2044 and cp2046 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341973 (https://phabricator.wikimedia.org/T436363) [17:38:58] (03PS2) 10Dzahn: zookeeper: pass through new feature flag to enable log4j on zuul hosts [puppet] - 10https://gerrit.wikimedia.org/r/1341974 (https://phabricator.wikimedia.org/T435503) [17:39:28] (03PS3) 10CDobbins: site: move cp2044 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341973 (https://phabricator.wikimedia.org/T436363) [17:39:31] (03CR) 10Ssingh: site: move cp2044 from upload to text (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341973 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [17:39:55] (03CR) 10Ssingh: [C:03+1] site: move cp2044 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341973 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [17:40:36] (03PS3) 10Dzahn: zookeeper: pass through new feature flag to enable log4j on zuul hosts [puppet] - 10https://gerrit.wikimedia.org/r/1341974 (https://phabricator.wikimedia.org/T435503) [17:40:51] (03PS1) 10Ssingh: site: Move cp6001 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341976 (https://phabricator.wikimedia.org/T436363) [17:40:53] (03CR) 10CDobbins: site: move cp2044 from upload to text (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341973 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [17:40:53] (03PS1) 10Ssingh: site: Move cp6002 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341977 (https://phabricator.wikimedia.org/T436363) [17:41:34] (03CR) 10CDobbins: [C:03+2] site: move cp2044 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341973 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [17:43:43] !log cdobbins@puppetserver1001 conftool action : set/pooled=no; selector: name=cp2044.codfw.wmnet [17:44:27] (03CR) 10Dzahn: [V:03+1 C:03+2] "https://puppet-compiler.wmflabs.org/output/1341974/9429/cloudcontrol2010-dev.codfw.wmnet/index.html" [puppet] - 10https://gerrit.wikimedia.org/r/1341974 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [17:45:25] sukhe: multiple, let's get it merged?:) [17:45:33] mutante: ha sorry [17:45:34] checking [17:45:44] yes plesae [17:45:55] doing! [17:46:39] and.. it's live. synced. [17:47:04] nice thanks [17:48:50] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host cp2044.codfw.wmnet with OS trixie [17:52:24] (03CR) 10BCornwall: [V:03+2 C:03+1] varnish: Refactor normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [17:52:54] !log brett@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage [17:53:26] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12324568 (10RobH) [17:53:29] (03CR) 10Ebernhardson: [C:03+1] apt: Add repo component for OpenSearch 3 [puppet] - 10https://gerrit.wikimedia.org/r/1341948 (https://phabricator.wikimedia.org/T433697) (owner: 10Bking) [17:54:39] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12324576 (10RobH) [17:55:08] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and restricted for cooltey - https://phabricator.wikimedia.org/T437658#12324577 (10cooltey) >>! In T437658#12320527, @thcipriani wrote: > Reason for access makes sense for `deployment` group membership. Approved! > > No reason to add to `restri... [17:56:03] !log eevans@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore1006.eqiad.wmnet with reason: Firmware upgrades — T437516 [17:56:19] !log eevans@cumin1004 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore1006.eqiad.wmnet [17:57:59] (03PS1) 10PipelineBot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341980 [17:58:24] (03PS1) 10BCornwall: preseed: Move all cacheproxy hosts to EFI config [puppet] - 10https://gerrit.wikimedia.org/r/1341982 [17:58:33] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage [17:59:28] (03CR) 10Ssingh: preseed: Move all cacheproxy hosts to EFI config (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341982 (owner: 10BCornwall) [17:59:53] (03CR) 10BCornwall: preseed: Move all cacheproxy hosts to EFI config (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341982 (owner: 10BCornwall) [18:00:05] jnuche and dduvall: Time to do the MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T1800). [18:01:28] !log brett@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp4046.ulsfo.wmnet with OS trixie [18:01:34] !log brett@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp4045.ulsfo.wmnet with OS trixie [18:02:29] (03CR) 10Ssingh: [C:03+1] preseed: Move all cacheproxy hosts to EFI config (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341982 (owner: 10BCornwall) [18:02:45] (03CR) 10BCornwall: [C:03+2] preseed: Move all cacheproxy hosts to EFI config [puppet] - 10https://gerrit.wikimedia.org/r/1341982 (owner: 10BCornwall) [18:03:46] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [18:03:51] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [18:03:55] !log sukhe@cumin1004 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp1101.eqiad.wmnet with OS trixie [18:04:14] !log sukhe@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on cp3074.esams.wmnet with reason: host reimage [18:04:57] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage [18:07:25] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3074.esams.wmnet with reason: host reimage [18:08:47] brett, godog, sukhe, cdobbins: stashbot was gone for ca. 10 minutes, you might want to re-!log those last few messages [18:08:50] !log eevans@cumin1004 START - Cookbook sre.hosts.reboot-single for host sessionstore1006.eqiad.wmnet [18:09:03] lucaswerkmeister: that's the cookbook doing it but thanks! [18:09:18] yeah probably harmless but I thought I’d mention it :) [18:09:27] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie [18:09:42] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2044.codfw.wmnet with reason: host reimage [18:10:05] sukhe@cumin1004 reimage (PID 2840122) is awaiting input [18:10:14] !log sukhe@cumin1004 START - Cookbook sre.hosts.reimage for host cp1101.eqiad.wmnet with OS trixie [18:12:07] FIRING: ProbeDown: Service sessionstore1006-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://wikitech.wikimedia.org/wiki/TLS/Runbook#sessionstore1006-a:9042 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:14:47] (03PS1) 10Dzahn: zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) [18:14:55] (03CR) 10Bking: [C:03+2] apt: Add repo component for OpenSearch 3 [puppet] - 10https://gerrit.wikimedia.org/r/1341948 (https://phabricator.wikimedia.org/T433697) (owner: 10Bking) [18:15:05] (03PS2) 10Dzahn: zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) [18:16:21] (03CR) 10CI reject: [V:04-1] zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [18:17:07] FIRING: [2x] ProbeDown: Service sessionstore1006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:17:54] (03PS3) 10Dzahn: zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) [18:18:09] (03PS4) 10Dzahn: zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) [18:19:07] (03CR) 10CI reject: [V:04-1] zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [18:19:15] (03PS5) 10Dzahn: zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) [18:19:23] (03PS6) 10Dzahn: zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) [18:19:26] !log eevans@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore1006.eqiad.wmnet [18:19:28] !log eevans@cumin1004 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore1006.eqiad.wmnet [18:19:44] !log eevans@cumin1004 START - Cookbook sre.hosts.provision for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [18:20:22] !log eevans@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [18:20:32] (03CR) 10CI reject: [V:04-1] zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [18:20:44] !log eevans@cumin1004 START - Cookbook sre.hosts.reimage for host sessionstore1006.eqiad.wmnet with OS bookworm [18:22:07] RESOLVED: [2x] ProbeDown: Service sessionstore1006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:22:16] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie [18:23:29] (03PS7) 10Dzahn: zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) [18:26:17] !log sukhe@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage [18:27:07] FIRING: [2x] ProbeDown: Service sessionstore1006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:29:04] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 15 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341271 (https://phabricator.wikimedia.org/T437916) (owner: 10Subramanya Sastry) [18:29:26] !log brett@cumin2003 START - Cookbook sre.hosts.provision for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [18:30:19] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1101.eqiad.wmnet with reason: host reimage [18:31:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:31:25] (03CR) 10Dzahn: [C:03+2] zookeeper: log4j is already enabled on all bookworm machines [puppet] - 10https://gerrit.wikimedia.org/r/1341997 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [18:31:33] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3074.esams.wmnet with OS trixie [18:32:18] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2044.codfw.wmnet with OS trixie [18:33:20] !log sukhe@puppetserver1001 conftool action : set/weight=1; selector: name=cp3074.esams.wmnet [18:36:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.87% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:36:40] !log sukhe@cumin1004 START - Cookbook sre.loadbalancer.admin config_reloading P{lvs3008.esams.wmnet} and A:liberica [18:36:58] !log sukhe@cumin1004 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading P{lvs3008.esams.wmnet} and A:liberica [18:38:34] (03PS1) 10CDobbins: site: move cp2046 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342004 (https://phabricator.wikimedia.org/T436363) [18:39:25] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet [18:40:23] !log eevans@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage [18:40:24] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4045.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [18:43:00] (03CR) 10Ssingh: "This needs a rebase since we are affecting other stuff?" [puppet] - 10https://gerrit.wikimedia.org/r/1342004 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [18:44:24] !log eevans@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1006.eqiad.wmnet with reason: host reimage [18:44:30] !log cdobbins@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp2044.codfw.wmnet [18:44:34] (03PS1) 10Dzahn: zookeeper: ensure the order of included logging jars stays the same [puppet] - 10https://gerrit.wikimedia.org/r/1342005 (https://phabricator.wikimedia.org/T435503) [18:45:14] !log cdobbins@puppetserver1001 conftool action : set/weight=1; selector: name=cp2044.codfw.wmnet [18:45:55] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12324794 (10ssingh) [18:47:44] (03CR) 10Dzahn: "this is ensuring that there is a true noop when puppet gets reenabled on hosts using zookeeper" [puppet] - 10https://gerrit.wikimedia.org/r/1342005 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [18:47:56] (03CR) 10Dzahn: [C:03+2] zookeeper: ensure the order of included logging jars stays the same [puppet] - 10https://gerrit.wikimedia.org/r/1342005 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [18:49:50] (03PS2) 10CDobbins: site: move cp2046 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342004 (https://phabricator.wikimedia.org/T436363) [18:50:07] (03Abandoned) 10Ssingh: site: Move cp6001 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341976 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [18:50:14] (03Abandoned) 10Ssingh: site: Move cp6002 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341977 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [18:52:49] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1101.eqiad.wmnet with OS trixie [18:52:55] !log brett@cumin2003 START - Cookbook sre.hosts.provision for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [18:54:21] (03CR) 10Ssingh: site: move cp2046 from upload to text (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1342004 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [18:55:10] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp4045.ulsfo.wmnet with OS trixie [18:55:37] !log sukhe@puppetserver1001 conftool action : set/weight=1; selector: name=cp1101.eqiad.wmnet [18:55:41] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp1101.eqiad.wmnet [18:57:58] !log sukhe@puppetserver1001 conftool action : set/pooled=no; selector: name=cp3075.esams.wmnet [18:58:15] !log brett@cumin2003 START - Cookbook sre.hosts.provision for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:00:08] !log sukhe@cumin1004 START - Cookbook sre.hosts.provision for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:00:59] (03CR) 10BCornwall: [C:04-1] "This is altering much more than cp2046!" [puppet] - 10https://gerrit.wikimedia.org/r/1342004 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [19:01:13] (03CR) 10BCornwall: [C:04-1] "marking unresolved." [puppet] - 10https://gerrit.wikimedia.org/r/1342004 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [19:01:16] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet [19:01:19] !log sukhe@puppetserver1001 conftool action : set/pooled=no; selector: name=cp1103.eqiad.wmnet [19:01:45] (03PS2) 10Ssingh: site.pp/conftool: move cp1103 to text from upload [puppet] - 10https://gerrit.wikimedia.org/r/1341959 (https://phabricator.wikimedia.org/T436363) [19:02:02] (03CR) 10Ssingh: "rebase, no code change" [puppet] - 10https://gerrit.wikimedia.org/r/1341959 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [19:02:24] PROBLEM - Host cp7009 is DOWN: PING CRITICAL - Packet loss = 100% [19:02:24] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp4046.mgmt.ulsfo.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:03:06] (03PS2) 10Ssingh: site: Move cp3075 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341971 (https://phabricator.wikimedia.org/T436363) [19:03:18] PROBLEM - Host cp3075 is DOWN: PING CRITICAL - Packet loss = 100% [19:03:40] (03PS1) 10Jforrester: abstractwiki: Add three new articles per community advice to show off the feature [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342008 (https://phabricator.wikimedia.org/T434227) [19:04:00] !log brett@puppetserver1001 conftool action : set/pooled=no; selector: name=cp3075.* [19:05:02] !log eevans@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1006.eqiad.wmnet with OS bookworm [19:05:55] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp4046.ulsfo.wmnet with OS trixie [19:08:13] !log sukhe@cumin1004 START - Cookbook sre.hosts.provision for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:08:46] !log sukhe@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp3075.esams.wmnet with reason: reimaging [19:09:11] (03PS1) 10Dzahn: zookeeper: add slf4j-api.jar to classpath in the exact same order [puppet] - 10https://gerrit.wikimedia.org/r/1342011 (https://phabricator.wikimedia.org/T435503) [19:09:48] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp7009.mgmt.magru.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:10:05] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp7009.magru.wmnet with OS trixie [19:11:04] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp3075.mgmt.esams.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:11:44] PROBLEM - Host cp1103 is DOWN: PING CRITICAL - Packet loss = 100% [19:12:13] !log ssastry@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [19:12:14] !log brett@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage [19:12:33] !log sukhe@cumin1004 START - Cookbook sre.hosts.reimage for host cp3075.esams.wmnet with OS trixie [19:12:44] !log ssastry@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [19:12:44] (03CR) 10Ssingh: [C:03+2] site: Move cp3075 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1341971 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [19:12:45] !log ssastry@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [19:13:15] !log ssastry@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [19:16:35] (03PS2) 10Dzahn: zookeeper: add slf4j-api.jar to classpath in the exact same order [puppet] - 10https://gerrit.wikimedia.org/r/1342011 (https://phabricator.wikimedia.org/T435503) [19:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:18:53] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4045.ulsfo.wmnet with reason: host reimage [19:19:02] !log sukhe@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp1103.eqiad.wmnet with reason: reimage [19:19:26] (03CR) 10Dzahn: [C:03+2] "https://puppet-compiler.wmflabs.org/output/1342011/9434/conf2006.codfw.wmnet/index.html" [puppet] - 10https://gerrit.wikimedia.org/r/1342011 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [19:19:34] (03CR) 10Dzahn: [V:03+1 C:03+2] "https://puppet-compiler.wmflabs.org/output/1342011/9434/conf2006.codfw.wmnet/index.html" [puppet] - 10https://gerrit.wikimedia.org/r/1342011 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [19:19:36] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp1103.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [19:20:01] (03CR) 10Ssingh: [C:03+2] site.pp/conftool: move cp1103 to text from upload [puppet] - 10https://gerrit.wikimedia.org/r/1341959 (https://phabricator.wikimedia.org/T436363) (owner: 10Ssingh) [19:20:19] sukhe: yes:) [19:20:29] thanks! [19:21:07] (03PS1) 10Arendpieter: CommonSettings: Send script-src-attr 'none' on auth.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342015 (https://phabricator.wikimedia.org/T419684) [19:21:35] !log sukhe@cumin1004 START - Cookbook sre.hosts.reimage for host cp1103.eqiad.wmnet with OS trixie [19:22:29] I am just making sure that no zookeeper on any conf* or other host is changed in any way - while also fixing logging and TLS for zookeeper used by zuul. It may look like a lot of changes but it's not. Re-enabling puppet once it's a NOOP everywhere else. [19:23:41] previous change was changing the order of things appearing in CLASSPATH. even when the contents are the same, just the order itself could also cause issues [19:23:55] !log brett@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage [19:25:04] (03PS1) 10Eevans: Revert "sessionstore1005: partition tables lost; full reimage needed" [puppet] - 10https://gerrit.wikimedia.org/r/1342017 [19:25:29] (03PS2) 10Eevans: Revert "sessionstore1005: partition tables lost; full reimage needed" [puppet] - 10https://gerrit.wikimedia.org/r/1342017 (https://phabricator.wikimedia.org/T437516) [19:26:59] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12324913 (10RobH) [19:27:07] RESOLVED: [2x] ProbeDown: Service sessionstore1006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:27:27] (03PS3) 10Eevans: Revert "sessionstore1005: partition tables lost; full reimage needed" [puppet] - 10https://gerrit.wikimedia.org/r/1342017 (https://phabricator.wikimedia.org/T437516) [19:27:47] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp4046.ulsfo.wmnet with reason: host reimage [19:28:54] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [19:30:02] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1341921 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [19:31:24] (03CR) 10Eevans: [C:03+2] Revert "sessionstore1005: partition tables lost; full reimage needed" [puppet] - 10https://gerrit.wikimedia.org/r/1342017 (https://phabricator.wikimedia.org/T437516) (owner: 10Eevans) [19:32:07] !log brett@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp7009.magru.wmnet with reason: host reimage [19:33:36] !log sukhe@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on cp3075.esams.wmnet with reason: host reimage [19:34:31] (03PS1) 10BCornwall: [beta] Update wgCdnServersNoPurge for new cp hosts [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) [19:35:22] (03CR) 10CI reject: [V:04-1] [beta] Update wgCdnServersNoPurge for new cp hosts [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) (owner: 10BCornwall) [19:35:58] (03PS2) 10BCornwall: [beta] Update wgCdnServersNoPurge for new cp hosts [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) [19:37:39] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp7009.magru.wmnet with reason: host reimage [19:39:04] !log sukhe@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage [19:39:44] (03PS1) 10CDobbins: site: move cp2046 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342027 (https://phabricator.wikimedia.org/T436363) [19:40:07] (03Abandoned) 10CDobbins: site: move cp2046 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342004 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [19:41:51] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3075.esams.wmnet with reason: host reimage [19:42:04] (03CR) 10Southparkfan: [C:03+1] [beta] Update wgCdnServersNoPurge for new cp hosts [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) (owner: 10BCornwall) [19:42:21] (03CR) 10Hashar: "The issue is not about the Zuul hosts per see but on any host having Trixie." [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [19:43:31] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4045.ulsfo.wmnet with OS trixie [19:45:11] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp1103.eqiad.wmnet with reason: host reimage [19:46:00] (03CR) 10BCornwall: [C:03+1] site: move cp2046 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342027 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [19:46:03] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 15 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) (owner: 10BCornwall) [19:48:31] (03CR) 10Southparkfan: [beta] Update wgCdnServersNoPurge for new cp hosts [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) (owner: 10BCornwall) [19:48:50] (03CR) 10CDobbins: [C:03+2] site: move cp2046 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342027 (https://phabricator.wikimedia.org/T436363) (owner: 10CDobbins) [19:48:54] (03PS3) 10BCornwall: [beta] Update wgCdnServersNoPurge for new cp hosts [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) [19:49:18] (03CR) 10Southparkfan: [C:03+1] [beta] Update wgCdnServersNoPurge for new cp hosts [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) (owner: 10BCornwall) [19:49:40] (03PS1) 10Dzahn: zookeeper: flatten/filter paths to avoid empty elements in class path [puppet] - 10https://gerrit.wikimedia.org/r/1342030 (https://phabricator.wikimedia.org/T435503) [19:49:53] (03PS2) 10Dzahn: zookeeper: flatten/filter paths to avoid empty elements in class path [puppet] - 10https://gerrit.wikimedia.org/r/1342030 (https://phabricator.wikimedia.org/T435503) [19:49:57] !log brett@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp4045.* [19:51:29] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp4046.ulsfo.wmnet with OS trixie [19:51:49] !log cdobbins@puppetserver1001 conftool action : set/pooled=no; selector: name=cp2046.codfw.wmnet [19:52:01] !log brett@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp4046.* [19:53:31] (03CR) 10Dzahn: [V:03+1] "https://puppet-compiler.wmflabs.org/output/1342030/9436/flink-zk2003.codfw.wmnet/index.html" [puppet] - 10https://gerrit.wikimedia.org/r/1342030 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [19:53:32] !log brett@puppetserver1001 conftool action : set/weight=1; selector: name=cp4046.* [19:53:44] !log brett@puppetserver1001 conftool action : set/weight=1; selector: name=cp4045.* [19:57:01] (03CR) 10BCornwall: "Seems the second entry is failing (loop is 0-index)" [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [19:57:05] (03CR) 10BCornwall: [V:04-1] varnish: Fix thumb.wm.o normalization to match upload.wm.o [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T2000). [20:00:05] lwatson, arlolra, and Southparkfan: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:44] (03CR) 10BCornwall: "@bblack@wikimedia.org This one is a little scary to me and I smell one of those awful edge-case scenarios that might be in your brain. Do " [puppet] - 10https://gerrit.wikimedia.org/r/1341937 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [20:01:06] o/ [20:01:13] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp7009.magru.wmnet with OS trixie [20:01:26] im here [20:01:35] same [20:02:12] !log brett@puppetserver1001 conftool action : set/weight=1; selector: name=cp7009.* [20:02:17] !log brett@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp7009.* [20:02:24] !log brett@puppetserver1001 conftool action : set/pooled=no; selector: name=cp7009.* [20:03:48] lwatson: do you need me to deploy or do you want to get started? [20:04:43] !log cdobbins@cumin1004 START - Cookbook sre.hosts.reimage for host cp2046.codfw.wmnet with OS trixie [20:04:54] if possible can you deploy it [20:04:58] !log brett@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp7009.* [20:05:01] sure [20:05:08] (03CR) 10Dzahn: [V:03+1 C:03+2] zookeeper: flatten/filter paths to avoid empty elements in class path [puppet] - 10https://gerrit.wikimedia.org/r/1342030 (https://phabricator.wikimedia.org/T435503) (owner: 10Dzahn) [20:05:10] do you need anything? [20:05:34] Just for you to be around to test anything when it's on the test servers [20:05:40] ok sure [20:05:53] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339741 (https://phabricator.wikimedia.org/T438009) (owner: 10LWatson) [20:06:35] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3075.esams.wmnet with OS trixie [20:06:52] (03Merged) 10jenkins-bot: Enable ReaderExperiments in eswiki, jawiki, and ptwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339741 (https://phabricator.wikimedia.org/T438009) (owner: 10LWatson) [20:07:17] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1339741|Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] [20:07:20] T438009: Enable ReaderExperiments on all target wikis - https://phabricator.wikimedia.org/T438009 [20:07:29] !log bking@ganeti1046 sudo gnt-instance modify -B memory=4g,vcpus=4 dse-k8s-etcd100[1-3].eqiad.wmnet T438084 [20:07:32] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:07:32] T438084: dse-k8s-etcd: Increase vCPU count - https://phabricator.wikimedia.org/T438084 [20:08:09] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp1103.eqiad.wmnet with OS trixie [20:08:32] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin config_reloading A:liberica-ulsfo (T436363) [20:08:35] T436363: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363 [20:10:33] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) config_reloading A:liberica-ulsfo (T436363) [20:11:38] !log arlolra@deploy1003 lwatson, arlolra: Backport for [[gerrit:1339741|Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:12:39] lwatson: let me know when you think it's good to proceed [20:13:32] ok let me check [20:16:23] PROBLEM - statsv Varnishkafka log producer on cp1114 is CRITICAL: PROCS CRITICAL: 2 processes with args /usr/bin/varnishkafka -S /etc/varnishkafka/statsv.conf https://wikitech.wikimedia.org/wiki/Analytics/Systems/Varnishkafka [20:16:32] !log brett@cumin2003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo (T436363) [20:16:36] T436363: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363 [20:16:38] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin depooling P{lvs4008.ulsfo.wmnet} and A:liberica [20:16:51] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs4008.ulsfo.wmnet} and A:liberica [20:16:57] !log sukhe@puppetserver1001 conftool action : set/weight=1; selector: name=cp1103.eqiad.wmnet [20:17:00] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin pooling P{lvs4008.ulsfo.wmnet} and A:liberica [20:17:03] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp1103.eqiad.wmnet [20:17:12] !log sukhe@puppetserver1001 conftool action : set/weight=1; selector: name=cp3075.esams.wmnet [20:17:17] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp3075.esams.wmnet [20:17:23] RECOVERY - statsv Varnishkafka log producer on cp1114 is OK: PROCS OK: 1 process with args /usr/bin/varnishkafka -S /etc/varnishkafka/statsv.conf https://wikitech.wikimedia.org/wiki/Analytics/Systems/Varnishkafka [20:17:23] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs4008.ulsfo.wmnet} and A:liberica [20:17:43] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin depooling P{lvs4009.ulsfo.wmnet} and A:liberica [20:17:55] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs4009.ulsfo.wmnet} and A:liberica [20:18:18] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin pooling P{lvs4009.ulsfo.wmnet} and A:liberica [20:18:30] looks fine on my end. ready to proceed [20:18:38] thanks [20:18:44] !log arlolra@deploy1003 lwatson, arlolra: Continuing with deployment [20:18:45] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs4009.ulsfo.wmnet} and A:liberica [20:19:05] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin depooling P{lvs4010.ulsfo.wmnet} and A:liberica [20:19:16] sorry for all that spam ._. [20:19:17] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs4010.ulsfo.wmnet} and A:liberica [20:19:26] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin pooling P{lvs4010.ulsfo.wmnet} and A:liberica [20:19:50] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs4010.ulsfo.wmnet} and A:liberica [20:19:53] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo (T436363) [20:20:57] !log cdobbins@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage [20:21:46] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [20:22:46] (03CR) 10Herron: "ran PS25 via test-cookbook on the eqiad cluster today:" [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [20:23:15] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1339741|Enable ReaderExperiments in eswiki, jawiki, and ptwiki (T438009)]] (duration: 15m 58s) [20:23:18] T438009: Enable ReaderExperiments on all target wikis - https://phabricator.wikimedia.org/T438009 [20:24:07] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2046.codfw.wmnet with reason: host reimage [20:25:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.38% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:25:35] Southparkfan: do you want to go next? [20:25:46] Yes, fine with me [20:25:56] do you need me to deploy? [20:26:31] Yes. I can perform the deploy directly on the Beta Cluster, and the change should not affect production at all (it's inside an if realm condition), but I cannot +2 myself [20:26:53] (03CR) 10Arlolra: Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341271 (https://phabricator.wikimedia.org/T437916) (owner: 10Subramanya Sastry) [20:27:27] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:28:25] (03PS1) 10BCornwall: site: Move cp6001 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342039 (https://phabricator.wikimedia.org/T436363) [20:28:30] ok, will do [20:29:29] (03PS1) 10BCornwall: site: Move cp6002 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342040 (https://phabricator.wikimedia.org/T436363) [20:29:35] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) (owner: 10BCornwall) [20:29:57] (03CR) 10Ssingh: [C:03+1] site: Move cp6001 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342039 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [20:30:39] (03Merged) 10jenkins-bot: [beta] Update wgCdnServersNoPurge for new cp hosts [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342025 (https://phabricator.wikimedia.org/T436468) (owner: 10BCornwall) [20:30:45] (03CR) 10Ssingh: [C:03+1] site: Move cp6002 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342040 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [20:30:56] !log brett@puppetserver1001 conftool action : set/pooled=no; selector: name=cp6001.* [20:33:41] (03CR) 10BCornwall: [C:03+2] site: Move cp6001 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342039 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [20:33:52] (03PS2) 10Arlolra: Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341271 (https://phabricator.wikimedia.org/T437916) (owner: 10Subramanya Sastry) [20:34:24] !log brett@cumin2003 START - Cookbook sre.hosts.provision for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [20:35:53] Southparkfan: all good? [20:36:02] Yes, works fine here [20:36:55] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341271 (https://phabricator.wikimedia.org/T437916) (owner: 10Subramanya Sastry) [20:37:03] PROBLEM - Host cp6001 is DOWN: PING CRITICAL - Packet loss = 100% [20:37:32] Testing a Beta-only change on production is hard, but if you wish to test basic read/edit functionality via the testservers before shipping to production, let me know. [20:38:19] Beta is OK. [20:38:29] (03Merged) 10jenkins-bot: Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341271 (https://phabricator.wikimedia.org/T437916) (owner: 10Subramanya Sastry) [20:38:51] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1341271|Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] [20:38:54] T437916: Parsoid Read Views: Deploy to wikitech - https://phabricator.wikimedia.org/T437916 [20:40:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.21% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:42:57] !log arlolra@deploy1003 ssastry, arlolra: Backport for [[gerrit:1341271|Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:43:33] !log arlolra@deploy1003 ssastry, arlolra: Continuing with deployment [20:44:53] Southparkfan: spiderpig didn't stop at any testservers for that patch, it's sync'ed to the production servers [20:45:42] Ah! Thanks for the info [20:47:56] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2046.codfw.wmnet with OS trixie [20:48:03] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341271|Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (T437916)]] (duration: 09m 11s) [20:48:06] T437916: Parsoid Read Views: Deploy to wikitech - https://phabricator.wikimedia.org/T437916 [20:49:10] (03PS1) 10Jdlrobson: [DONOTMERGE - demoing deployment process only] Refactor vetags to not be a single saveField [extensions/VisualEditor] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342046 (https://phabricator.wikimedia.org/T437736) [20:50:59] !log cdobbins@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp2046.codfw.wmnet [20:51:13] !log cdobbins@puppetserver1001 conftool action : set/weight=1; selector: name=cp2046.codfw.wmnet [20:52:40] brett@cumin2003 provision (PID 2322022) is awaiting input [20:54:31] We're done with the backport window [20:59:35] (03PS1) 10Cklimas: MobileFrontend: Add app icons [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342049 (https://phabricator.wikimedia.org/T434258) [20:59:44] thanks arlolra I'm going to take the reader window shortly [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260915T2100) [21:00:13] (03CR) 10Dzahn: [V:03+1 C:03+2] "https://phabricator.wikimedia.org/T435503#12325210" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [21:00:47] thank you arlolra! [21:03:23] (03PS10) 10CDobbins: prometheus: fix prometheus-ferm-mss.py [puppet] - 10https://gerrit.wikimedia.org/r/1333244 (https://phabricator.wikimedia.org/T433672) [21:04:35] brett@cumin2003 provision (PID 2322022) is awaiting input [21:05:44] (03CR) 10Dzahn: [V:03+1 C:03+2] "Even an ordering change can theoretically break things. My entire goal is to not change hosts that are owned by others while fixing the is" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [21:06:06] (03CR) 10CDobbins: "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1333244 (https://phabricator.wikimedia.org/T433672) (owner: 10CDobbins) [21:06:39] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342049 (https://phabricator.wikimedia.org/T434258) (owner: 10Cklimas) [21:08:05] (03CR) 10Jdlrobson: [C:03+1] MobileFrontend: Add app icons [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342049 (https://phabricator.wikimedia.org/T434258) (owner: 10Cklimas) [21:08:06] (03Merged) 10jenkins-bot: MobileFrontend: Add app icons [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342049 (https://phabricator.wikimedia.org/T434258) (owner: 10Cklimas) [21:08:12] (03CR) 10Dzahn: rsyslog: send php8.5-fpm logs to Logstash [puppet] - 10https://gerrit.wikimedia.org/r/1341351 (https://phabricator.wikimedia.org/T435393) (owner: 10Southparkfan) [21:08:30] !log jdlrobson@deploy1003 Started scap sync-world: Backport for [[gerrit:1342049|MobileFrontend: Add app icons (T434258)]] [21:08:34] T434258: [REQUEST] Add a blue banner to visual editor after publishing or abandoning an app edit - https://phabricator.wikimedia.org/T434258 [21:12:13] (03PS3) 10Dzahn: admin: add anilk to analytics-privatedata-users, remove anil [puppet] - 10https://gerrit.wikimedia.org/r/1339129 (https://phabricator.wikimedia.org/T437611) (owner: 10Ssingh) [21:12:38] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6001.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [21:12:45] (03PS4) 10Dzahn: admin: add anilk to analytics-privatedata-users, remove anil [puppet] - 10https://gerrit.wikimedia.org/r/1339129 (https://phabricator.wikimedia.org/T437611) (owner: 10Ssingh) [21:12:47] !log jdlrobson@deploy1003 jdlrobson, cklimas: Backport for [[gerrit:1342049|MobileFrontend: Add app icons (T434258)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:13:08] (03CR) 10Dzahn: [C:03+1] admin: add anilk to analytics-privatedata-users, remove anil [puppet] - 10https://gerrit.wikimedia.org/r/1339129 (https://phabricator.wikimedia.org/T437611) (owner: 10Ssingh) [21:13:08] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp6001.drmrs.wmnet with OS trixie [21:13:59] (03CR) 10Dzahn: [C:03+1] "I made conflicting changes to this. Let me amend to this to rebase it on top with the same intention to remove the special case." [puppet] - 10https://gerrit.wikimedia.org/r/1322828 (https://phabricator.wikimedia.org/T424266) (owner: 10Scott French) [21:15:44] checked and looks good proceeding with sync [21:15:50] !log jdlrobson@deploy1003 jdlrobson, cklimas: Continuing with deployment [21:17:09] (03PS2) 10Dzahn: zookeeper::server: Use default zookeeper CLASSPATH when using 3.4 [puppet] - 10https://gerrit.wikimedia.org/r/1322828 (https://phabricator.wikimedia.org/T424266) (owner: 10Scott French) [21:17:58] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9438/co" [puppet] - 10https://gerrit.wikimedia.org/r/1333244 (https://phabricator.wikimedia.org/T433672) (owner: 10CDobbins) [21:18:35] RECOVERY - Host cp6001 is UP: PING WARNING - Packet loss = 90%, RTA = 84.87 ms [21:20:18] !log jdlrobson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342049|MobileFrontend: Add app icons (T434258)]] (duration: 11m 47s) [21:20:21] T434258: [REQUEST] Add a blue banner to visual editor after publishing or abandoning an app edit - https://phabricator.wikimedia.org/T434258 [21:22:25] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9439/co" [puppet] - 10https://gerrit.wikimedia.org/r/1333244 (https://phabricator.wikimedia.org/T433672) (owner: 10CDobbins) [21:26:09] 06SRE, 10LDAP-Access-Requests: Request for LDAP NDA Access for CentralNotice Metrics for Yahya - https://phabricator.wikimedia.org/T436294#12325257 (10NAramayo-WMF) @Yahya I'll be working on getting the NDA over to you for signing. Could you please provide me with the below information via email to naramayo@wi... [21:30:04] !log brett@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage [21:31:22] Jdlrobson: I have some some infra stuff to do whenever you're finished, but no rush, still your window :) [21:31:32] hey rzl yep done please feel free [21:31:37] thanks! [21:31:41] (03CR) 10RLazarus: [C:03+2] mcrouter-wancache: Shift traffic from mc-wf1002 to mc-wf1001 [puppet] - 10https://gerrit.wikimedia.org/r/1339855 (https://phabricator.wikimedia.org/T421711) (owner: 10RLazarus) [21:33:19] 06SRE, 10LDAP-Access-Requests: Request for LDAP NDA Access for CentralNotice Metrics for Yahya - https://phabricator.wikimedia.org/T436294#12325297 (10Yahya) Just sent the email. [21:33:23] (03CR) 10Bking: [C:03+1] "Overall LGTM, a couple of suggestions though:" [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [21:34:00] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6001.drmrs.wmnet with reason: host reimage [21:48:25] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:48:32] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:52:45] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [21:52:49] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:53:26] !log rzl@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:53:31] !log rzl@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:53:51] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply [21:53:57] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply [21:55:26] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply [21:55:31] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply [21:56:20] FIRING: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 27.77777777777777 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [21:57:52] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6001.drmrs.wmnet with OS trixie [22:01:01] !log rzl@cumin2003 START - Cookbook sre.hosts.reimage for host mc-wf1002.eqiad.wmnet with OS trixie [22:01:20] RESOLVED: CirrusSearchBackendMemoryIssue: CirrusSearch backend failed 25.55556172842935 times in the last 10 minutes due to memory usage (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchBackendMemoryIssue [22:01:32] !log rzl@cumin2003 START - Cookbook sre.hosts.move-vlan for host mc-wf1002 [22:01:33] (should have said -- finished deploying) [22:02:06] !log rzl@cumin2003 START - Cookbook sre.dns.netbox [22:05:02] (03PS1) 10RLazarus: mcrouter-wancache: Restore mc-wf1002 [puppet] - 10https://gerrit.wikimedia.org/r/1342060 (https://phabricator.wikimedia.org/T421711) [22:07:13] !log rzl@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003" [22:07:17] !log rzl@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc-wf1002 - rzl@cumin2003" [22:07:17] !log rzl@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [22:07:17] !log rzl@cumin2003 START - Cookbook sre.dns.wipe-cache mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:07:20] !log rzl@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-wf1002.eqiad.wmnet 142.48.64.10.in-addr.arpa 2.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:07:21] !log rzl@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host mc-wf1002 [22:07:53] !log rzl@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-wf1002 [22:07:53] !log rzl@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc-wf1002 [22:09:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:14:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:15:46] (03CR) 10Jasmine: [C:03+1] mcrouter-wancache: Restore mc-wf1002 [puppet] - 10https://gerrit.wikimedia.org/r/1342060 (https://phabricator.wikimedia.org/T421711) (owner: 10RLazarus) [22:15:51] (03CR) 10Ryan Kemper: sre.opensearch.roll-restart-reboot: include checklist items (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [22:18:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 19.77% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:19:46] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 499542336 and 38 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [22:23:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:26:36] !log rzl@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage [22:32:52] !log brett@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp6001.* [22:33:19] !log brett@puppetserver1001 conftool action : set/weight=100; selector: name=cp6001.* [22:33:27] !log rzl@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-wf1002.eqiad.wmnet with reason: host reimage [22:39:54] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1331830 (https://phabricator.wikimedia.org/T436398) (owner: 10C. Scott Ananian) [22:41:29] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1331830 (https://phabricator.wikimedia.org/T436398) (owner: 10C. Scott Ananian) [22:42:09] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340216 (https://phabricator.wikimedia.org/T328012) (owner: 10C. Scott Ananian) [22:42:36] (03CR) 10C. Scott Ananian: [C:03+1] Parsoid Read Views: Enable on 60 wikiquote wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341949 (https://phabricator.wikimedia.org/T437917) (owner: 10Subramanya Sastry) [22:42:45] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341949 (https://phabricator.wikimedia.org/T437917) (owner: 10Subramanya Sastry) [22:43:20] !log brett@cumin2003 START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs (T436363) [22:43:23] T436363: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363 [22:43:37] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin depooling P{lvs6001.drmrs.wmnet} and A:liberica [22:43:52] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs6001.drmrs.wmnet} and A:liberica [22:44:07] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin pooling P{lvs6001.drmrs.wmnet} and A:liberica [22:44:34] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs6001.drmrs.wmnet} and A:liberica [22:44:35] (03PS1) 10Samwilson: Thunbor: Switch from CONVERT_PATH to MAGICK_PATH [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342065 (https://phabricator.wikimedia.org/T438089) [22:44:54] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin depooling P{lvs6002.drmrs.wmnet} and A:liberica [22:45:09] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs6002.drmrs.wmnet} and A:liberica [22:45:24] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin pooling P{lvs6002.drmrs.wmnet} and A:liberica [22:45:50] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs6002.drmrs.wmnet} and A:liberica [22:46:00] (03PS2) 10Samwilson: Thunbor: Switch from CONVERT_PATH to MAGICK_PATH [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342065 (https://phabricator.wikimedia.org/T438089) [22:46:09] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin depooling P{lvs6003.drmrs.wmnet} and A:liberica [22:46:25] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs6003.drmrs.wmnet} and A:liberica [22:46:40] !log brett@cumin2003 START - Cookbook sre.loadbalancer.admin pooling P{lvs6003.drmrs.wmnet} and A:liberica [22:46:54] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs6003.drmrs.wmnet} and A:liberica [22:46:57] !log brett@cumin2003 END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs (T436363) [22:47:57] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [22:50:21] !log rzl@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-wf1002.eqiad.wmnet with OS trixie [22:50:37] (03CR) 10RLazarus: [C:03+2] mcrouter-wancache: Restore mc-wf1002 [puppet] - 10https://gerrit.wikimedia.org/r/1342060 (https://phabricator.wikimedia.org/T421711) (owner: 10RLazarus) [22:52:57] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) #page - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [22:53:09] yeehaw [22:57:48] !log sukhe@puppetserver1001 conftool action : set/weight=1; selector: name=cp6001.drmrs.wmnet,service=cdn [22:57:57] RESOLVED: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) #page - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [22:57:57] FIRING: [2x] ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [22:58:56] 10ops-eqiad, 06SRE, 06DC-Ops: Q4: eqiad: (12) PDUs for ML expansion - https://phabricator.wikimedia.org/T400778#12325553 (10wiki_willy) Hi @VRiley-WMF - it looks like the PDUs in E9 thru E14 aren't showing up in Grafana graphs below. After you configure the PDUs in E13 and E14, can you check to see if somet... [23:02:57] RESOLVED: [2x] ProbeDown: Service text-https:443 has failed probes (http_text-https_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [23:05:47] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 2356344 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [23:06:25] !log rzl@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [23:06:33] !log rzl@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [23:07:09] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [23:07:18] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [23:07:22] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [23:07:27] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [23:08:40] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply [23:08:48] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply [23:08:55] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply [23:09:00] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply [23:16:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:40:57] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1342070 [23:40:57] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1342070 (owner: 10TrainBranchBot) [23:50:03] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1342070 (owner: 10TrainBranchBot) [23:57:09] (03CR) 10Subramanya Sastry: Parsoid Read Views: Enable on all namespaces on wikitech (labswiki) (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341271 (https://phabricator.wikimedia.org/T437916) (owner: 10Subramanya Sastry)