[00:03:05] (03PS1) 10RLazarus: build-production-images: Remove buster from the rest of the base images [puppet] - 10https://gerrit.wikimedia.org/r/1342072 (https://phabricator.wikimedia.org/T438101) [00:03:07] (03PS1) 10RLazarus: build-production-images: Remove seed_image field, unused since 2023 [puppet] - 10https://gerrit.wikimedia.org/r/1342073 (https://phabricator.wikimedia.org/T438101) [00:03:38] (03CR) 10RLazarus: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342072 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [00:03:41] (03CR) 10RLazarus: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342073 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [00:06:27] (03PS1) 10RLazarus: Remove seed_image config field, unused since 2023 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342074 (https://phabricator.wikimedia.org/T438101) [00:19:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.35% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:21:46] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [00:27:27] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:42:43] PROBLEM - snapshot of s4 in codfw on backupmon1001 is CRITICAL: Last snapshot for s4 at codfw (db2239) taken on 2026-09-15 23:44:29 is 1179 GiB, but the previous one was 2151 GiB, a change of -45.2 % https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [01:11:03] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1342088 [01:11:03] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1342088 (owner: 10TrainBranchBot) [01:17:38] (03PS1) 10RLazarus: Depool poolcounter[1006,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342089 (https://phabricator.wikimedia.org/T435163) [01:17:38] (03PS1) 10RLazarus: Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) [01:17:39] (03PS1) 10RLazarus: Repool poolcounter[1007,2006] [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342091 (https://phabricator.wikimedia.org/T435163) [01:18:39] (03PS2) 10RLazarus: Depool poolcounter[1007,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342089 (https://phabricator.wikimedia.org/T435163) [01:18:39] (03PS2) 10RLazarus: Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) [01:18:39] (03PS2) 10RLazarus: Repool poolcounter[1007,2006] [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342091 (https://phabricator.wikimedia.org/T435163) [01:18:43] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1342088 (owner: 10TrainBranchBot) [01:20:09] (03PS3) 10RLazarus: Depool poolcounter[1006,2005] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342089 (https://phabricator.wikimedia.org/T435163) [01:20:09] (03PS3) 10RLazarus: Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) [01:20:09] (03PS3) 10RLazarus: Repool poolcounter[1007,2006] [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342091 (https://phabricator.wikimedia.org/T435163) [01:32:33] (03CR) 10Tim Starling: [C:03+1] "Fine as a temporary measure while we wait for my PR https://github.com/debuerreotype/debuerreotype/pull/202" [puppet] - 10https://gerrit.wikimedia.org/r/1341110 (https://phabricator.wikimedia.org/T437829) (owner: 10Muehlenhoff) [01:54:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:58:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.56% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:00:05] Deploy window Automatic deployment of MediaWiki to pretrain wikis - see mw:Pretrain (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T0200) [02:01:06] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:06:25] FIRING: [2x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:08:42] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 36s) [02:12:12] FIRING: [2x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:12:27] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:13:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.8% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:17:12] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:06:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:11:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.38% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:41:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.63% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:56:25] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342102 [03:56:25] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342103 [03:59:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.28% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:59:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [04:04:44] (03CR) 10Giuseppe Lavagetto: [C:03+2] wikikube: enable gVisor [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341884 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [04:14:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [04:15:46] (03Merged) 10jenkins-bot: wikikube: enable gVisor [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341884 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [04:21:46] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [04:22:52] !log oblivian@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [04:22:55] !log oblivian@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [04:24:35] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [04:24:37] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [04:29:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [04:34:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [04:34:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.72% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [04:36:34] Updating Apertium [04:39:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [04:40:26] (03CR) 10KartikMistry: [C:03+2] Update Apertium to 2026-09-15-084320-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341866 (https://phabricator.wikimedia.org/T437213) (owner: 10KartikMistry) [04:42:53] (03Merged) 10jenkins-bot: Update Apertium to 2026-09-15-084320-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341866 (https://phabricator.wikimedia.org/T437213) (owner: 10KartikMistry) [04:45:09] !log kartik@deploy1003 helmfile [staging] START helmfile.d/services/apertium: apply [04:45:32] !log kartik@deploy1003 helmfile [staging] DONE helmfile.d/services/apertium: apply [04:49:50] !log kartik@deploy1003 helmfile [codfw] START helmfile.d/services/apertium: apply [04:50:23] !log kartik@deploy1003 helmfile [codfw] DONE helmfile.d/services/apertium: apply [04:52:39] (03CR) 10Samwilson: [C:03+1] Set JPGs to be thumbnailed via VIPS if md5 starts with a (032 comments) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [04:54:21] !log kartik@deploy1003 helmfile [eqiad] START helmfile.d/services/apertium: apply [04:54:55] !log kartik@deploy1003 helmfile [eqiad] DONE helmfile.d/services/apertium: apply [04:56:15] !log Updated Apertium to 2026-09-15-084320-production (T437213) [04:56:18] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [04:56:19] T437213: Review and update LPL service base images – 2026Q3 - https://phabricator.wikimedia.org/T437213 [05:11:26] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Remove obsolete buildkitd image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339620 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [05:12:30] !log pruned obsolete Bullseye image buildkitd from the docker registry T416452 [05:12:33] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:12:34] T416452: Migrate Docker images running in Production away from Bullseye - https://phabricator.wikimedia.org/T416452 [05:12:54] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] "buildkitd has been pruned from the Docker registry" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339620 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [05:14:01] (03PS2) 10Muehlenhoff: amd::gpu: Remove support for bullseye [puppet] - 10https://gerrit.wikimedia.org/r/1337885 [05:23:36] (03CR) 10Muehlenhoff: [C:03+2] amd::gpu: Remove support for bullseye [puppet] - 10https://gerrit.wikimedia.org/r/1337885 (owner: 10Muehlenhoff) [05:26:51] (03PS6) 10Muehlenhoff: Move DB backups from cumin1003 to cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1339778 (https://phabricator.wikimedia.org/T427897) [05:27:55] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1339778 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [05:35:28] (03CR) 10Muehlenhoff: "Ready for review now (the breakage in PCC was due to an related prep change for Puppet 8)" [puppet] - 10https://gerrit.wikimedia.org/r/1339778 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [05:41:27] 10ops-codfw, 06DC-Ops: Unresponsive management for ms-be2094.mgmt:22 - https://phabricator.wikimedia.org/T438111 (10phaultfinder) 03NEW [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T0600) [06:04:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:06:40] FIRING: [2x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:14:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:15:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:17:12] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:30:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.07% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:32:57] (03CR) 10Giuseppe Lavagetto: [C:03+2] shellbox-timeline: enable gVisor everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341600 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [06:35:35] (03PS1) 10Matthias Mullie: Adds an instrument for pre-image-carousel-retest [extensions/WikimediaEvents] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342112 (https://phabricator.wikimedia.org/T437076) [06:35:43] (03Merged) 10jenkins-bot: shellbox-timeline: enable gVisor everywhere [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341600 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [06:35:55] (03PS1) 10Matthias Mullie: Adds an instrument for pre-image-carousel-retest [extensions/WikimediaEvents] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342113 (https://phabricator.wikimedia.org/T437076) [06:36:12] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/WikimediaEvents] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342112 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [06:36:23] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/WikimediaEvents] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342113 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [06:37:08] (03PS1) 10Slyngshede: site.pp move cp5025 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342114 (https://phabricator.wikimedia.org/T436363) [06:37:27] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply [06:38:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.42% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:40:21] (03PS1) 10Slyngshede: site.pp move cp5026 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342115 (https://phabricator.wikimedia.org/T436363) [06:43:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:43:24] (03PS1) 10Matthias Mullie: Set up instrument for 5-arm test [extensions/MultimediaViewer] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1342116 (https://phabricator.wikimedia.org/T437076) [06:44:56] (03Abandoned) 10Matthias Mullie: Set up instrument for 5-arm test [extensions/MultimediaViewer] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1342116 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [06:45:10] (03PS1) 10Matthias Mullie: Set up instrument for 5-arm test [extensions/MultimediaViewer] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342117 (https://phabricator.wikimedia.org/T437076) [06:45:29] (03PS1) 10Matthias Mullie: Set up instrument for 5-arm test [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342118 (https://phabricator.wikimedia.org/T437076) [06:47:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.79% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:47:55] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply [06:50:10] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342117 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [06:50:18] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342118 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [06:50:25] !log installing sudo security updates [06:50:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:56:21] (03PS1) 10Muehlenhoff: Stop tracking buster OS deprecations [puppet] - 10https://gerrit.wikimedia.org/r/1342123 [06:57:46] (03PS1) 10Giuseppe Lavagetto: gvisor: do not mark as enabled on bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1342124 [06:59:07] (03CR) 10Hashar: "That is not an issue with Zuul, it is an issue with Zookeeper on Trixie (which ended up affecting the Zuul infra) and any Zookeeper on Tri" [puppet] - 10https://gerrit.wikimedia.org/r/1327569 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [07:00:05] Amir1, urbanecm, and awight: OwO what's this, a deployment window?? UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T0700). nyaa~ [07:00:05] matthiasmullie: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:08] o/ [07:00:20] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (NOOP 1 CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1342124 (owner: 10Giuseppe Lavagetto) [07:01:29] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12326132 (10SLyngshede-WMF) [07:03:34] (03PS1) 10Muehlenhoff: Remove the seed_image from the docker build config files [puppet] - 10https://gerrit.wikimedia.org/r/1342126 [07:04:10] (03CR) 10Giuseppe Lavagetto: [V:03+1 C:03+2] gvisor: do not mark as enabled on bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1342124 (owner: 10Giuseppe Lavagetto) [07:04:44] (03PS1) 10Slyngshede: site.pp: move cp6002 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342127 (https://phabricator.wikimedia.org/T436363) [07:04:48] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mlitn@deploy1003 using scap backport" [extensions/WikimediaEvents] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342112 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [07:04:48] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mlitn@deploy1003 using scap backport" [extensions/WikimediaEvents] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342113 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [07:04:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mlitn@deploy1003 using scap backport" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342117 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [07:04:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mlitn@deploy1003 using scap backport" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342118 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [07:06:22] (03Merged) 10jenkins-bot: Adds an instrument for pre-image-carousel-retest [extensions/WikimediaEvents] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342112 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [07:06:37] (03Merged) 10jenkins-bot: Adds an instrument for pre-image-carousel-retest [extensions/WikimediaEvents] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342113 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [07:06:39] (03Merged) 10jenkins-bot: Set up instrument for 5-arm test [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342118 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [07:06:42] (03CR) 10JMeybohm: "I would suggest to also absent gvisor in `modules/profile/manifests/containerd.pp` for consistency" [puppet] - 10https://gerrit.wikimedia.org/r/1342124 (owner: 10Giuseppe Lavagetto) [07:07:35] (03Merged) 10jenkins-bot: Set up instrument for 5-arm test [extensions/MultimediaViewer] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342117 (https://phabricator.wikimedia.org/T437076) (owner: 10Matthias Mullie) [07:08:18] (03PS1) 10Muehlenhoff: Remove buster from base build config [puppet] - 10https://gerrit.wikimedia.org/r/1342128 [07:09:52] !log mlitn@deploy1003 Started scap sync-world: Backport for [[gerrit:1342112|Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113|Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117|Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118|Set up instrument for 5-arm test (T437076)]] [07:09:55] T437076: Predicting power for five-arm test - https://phabricator.wikimedia.org/T437076 [07:13:46] (03Abandoned) 10Dpogorzelski: liftwing-studio: OIDC auth against idp.wikimedia.org [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341909 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [07:15:39] !log mlitn@deploy1003 mlitn: Backport for [[gerrit:1342112|Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113|Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117|Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118|Set up instrument for 5-arm test (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri [07:15:39] fied there. [07:15:42] T437076: Predicting power for five-arm test - https://phabricator.wikimedia.org/T437076 [07:16:11] !log mlitn@deploy1003 mlitn: Continuing with deployment [07:19:57] (03CR) 10Muehlenhoff: [C:03+2] Stop tracking buster OS deprecations [puppet] - 10https://gerrit.wikimedia.org/r/1342123 (owner: 10Muehlenhoff) [07:20:48] !log mlitn@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342112|Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342113|Adds an instrument for pre-image-carousel-retest (T437076)]], [[gerrit:1342117|Set up instrument for 5-arm test (T437076)]], [[gerrit:1342118|Set up instrument for 5-arm test (T437076)]] (duration: 10m 56s) [07:20:52] T437076: Predicting power for five-arm test - https://phabricator.wikimedia.org/T437076 [07:21:00] (03CR) 10Daniel Kertesz: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1342114 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [07:21:27] Done. I believe that concludes this window's deployments. [07:22:15] (03CR) 10Daniel Kertesz: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1342115 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [07:23:27] (03PS1) 10Muehlenhoff: os-reports: Also remove buster from the config [puppet] - 10https://gerrit.wikimedia.org/r/1342132 [07:29:36] (03CR) 10Elukey: [C:03+1] Remove buster from base build config [puppet] - 10https://gerrit.wikimedia.org/r/1342128 (owner: 10Muehlenhoff) [07:29:49] (03PS1) 10Dpogorzelski: idp: add the liftwing_studio service [puppet] - 10https://gerrit.wikimedia.org/r/1342135 (https://phabricator.wikimedia.org/T437706) [07:32:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:34:19] !log oblivian@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply [07:34:51] !log oblivian@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply [07:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.35% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:37:53] (03CR) 10Elukey: [C:03+1] os-reports: Also remove buster from the config [puppet] - 10https://gerrit.wikimedia.org/r/1342132 (owner: 10Muehlenhoff) [07:38:41] (03CR) 10Elukey: [C:03+1] Remove the seed_image from the docker build config files [puppet] - 10https://gerrit.wikimedia.org/r/1342126 (owner: 10Muehlenhoff) [07:39:54] (03CR) 10Muehlenhoff: [C:03+2] os-reports: Also remove buster from the config [puppet] - 10https://gerrit.wikimedia.org/r/1342132 (owner: 10Muehlenhoff) [07:41:55] (03CR) 10Fabfur: [C:03+1] site.pp move cp5025 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342114 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [07:43:02] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.13 point update - https://phabricator.wikimedia.org/T414205#12326201 (10MoritzMuehlenhoff) [07:43:31] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.13 point update - https://phabricator.wikimedia.org/T414205#12326202 (10MoritzMuehlenhoff) 05Open→03Resolved a:03MoritzMuehlenhoff All done! [07:44:01] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.14 point update - https://phabricator.wikimedia.org/T426759#12326205 (10MoritzMuehlenhoff) [07:47:21] (03CR) 10CWilliams: mediabackups: Update versitygw systemd unit file to SIGHUP on reload (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [07:48:18] (03PS2) 10CWilliams: admin: Updated aliases for cwilliams [puppet] - 10https://gerrit.wikimedia.org/r/1341871 [07:48:27] 06SRE, 06Infrastructure-Foundations, 07Epic, 07Kubernetes: aux-k8s: eqiad expansion, codfw creation, & future hopes and dreams - https://phabricator.wikimedia.org/T378742#12326242 (10LSobanski) [07:52:05] (03CR) 10CWilliams: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [07:52:56] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for recommendation-api,sessionstore,tegola,termbox [puppet] - 10https://gerrit.wikimedia.org/r/1341131 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [07:54:05] (03CR) 10CWilliams: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1339778 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [07:54:13] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw [07:58:32] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [07:59:30] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [07:59:30] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw [07:59:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [08:00:06] jnuche and dduvall: MediaWiki train - Utc-0+Utc-7 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T0800). Please do the needful. [08:00:15] morning, the train will roll out in a few minutes [08:00:46] (03PS2) 10Jelto: service::catalog: Set ipip for recommendation-api,sessionstore,tegola,termbox [puppet] - 10https://gerrit.wikimedia.org/r/1341132 (https://phabricator.wikimedia.org/T420436) [08:01:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:01:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.35% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:03:01] 06SRE, 06Infrastructure-Foundations, 07Epic, 07Kubernetes: aux-k8s: eqiad expansion, codfw creation, & future hopes and dreams - https://phabricator.wikimedia.org/T378742#12326316 (10LSobanski) [08:04:23] (03CR) 10CWilliams: [C:03+2] admin: Updated aliases for cwilliams [puppet] - 10https://gerrit.wikimedia.org/r/1341871 (owner: 10CWilliams) [08:04:28] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342177 (https://phabricator.wikimedia.org/T430839) [08:04:31] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342177 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [08:04:53] 10ops-eqiad, 06SRE, 06DC-Ops, 10Kafka-Infrastructure, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12326337 (10brouberol) Let's wait a few weeks, no problem! [08:05:31] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342177 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [08:05:38] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for recommendation-api,sessionstore,tegola,termbox [puppet] - 10https://gerrit.wikimedia.org/r/1341132 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [08:07:21] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [08:07:46] (03CR) 10Ayounsi: [C:03+2] move BGPalerter to netmon hosts [puppet] - 10https://gerrit.wikimedia.org/r/1341845 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [08:08:57] (03PS1) 10Muehlenhoff: Record LDAP acccess for elliottetzkorn [puppet] - 10https://gerrit.wikimedia.org/r/1342180 [08:10:03] PROBLEM - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [08:10:36] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [08:11:28] !log jnuche@deploy1003 Rolling back deployment [08:11:43] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [08:11:43] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad [08:11:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.62% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:14:04] !log jnuche@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs T430839 [08:14:07] T430839: 1.47.0-wmf.20 deployment blockers - https://phabricator.wikimedia.org/T430839 [08:14:39] (03CR) 10Muehlenhoff: [C:03+2] Record LDAP acccess for elliottetzkorn [puppet] - 10https://gerrit.wikimedia.org/r/1342180 (owner: 10Muehlenhoff) [08:14:52] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342183 (https://phabricator.wikimedia.org/T430839) [08:14:55] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342183 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [08:14:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [08:15:56] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342183 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [08:21:46] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [08:22:08] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:22:37] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:22:48] 06SRE, 06Infrastructure-Foundations: Upgrade rpki hosts to Trixie - https://phabricator.wikimedia.org/T438122 (10MoritzMuehlenhoff) 03NEW [08:23:47] (03PS1) 10Jelto: service::catalog: Set ipip for thumbor toolhub wikifeeds zotero codfw [puppet] - 10https://gerrit.wikimedia.org/r/1342185 (https://phabricator.wikimedia.org/T420436) [08:23:50] (03PS1) 10Jelto: service::catalog: Set ipip for thumbor toolhub wikifeeds zotero eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1342186 (https://phabricator.wikimedia.org/T420436) [08:24:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.94% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:24:49] !log jnuche@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs T430839 [08:24:53] T430839: 1.47.0-wmf.20 deployment blockers - https://phabricator.wikimedia.org/T430839 [08:29:12] (03PS2) 10Jelto: service::catalog: Set ipip for thumbor toolhub wikifeeds zotero codfw [puppet] - 10https://gerrit.wikimedia.org/r/1342185 (https://phabricator.wikimedia.org/T420436) [08:29:40] (03PS1) 10Cathal Mooney: Re-enable ulsfo path over HE [homer/public] - 10https://gerrit.wikimedia.org/r/1342187 (https://phabricator.wikimedia.org/T435543) [08:31:19] (03CR) 10Ayounsi: [C:03+1] Re-enable ulsfo path over HE [homer/public] - 10https://gerrit.wikimedia.org/r/1342187 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [08:31:30] (03CR) 10Cathal Mooney: [C:03+2] Re-enable ulsfo path over HE [homer/public] - 10https://gerrit.wikimedia.org/r/1342187 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [08:31:50] (03PS2) 10Jelto: cache-text: set caching for etherpad-next "websockets" [puppet] - 10https://gerrit.wikimedia.org/r/1341894 (https://phabricator.wikimedia.org/T435509) [08:32:13] train blocked at T438125 [08:32:14] T438125: Error: Interface "MediaWiki\Extension\VisualEditor\VisualEditorRegisterChangeTagsHook" not found - https://phabricator.wikimedia.org/T438125 [08:32:33] (03CR) 10Jelto: cache-text: set caching for etherpad-next "websockets" (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341894 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [08:32:41] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, and 2 others: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12326440 (10fgiunchedi) I did a quick audit on the fleet (see below), my understanding is that we want to do the following... [08:32:57] (03Merged) 10jenkins-bot: Re-enable ulsfo path over HE [homer/public] - 10https://gerrit.wikimedia.org/r/1342187 (https://phabricator.wikimedia.org/T435543) (owner: 10Cathal Mooney) [08:33:58] (03CR) 10Brouberol: [C:03+1] liftwing-studio: add external-services egress to the chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341903 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:34:31] (03CR) 10Brouberol: [C:03+1] liftwing-studio: OIDC auth against idp.wikimedia.org [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341904 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:34:33] !log start depooling eqsin (T438052) [08:34:36] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:34:36] (03CR) 10Brouberol: [C:03+1] idp: add the liftwing_studio service [puppet] - 10https://gerrit.wikimedia.org/r/1342135 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:34:36] T438052: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052 [08:34:59] (03CR) 10Dpogorzelski: [C:03+2] liftwing-studio: add external-services egress to the chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341903 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:35:02] !log slyngshede@cumin1003 conftool action : set/pooled=no; selector: cluster=dnsbox,dc=eqsin [08:35:14] !log slyngshede@cumin1003 START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: no reason specified, no task ID specified] [08:35:38] !log slyngshede@cumin1003 END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool eqsin [reason: no reason specified, no task ID specified] [08:36:15] !log slyngshede@cumin1003 START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: depooling for maintainance, T438052] [08:36:24] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: depooling for maintainance, T438052] [08:37:23] (03Merged) 10jenkins-bot: liftwing-studio: add external-services egress to the chart [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341903 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:38:22] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [08:40:02] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:40:07] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply [08:43:22] RESOLVED: CertAlmostExpired: gNMI TLS certificate for lsw1-c2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [08:43:51] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12326457 (10SLyngshede-WMF) Traffic is moving to ULSFO: https://grafana.wikimedia.org/d/000000093/cdn-frontend-network?from=now-30m&to=now&timezone... [08:44:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.79% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:45:53] (03CR) 10CWilliams: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [08:47:59] (03PS1) 10Hashar: specs: change explicit test_on to defaults [puppet] - 10https://gerrit.wikimedia.org/r/1342189 (https://phabricator.wikimedia.org/T435917) [08:48:01] (03PS1) 10Hashar: jenkins: run specs against all default Debians [puppet] - 10https://gerrit.wikimedia.org/r/1342190 (https://phabricator.wikimedia.org/T435917) [08:48:15] (03CR) 10Klausman: [C:03+1] liftwing-studio: OIDC auth against idp.wikimedia.org [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341904 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:50:02] (03CR) 10CWilliams: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [08:50:52] (03CR) 10Hashar: "That is a follow up to I605a7d6450bde982831cd67722b24b41159f1562 . I think whenever we update `test_on` to include Debian 14, we should ha" [puppet] - 10https://gerrit.wikimedia.org/r/1342189 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [08:52:52] (03CR) 10Hashar: "When we relocated the Jenkins instances to Trixie (13) hosts, I forgot to update `test_on` to match reality. I think it is better to foll" [puppet] - 10https://gerrit.wikimedia.org/r/1342190 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [08:53:33] (03PS1) 10Ayounsi: BGPalerter: fix RPKI config [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) [08:54:09] (03CR) 10Muehlenhoff: [C:03+2] Move DB backups from cumin1003 to cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1339778 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [08:54:20] (03CR) 10Ayounsi: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [08:54:41] !log btullis@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [08:55:14] !log btullis@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [08:57:30] (03CR) 10Dpogorzelski: [C:03+2] liftwing-studio: OIDC auth against idp.wikimedia.org [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341904 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [08:58:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.07% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [08:58:48] (03CR) 10Arnaudb: [C:03+1] "thanks for the rebase! lgtm!" [puppet] - 10https://gerrit.wikimedia.org/r/1341894 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [09:00:02] (03Merged) 10jenkins-bot: liftwing-studio: OIDC auth against idp.wikimedia.org [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341904 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [09:00:03] RECOVERY - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [09:03:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:06:06] (03PS2) 10Ayounsi: BGPalerter: fix RPKI config [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) [09:06:50] (03CR) 10Ayounsi: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [09:07:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:09:22] !log elukey@deploy1003 helmfile [eqiad] START helmfile.d/admin 'sync'. [09:09:24] !log elukey@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'sync'. [09:10:17] !log elukey@deploy1003 helmfile [codfw] START helmfile.d/admin 'sync'. [09:10:21] !log elukey@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'sync'. [09:11:35] (03PS3) 10Ayounsi: BGPalerter: fix RPKI config [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) [09:11:52] !log elukey@deploy1003 helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'. [09:11:58] !log elukey@deploy1003 helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'. [09:12:05] (03CR) 10Ayounsi: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [09:12:11] (03CR) 10DCausse: [C:03+2] search: add semantic-highlighter-staging to liftwing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341239 (https://phabricator.wikimedia.org/T433872) (owner: 10DCausse) [09:12:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.87% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:12:47] !log elukey@deploy1003 helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'. [09:12:50] !log elukey@deploy1003 helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'. [09:13:56] FIRING: TransportLinksInUsageNoRedundancy: ulsfo inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=ulsfo - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [09:14:39] (03Merged) 10jenkins-bot: search: add semantic-highlighter-staging to liftwing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341239 (https://phabricator.wikimedia.org/T433872) (owner: 10DCausse) [09:15:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:15:28] (03PS4) 10Ayounsi: BGPalerter: fix RPKI config [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) [09:15:36] (03CR) 10Ayounsi: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [09:15:47] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [09:18:14] !log slyngshede@cumin1003 conftool action : set/pooled=no; selector: name=cp3074.esams.wmnet [09:18:28] !log slyngshede@cumin1003 conftool action : set/pooled=yes; selector: name=cp3074.esams.wmnet [09:19:37] !log btullis@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'. [09:19:41] !log slyngshede@cumin1003 conftool action : set/pooled=no; selector: name=cp5025.eqsin.wmnet [09:19:53] !log btullis@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'. [09:20:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [09:22:56] (03CR) 10Slyngshede: [C:03+2] site.pp move cp5025 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342114 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [09:26:27] (03CR) 10Muehlenhoff: [C:03+1] "Looks good!" [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [09:26:49] (03CR) 10Ayounsi: [C:03+2] BGPalerter: fix RPKI config [puppet] - 10https://gerrit.wikimedia.org/r/1342191 (https://phabricator.wikimedia.org/T437024) (owner: 10Ayounsi) [09:29:01] !log slyngshede@cumin1003 START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [09:29:19] PROBLEM - Host mr1-eqsin is DOWN: CRITICAL - Time to live exceeded (103.102.166.128) [09:29:25] PROBLEM - Host mr1-eqsin IPv6 is DOWN: CRITICAL - Time to live exceeded (2001:df2:e500:ffff::1) [09:29:35] 06SRE, 10ServiceOps-Upgrades-Hardware, 07Essential-Work, 06ServiceOps (Next quarter): mc20[56-73] implementation tracking - https://phabricator.wikimedia.org/T436273#12326628 (10MLechvien-WMF) [09:29:42] (03CR) 10Marostegui: [C:03+1] mediabackups: Update versitygw systemd unit file to SIGHUP on reload (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [09:29:47] PROBLEM - Host ps1-603-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [09:30:01] PROBLEM - Host ps1-604-eqsin is DOWN: PING CRITICAL - Packet loss = 100% [09:30:11] PROBLEM - Host mr1-eqsin.oob IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [09:30:37] (03CR) 10Marostegui: "We have to double check the next iteration of backups look good! Thank you" [puppet] - 10https://gerrit.wikimedia.org/r/1339778 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [09:30:39] !log slyngshede@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [09:30:43] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1342189 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [09:33:57] FIRING: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:34:22] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [09:35:57] FIRING: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:36:35] (03CR) 10Marostegui: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [09:36:51] FIRING: SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://wikitech.wikimedia.org/wiki/Runbook#https://wikifeeds.svc.eqiad.wmnet:4101 - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [09:38:15] !log imported routinator 0.15.2-1trixie to thirdparty/routinator T438122 [09:38:18] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:38:19] T438122: Upgrade rpki hosts to Trixie - https://phabricator.wikimedia.org/T438122 [09:38:37] (03PS1) 10Hasan Akgün (WMDE): wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342195 [09:40:57] RESOLVED: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:41:32] !log oblivian@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply [09:41:43] (03CR) 10Arthur taylor: wikidata-query-builder: bump image version (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342195 (owner: 10Hasan Akgün (WMDE)) [09:41:57] !log oblivian@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply [09:42:12] FIRING: [2x] JobUnavailable: Reduced availability for job pdu_sentry4 in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:42:34] (03CR) 10CWilliams: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [09:42:47] RECOVERY - Host mr1-eqsin is UP: PING OK - Packet loss = 0%, RTA = 222.61 ms [09:42:49] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18), 07Essential-Work: Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12326738 (10Gehel) [09:42:51] RECOVERY - Host ps1-603-eqsin is UP: PING OK - Packet loss = 0%, RTA = 211.88 ms [09:42:51] RECOVERY - Host ps1-604-eqsin is UP: PING OK - Packet loss = 0%, RTA = 212.48 ms [09:43:13] 07sre-alert-triage, 06Data-Platform-SRE (2026-08-28 - 2026-09-18), 07Essential-Work: Alert in need of triage: AlertLintProblem (instance localhost:9123) - https://phabricator.wikimedia.org/T430139#12326743 (10Gehel) [09:43:57] RESOLVED: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:44:14] (03CR) 10Marostegui: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [09:44:22] RESOLVED: [2x] CertAlmostExpired: gNMI TLS certificate for asw1-603-eqsin.mgmt.eqsin.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [09:44:32] 06SRE, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-28 - 2026-09-18), 07Essential-Work: install1005 running out of disk due to squid log volume from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12326771 (10Gehel) [09:44:33] PROBLEM - HAProxy HTTPS measure-eqiad.wikimedia.org ECDSA on cp5025 is CRITICAL: SSL CRITICAL - Certificate *.wikipedia.org SAN measure-eqiad.wikimedia.org not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, *.m.wikidata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m.wikiquote.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, *.m.wiktionary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wik [09:44:33] rg, *.wikidata.org, *.wikifunctions.org, *.wikimedia.org, *.wikimediafoundation.org, *.wikinews.org, *.wikipedia.org, *.wikiquote.org, *.wikisource.org, *.wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfusercontent.org, mediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, wikimedia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquote.org, wikisource.org, wikiversity.org, wikivoyage.org, wiktiona [09:44:33] wmfusercontent.org:Certificate *.wikipedia.org SAN measure-codfw.wikimedia.org not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, *.m.wikidata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m.wikiquote.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, *.m.wiktionary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wikibooks.org, *.wikidata.org, *.wikifunctions.org, *.wikimedia.org, *.wikimedia [09:44:33] on.org, *.wikinews.org, *.wikipedia.org, *.wikiquote.org, *.wikisource.org, *.wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfusercontent.org, mediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, wikimedia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquote.org, wikisource.org, wikiversity.org, wikivoyage.org, wiktionary.org, wmfusercontent.org:Certificate *.wikipedia.org SAN measure-esams.wiki [09:44:33] g not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, *.m.wikidata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m.wikiquote.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, *.m.wiktionary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wikibooks.org, *.wikidata.org, *.wikifunctions.org, *.wikimedia.org, *.wikimediafoundation.org, *.wikinews.org, *.wikipedia.org, *.wikiquote.org, *.wikisourc [09:44:34] .wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfusercontent.org, mediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, wikimedia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquote.org, wikisource.org, wikiversity.org, wikivoyage.org, wiktionary.org, wmfusercontent.org:Certificate *.wikipedia.org SAN measure-ulsfo.wikimedia.org not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, * [09:44:34] ata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m.wikiquote.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, *.m.wiktionary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wikibooks.org, *.wikidata.org, *.wikifunctions.org, *.wikimedia.org, *.wikimediafoundation.org, *.wikinews.org, *.wikipedia.org, *.wikiquote.org, *.wikisource.org, *.wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfuserconten [09:44:35] ediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, wikimedia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquote.org, wikisource.org, wikiversity.org, wikivoyage.org, wiktionary.org, wmfusercontent.org:Certificate *.wikipedia.org SAN measure-eqsin.wikimedia.org not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, *.m.wikidata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m. [09:44:35] e.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, *.m.wiktionary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wikibooks.org, *.wikidata.org, *.wikifunctions.org, *.wikimedia.org, *.wikimediafoundation.org, *.wikinews.org, *.wikipedia.org, *.wikiquote.org, *.wikisource.org, *.wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfusercontent.org, mediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, [09:44:36] ia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquote.org, wikisource.org, wikiversity.org, wikivoyage.org, wiktionary.org, wmfusercontent.org:Certificate *.wikipedia.org SAN measure-drmrs.wikimedia.org not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, *.m.wikidata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m.wikiquote.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, * [09:44:36] onary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wikibooks.org, *.wikidata.org, *.wikifunctions.org, *.wikimedia.org, *.wikimediafoundation.org, *.wikinews.org, *.wikipedia.org, *.wikiquote.org, *.wikisource.org, *.wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfusercontent.org, mediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, wikimedia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquo [09:44:37] wikisource.org, wikiversity.org, wikivoyage.org, wiktionary.org, wmfusercontent.org:Certificate *.wikipedia.org SAN measure-magru.wikimedia.org not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, *.m.wikidata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m.wikiquote.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, *.m.wiktionary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wikibooks.org, [09:44:37] ta.org, *.wikifunctions.org, *.wikimedia.org, *.wikimediafoundation.org, *.wikinews.org, *.wikipedia.org, *.wikiquote.org, *.wikisource.org, *.wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfusercontent.org, mediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, wikimedia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquote.org, wikisource.org, wikiversity.org, wikivoyage.org, wiktionary.org, wmfu [09:44:38] nt.org https://wikitech.wikimedia.org/wiki/HTTPS [09:44:38] PROBLEM - Confd vcl based reload on cp5025 is CRITICAL: reload-vcl failed to run since 0h, 1 minutes. https://wikitech.wikimedia.org/wiki/Varnish [09:44:42] 10ops-eqiad, 06SRE, 06DC-Ops, 10Kafka-Infrastructure, and 2 others: Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12326773 (10Gehel) [09:44:49] RECOVERY - Host mr1-eqsin IPv6 is UP: PING OK - Packet loss = 0%, RTA = 211.12 ms [09:44:59] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18), 07Essential-Work: Q1:rack/setup/install cirrussearch11[26-30] - https://phabricator.wikimedia.org/T436285#12326776 (10Gehel) [09:45:11] 10ops-codfw, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18), 07Essential-Work: Q1:rack/setup/install cirrussearch21[16-20] - https://phabricator.wikimedia.org/T436287#12326777 (10Gehel) [09:45:29] RECOVERY - Host mr1-eqsin.oob IPv6 is UP: PING OK - Packet loss = 0%, RTA = 245.52 ms [09:45:31] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18), 07Essential-Work: Q1:rack/setup/install dse-k8s-worker10[29-38] - https://phabricator.wikimedia.org/T436964#12326781 (10Gehel) [09:45:31] PROBLEM - HAProxy HTTPS upload.wikimedia.org ECDSA on cp5025 is CRITICAL: SSL CRITICAL - Certificate *.wikipedia.org SAN upload.wikimedia.org not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, *.m.wikidata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m.wikiquote.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, *.m.wiktionary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wikibooks.org, *. [09:45:31] .org, *.wikifunctions.org, *.wikimedia.org, *.wikimediafoundation.org, *.wikinews.org, *.wikipedia.org, *.wikiquote.org, *.wikisource.org, *.wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfusercontent.org, mediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, wikimedia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquote.org, wikisource.org, wikiversity.org, wikivoyage.org, wiktionary.org, wmfuse [09:45:31] .org:Certificate *.wikipedia.org SAN maps.wikimedia.org not found in cert SAN list: *.m.mediawiki.org, *.m.wikibooks.org, *.m.wikidata.org, *.m.wikimedia.org, *.m.wikinews.org, *.m.wikipedia.org, *.m.wikiquote.org, *.m.wikisource.org, *.m.wikiversity.org, *.m.wikivoyage.org, *.m.wiktionary.org, *.mediawiki.org, *.planet.wikimedia.org, *.wikibooks.org, *.wikidata.org, *.wikifunctions.org, *.wikimedia.org, *.wikimediafoundation.org, *.wikin [09:45:31] *.wikipedia.org, *.wikiquote.org, *.wikisource.org, *.wikiversity.org, *.wikivoyage.org, *.wiktionary.org, *.wmfusercontent.org, mediawiki.org, w.wiki, wikibooks.org, wikidata.org, wikifunctions.org, wikimedia.org, wikimediafoundation.org, wikinews.org, wikipedia.org, wikiquote.org, wikisource.org, wikiversity.org, wikivoyage.org, wiktionary.org, wmfusercontent.org https://wikitech.wikimedia.org/wiki/HTTPS [09:45:53] (03PS7) 10CWilliams: mediabackups: Update versitygw systemd unit file to SIGHUP on reload [puppet] - 10https://gerrit.wikimedia.org/r/1321123 (https://phabricator.wikimedia.org/T430023) (owner: 10Jcrespo) [09:46:50] 06SRE, 06Data-Engineering, 10Kafka-Infrastructure, 06serviceops-radar, and 4 others: Configuration Management for Kafka settings - https://phabricator.wikimedia.org/T276088#12326792 (10Gehel) [09:47:09] RESOLVED: SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://wikitech.wikimedia.org/wiki/Runbook#https://wikifeeds.svc.eqiad.wmnet:4101 - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [09:47:16] FIRING: [2x] JobUnavailable: Reduced availability for job pdu_sentry4 in ops@eqsin - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:48:16] (03PS2) 10Hasan Akgün (WMDE): wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342195 [09:48:21] !log slyngshede@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5025.eqsin.wmnet with reason: reimaging [09:48:35] (03CR) 10Hasan Akgün (WMDE): wikidata-query-builder: bump image version (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342195 (owner: 10Hasan Akgün (WMDE)) [09:48:43] (03CR) 10Arthur taylor: [C:03+1] "Looks good to me - can be deployed" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342195 (owner: 10Hasan Akgün (WMDE)) [09:48:47] (03PS1) 10Marostegui: Revert "dbbackups: Pool db2250:s5 to replace db2201:s5 for backups" [puppet] - 10https://gerrit.wikimedia.org/r/1342197 [09:48:55] (03PS2) 10Marostegui: Revert "dbbackups: Pool db2250:s5 to replace db2201:s5 for backups" [puppet] - 10https://gerrit.wikimedia.org/r/1342197 [09:49:03] (03PS1) 10Btullis: admin_ng: add analytics namespaces as storage tenants [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342198 [09:49:16] !log slyngshede@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on cp5025.eqsin.wmnet with reason: reimaging [09:49:17] (03CR) 10Marostegui: "@cwilliams@wikimedia.org I am reading to repool db2201:s5, this is a direct revert from Jaime's commit, so it should be fine, but if you h" [puppet] - 10https://gerrit.wikimedia.org/r/1342197 (owner: 10Marostegui) [09:49:21] (03PS2) 10Btullis: admin_ng: add analytics namespaces as storage tenants [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342198 (https://phabricator.wikimedia.org/T437447) [09:49:47] (03PS1) 10Gkyziridis: ml-services: Deployment of Qwen3.8-27B-FP8 model under llm namespace. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342199 (https://phabricator.wikimedia.org/T436644) [09:54:11] (03PS1) 10DCausse: search: semantic-highlighting fix model name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342200 (https://phabricator.wikimedia.org/T433872) [09:54:54] (03PS5) 10Ayounsi: Add Arelion Prometheus integration [puppet] - 10https://gerrit.wikimedia.org/r/1341098 (https://phabricator.wikimedia.org/T311005) [09:54:56] (03CR) 10Ozge: [C:03+1] search: semantic-highlighting fix model name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342200 (https://phabricator.wikimedia.org/T433872) (owner: 10DCausse) [09:55:10] (03CR) 10DCausse: [C:03+2] search: semantic-highlighting fix model name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342200 (https://phabricator.wikimedia.org/T433872) (owner: 10DCausse) [09:55:34] (03PS6) 10Ayounsi: Add Arelion Prometheus integration [puppet] - 10https://gerrit.wikimedia.org/r/1341098 (https://phabricator.wikimedia.org/T311005) [09:55:43] (03CR) 10Ayounsi: Add Arelion Prometheus integration (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341098 (https://phabricator.wikimedia.org/T311005) (owner: 10Ayounsi) [09:57:43] (03Merged) 10jenkins-bot: search: semantic-highlighting fix model name [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342200 (https://phabricator.wikimedia.org/T433872) (owner: 10DCausse) [09:58:43] (03CR) 10CI reject: [V:04-1] Add Arelion Prometheus integration [puppet] - 10https://gerrit.wikimedia.org/r/1341098 (https://phabricator.wikimedia.org/T311005) (owner: 10Ayounsi) [09:58:53] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1000) [10:00:33] (03PS7) 10Ayounsi: Add Arelion Prometheus integration [puppet] - 10https://gerrit.wikimedia.org/r/1341098 (https://phabricator.wikimedia.org/T311005) [10:02:17] !log btullis@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [10:02:37] !log btullis@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [10:05:45] FIRING: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [10:06:40] FIRING: [2x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:07:23] (03CR) 10Muehlenhoff: [C:03+2] Remove the seed_image from the docker build config files [puppet] - 10https://gerrit.wikimedia.org/r/1342126 (owner: 10Muehlenhoff) [10:10:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [10:10:58] !log slyngshede@cumin1003 START - Cookbook sre.hosts.provision for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [10:13:25] (03CR) 10Muehlenhoff: [C:03+2] Remove buster from base build config [puppet] - 10https://gerrit.wikimedia.org/r/1342128 (owner: 10Muehlenhoff) [10:13:58] (03PS1) 10Muehlenhoff: Remove builder role from build2001 [puppet] - 10https://gerrit.wikimedia.org/r/1342206 (https://phabricator.wikimedia.org/T417389) [10:15:54] (03PS1) 10Muehlenhoff: wmflib: Remove buster from debian_php_version/debian_postgresql_version [puppet] - 10https://gerrit.wikimedia.org/r/1342207 [10:17:12] (03PS1) 10Jelto: gitlab: update gitlab-settings tag to v1.13.0 [puppet] - 10https://gerrit.wikimedia.org/r/1342208 (https://phabricator.wikimedia.org/T423984) [10:20:05] (03PS2) 10Muehlenhoff: Remove builder role from build2001 [puppet] - 10https://gerrit.wikimedia.org/r/1342206 (https://phabricator.wikimedia.org/T417389) [10:22:06] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5025.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [10:23:19] PROBLEM - ps1-603-eqsin-infeed-load-tower-B-single-phase on ps1-603-eqsin is CRITICAL: SNMP CRITICAL - ps1-603-eqsin-infeed-load-tower-B-single-phase *-1* https://wikitech.wikimedia.org/wiki/Dc-operations/Hardware_Troubleshooting_Runbook [10:25:25] (03CR) 10Arnaudb: [C:03+1] "looks good to me, thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1342208 (https://phabricator.wikimedia.org/T423984) (owner: 10Jelto) [10:25:37] FIRING: NetworkDeviceAlarmActive: Alarm active on cr3-eqsin - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr3-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [10:26:24] (03PS1) 10Jaime Nuche: Revert "Tell VisualEditor about the app web edit tags" [extensions/MobileApp] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342210 (https://phabricator.wikimedia.org/T437736) [10:27:19] (03CR) 10Muehlenhoff: [C:03+2] Remove builder role from build2001 [puppet] - 10https://gerrit.wikimedia.org/r/1342206 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [10:29:16] (03CR) 10Brouberol: [C:03+1] admin_ng: add analytics namespaces as storage tenants [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342198 (https://phabricator.wikimedia.org/T437447) (owner: 10Btullis) [10:33:19] RECOVERY - ps1-603-eqsin-infeed-load-tower-B-single-phase on ps1-603-eqsin is OK: SNMP OK - ps1-603-eqsin-infeed-load-tower-B-single-phase 0 https://wikitech.wikimedia.org/wiki/Dc-operations/Hardware_Troubleshooting_Runbook [10:34:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jnuche@deploy1003 using scap backport" [extensions/MobileApp] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342210 (https://phabricator.wikimedia.org/T437736) (owner: 10Jaime Nuche) [10:36:31] (03PS1) 10Marostegui: installserver: Do not format db1280 [puppet] - 10https://gerrit.wikimedia.org/r/1342213 [10:37:08] (03Merged) 10jenkins-bot: Revert "Tell VisualEditor about the app web edit tags" [extensions/MobileApp] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342210 (https://phabricator.wikimedia.org/T437736) (owner: 10Jaime Nuche) [10:37:33] !log jnuche@deploy1003 Started scap sync-world: Backport for [[gerrit:1342210|Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] [10:37:39] T437736: Visual Editor: Create public edit tag highlighting an edit came from VE in the app - https://phabricator.wikimedia.org/T437736 [10:37:39] T438125: Error: Interface "MediaWiki\Extension\VisualEditor\VisualEditorRegisterChangeTagsHook" not found - https://phabricator.wikimedia.org/T438125 [10:40:08] (03CR) 10Marostegui: [C:03+2] installserver: Do not format db1280 [puppet] - 10https://gerrit.wikimedia.org/r/1342213 (owner: 10Marostegui) [10:40:37] (03PS2) 10Elukey: Move the pki service entry to service_setup [puppet] - 10https://gerrit.wikimedia.org/r/1338937 (https://phabricator.wikimedia.org/T436809) [10:40:37] (03PS3) 10Elukey: role::pki: add lvs configurations [puppet] - 10https://gerrit.wikimedia.org/r/1338936 (https://phabricator.wikimedia.org/T436809) [10:40:37] (03PS1) 10Elukey: Move pki.discovery.wmnet to lvs_setup state [puppet] - 10https://gerrit.wikimedia.org/r/1342214 (https://phabricator.wikimedia.org/T436809) [10:40:37] RESOLVED: NetworkDeviceAlarmActive: Alarm active on cr3-eqsin - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr3-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [10:41:39] (03CR) 10Elukey: Move the pki service entry to service_setup (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338937 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [10:41:52] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1338936 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [10:41:56] (03PS1) 10Muehlenhoff: Remove wmflib::wmf_php_version [puppet] - 10https://gerrit.wikimedia.org/r/1342215 [10:42:26] (03CR) 10Elukey: [C:03+1] wmflib: Remove buster from debian_php_version/debian_postgresql_version [puppet] - 10https://gerrit.wikimedia.org/r/1342207 (owner: 10Muehlenhoff) [10:43:10] (03CR) 10Btullis: [C:03+2] admin_ng: add analytics namespaces as storage tenants [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342198 (https://phabricator.wikimedia.org/T437447) (owner: 10Btullis) [10:43:29] (03CR) 10Muehlenhoff: [C:03+2] wmflib: Remove buster from debian_php_version/debian_postgresql_version [puppet] - 10https://gerrit.wikimedia.org/r/1342207 (owner: 10Muehlenhoff) [10:44:25] (03CR) 10Ladsgroup: "The chart version needs to be bumped. I can take care of it while deploying I72f56caeab3d" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342065 (https://phabricator.wikimedia.org/T438089) (owner: 10Samwilson) [10:53:58] (03PS1) 10Muehlenhoff: Remove python-bullseye container image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342219 (https://phabricator.wikimedia.org/T416452) [10:54:24] (03Merged) 10jenkins-bot: admin_ng: add analytics namespaces as storage tenants [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342198 (https://phabricator.wikimedia.org/T437447) (owner: 10Btullis) [10:56:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 17.67% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [10:56:59] jouncebot: nowandnext [10:56:59] For the next 0 hour(s) and 3 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1000) [10:57:00] In 0 hour(s) and 3 minute(s): Services – Citoid / Zotero (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1100) [10:57:25] !log jnuche@deploy1003 jnuche: Backport for [[gerrit:1342210|Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [10:57:30] T437736: Visual Editor: Create public edit tag highlighting an edit came from VE in the app - https://phabricator.wikimedia.org/T437736 [10:57:30] T438125: Error: Interface "MediaWiki\Extension\VisualEditor\VisualEditorRegisterChangeTagsHook" not found - https://phabricator.wikimedia.org/T438125 [10:57:56] !log jnuche@deploy1003 jnuche: Continuing with deployment [11:00:05] mvolz: I, the Bot under the Fountain, call upon thee, The Deployer, to do Services – Citoid / Zotero deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1100). [11:01:00] !log btullis@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [11:01:42] !log btullis@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [11:01:55] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12327258 (10Fabfur) [11:02:08] (03PS1) 10Marostegui: instances.yaml: Remove db1180 [puppet] - 10https://gerrit.wikimedia.org/r/1342220 (https://phabricator.wikimedia.org/T437222) [11:03:57] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [11:03:57] (03CR) 10Marostegui: [C:03+2] instances.yaml: Remove db1180 [puppet] - 10https://gerrit.wikimedia.org/r/1342220 (https://phabricator.wikimedia.org/T437222) (owner: 10Marostegui) [11:04:30] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [11:04:30] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12327268 (10cmooney) [11:05:02] !log marostegui@cumin1004 dbctl commit (dc=all): 'Remove db1180 from dbctl T437222', diff saved to https://phabricator.wikimedia.org/P96459 and previous config saved to /var/cache/conftool/dbconfig/20260916-110502-marostegui.json [11:05:06] T437222: decommission db1180.eqiad.wmnet - https://phabricator.wikimedia.org/T437222 [11:06:34] (03PS1) 10Marostegui: db1180: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1342222 (https://phabricator.wikimedia.org/T437222) [11:07:53] (03CR) 10Marostegui: [C:03+2] db1180: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1342222 (https://phabricator.wikimedia.org/T437222) (owner: 10Marostegui) [11:09:58] (03CR) 10Mvolz: [C:03+2] zotero: update to latest [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341867 (https://phabricator.wikimedia.org/T435179) (owner: 10Mvolz) [11:10:30] !log slyngshede@cumin1003 START - Cookbook sre.hosts.reimage for host cp5025.eqsin.wmnet with OS trixie [11:10:45] (03CR) 10Hnowlan: [C:03+1] thumbor: Enable cache expiry TTL on two containers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341952 (https://phabricator.wikimedia.org/T433964) (owner: 10Ladsgroup) [11:10:48] !log jnuche@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342210|Revert "Tell VisualEditor about the app web edit tags" (T437736 T438125)]] (duration: 33m 14s) [11:10:53] T437736: Visual Editor: Create public edit tag highlighting an edit came from VE in the app - https://phabricator.wikimedia.org/T437736 [11:10:53] T438125: Error: Interface "MediaWiki\Extension\VisualEditor\VisualEditorRegisterChangeTagsHook" not found - https://phabricator.wikimedia.org/T438125 [11:11:46] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/zotero: apply [11:11:54] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/zotero: apply [11:11:59] blocker should be solved, I'm rolling the train forward [11:12:22] (03CR) 10Ladsgroup: Set JPGs to be thumbnailed via VIPS if md5 starts with a (032 comments) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [11:12:24] (03Merged) 10jenkins-bot: zotero: update to latest [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341867 (https://phabricator.wikimedia.org/T435179) (owner: 10Mvolz) [11:12:30] 10ops-eqiad, 06DC-Ops: Unresponsive management for db1245.mgmt:22 - https://phabricator.wikimedia.org/T438150 (10phaultfinder) 03NEW [11:12:34] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342223 (https://phabricator.wikimedia.org/T430839) [11:12:36] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342223 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [11:13:39] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342223 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [11:16:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.33% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:16:40] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/zotero: apply [11:17:02] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/zotero: apply [11:17:29] (03PS1) 10Urbanecm: mw::maintenance::growthexperiment: Do not use topictype=ores [puppet] - 10https://gerrit.wikimedia.org/r/1342224 (https://phabricator.wikimedia.org/T437889) [11:19:27] !log kicked off a new run of production-images-weekly-rebuild.service on build2004 (previously some leftovers of buster in the config prevented a complete run) [11:19:33] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:19:34] (03PS2) 10Urbanecm: mw::maintenance::growthexperiment: Do not use topictype=ores [puppet] - 10https://gerrit.wikimedia.org/r/1342224 (https://phabricator.wikimedia.org/T437889) [11:19:56] (03CR) 10Federico Ceratto: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:20:05] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/zotero: apply [11:20:31] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/zotero: apply [11:21:05] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/zotero: apply [11:21:25] FIRING: [2x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:21:33] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/zotero: apply [11:22:53] (03CR) 10Mvolz: [C:03+2] citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342103 (owner: 10PipelineBot) [11:24:56] !log jnuche@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.20 refs T430839 [11:25:01] T430839: 1.47.0-wmf.20 deployment blockers - https://phabricator.wikimedia.org/T430839 [11:25:22] (03Merged) 10jenkins-bot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342103 (owner: 10PipelineBot) [11:25:50] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [11:26:57] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [11:27:25] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [11:27:46] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [11:31:28] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/citoid: apply [11:31:58] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/citoid: apply [11:32:09] (03CR) 10CWilliams: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:33:50] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/citoid: apply [11:34:13] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/citoid: apply [11:35:04] (03PS5) 10Samtar: IS/IS-labs: Set wmgUseModeratorToolkit default false [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341284 (https://phabricator.wikimedia.org/T431000) [11:35:14] (03CR) 10CWilliams: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:35:43] !log slyngshede@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage [11:35:48] jouncebot: nowandnext [11:35:48] For the next 0 hour(s) and 24 minute(s): Services – Citoid / Zotero (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1100) [11:35:48] In 1 hour(s) and 24 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1300) [11:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.94% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:38:03] (03CR) 10CWilliams: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:38:21] (03PS4) 10Ladsgroup-claude: Set JPGs to be thumbnailed via VIPS if md5 starts with a [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) [11:38:32] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341284 (https://phabricator.wikimedia.org/T431000) (owner: 10Samtar) [11:39:07] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5025.eqsin.wmnet with reason: host reimage [11:39:29] (03Merged) 10jenkins-bot: IS/IS-labs: Set wmgUseModeratorToolkit default false [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341284 (https://phabricator.wikimedia.org/T431000) (owner: 10Samtar) [11:39:44] (03CR) 10Ladsgroup: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [11:39:54] !log samtar@deploy1003 Started scap sync-world: Backport for [[gerrit:1341284|IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] [11:39:58] T431000: Deploy the ModeratorToolkit extension to Beta Cluster - https://phabricator.wikimedia.org/T431000 [11:41:31] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342102 (owner: 10PipelineBot) [11:41:34] !log klausman@cumin1003 START - Cookbook sre.hosts.reboot-single for host ml-lab1002.eqiad.wmnet [11:41:39] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341883 (owner: 10PipelineBot) [11:41:48] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341692 (owner: 10PipelineBot) [11:41:56] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339757 (owner: 10PipelineBot) [11:42:09] (03CR) 10Majavah: [C:03+1] "not used for Toolforge" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342219 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [11:42:58] (03PS4) 10Krinkle: varnish: Move query string strip for upload.wm.o to pre-purge [puppet] - 10https://gerrit.wikimedia.org/r/1341937 (https://phabricator.wikimedia.org/T425216) [11:44:05] (03CR) 10Marostegui: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:44:23] !log samtar@deploy1003 samtar: Backport for [[gerrit:1341284|IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:45:57] !log samtar@deploy1003 samtar: Continuing with deployment [11:46:18] !log klausman@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-lab1002.eqiad.wmnet [11:46:43] Another blocker: T430839 [11:46:44] T430839: 1.47.0-wmf.20 deployment blockers - https://phabricator.wikimedia.org/T430839 [11:46:46] I need to roll back [11:47:16] (03CR) 10CI reject: [V:04-1] Set JPGs to be thumbnailed via VIPS if md5 starts with a [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [11:47:26] (03PS2) 10Muehlenhoff: Remove python-bullseye container image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342219 (https://phabricator.wikimedia.org/T416452) [11:48:36] I can't roll back, an unscheduled backport is happening: https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1341284 [11:49:07] jnuche: ah, I am deploying at the moment, apologies - I saw you were finished. The deploy is almost complete [11:50:27] !log samtar@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341284|IS/IS-labs: Set wmgUseModeratorToolkit default false (T431000)]] (duration: 10m 32s) [11:50:34] (03CR) 10CWilliams: backups-disk-space.yaml: backups disk space (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:50:37] T431000: Deploy the ModeratorToolkit extension to Beta Cluster - https://phabricator.wikimedia.org/T431000 [11:50:41] jnuche: done, sorry about that again [11:50:58] TheresNoTime: thanks, rolling back now [11:51:25] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342230 (https://phabricator.wikimedia.org/T430839) [11:51:28] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342230 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [11:53:00] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342230 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [11:54:25] (03CR) 10Jelto: [C:03+2] gitlab: update gitlab-settings tag to v1.13.0 [puppet] - 10https://gerrit.wikimedia.org/r/1342208 (https://phabricator.wikimedia.org/T423984) (owner: 10Jelto) [11:55:28] (03CR) 10Marostegui: "@tfogli@wikimedia.org any thoughts on the discussion? thank you!" [alerts] - 10https://gerrit.wikimedia.org/r/1341852 (https://phabricator.wikimedia.org/T438003) (owner: 10Federico Ceratto) [11:57:02] (03PS5) 10Ladsgroup-claude: Set JPGs to be thumbnailed via VIPS if md5 starts with a [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) [11:57:04] (03PS2) 10Samtar: IS-labs: Set wmgUseModeratorToolkit true for beta cluster enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341285 (https://phabricator.wikimedia.org/T431000) [11:58:19] (03PS1) 10STran: SI: Unset all filters on links to cases [extensions/CheckUser] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342231 (https://phabricator.wikimedia.org/T434530) [11:58:20] (03PS1) 10PipelineBot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342233 [11:58:32] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Remove python-bullseye container image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342219 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [11:58:44] (03PS1) 10STran: SI: Unset all filters on links to cases [extensions/CheckUser] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342234 (https://phabricator.wikimedia.org/T434530) [11:59:20] !log pruned obsolete Bullseye image python3-bullseye from the docker registry T416452 [11:59:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:59:27] T416452: Migrate Docker images running in Production away from Bullseye - https://phabricator.wikimedia.org/T416452 [11:59:40] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] "python3-bullseye has been removed from the docker registry" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342219 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [12:00:17] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for db1245.mgmt:22 - https://phabricator.wikimedia.org/T438150#12327590 (10Jclark-ctr) a:03Jclark-ctr [12:01:09] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for db1245.mgmt:22 - https://phabricator.wikimedia.org/T438150#12327595 (10Jclark-ctr) This is a repeat issue leaving ticket open till main issue is resolved [12:01:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 19.71% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:01:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.37% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:01:47] !log jnuche@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.20 refs T430839 [12:01:51] T430839: 1.47.0-wmf.20 deployment blockers - https://phabricator.wikimedia.org/T430839 [12:04:57] FIRING: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:05:00] (03PS1) 10Muehlenhoff: Switch Docker reporting from build2002 to build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1342237 (https://phabricator.wikimedia.org/T435314) [12:05:08] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#de" [extensions/CheckUser] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342231 (https://phabricator.wikimedia.org/T434530) (owner: 10STran) [12:05:28] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#de" [extensions/CheckUser] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342234 (https://phabricator.wikimedia.org/T434530) (owner: 10STran) [12:05:51] FIRING: SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://wikitech.wikimedia.org/wiki/Runbook#https://wikifeeds.svc.eqiad.wmnet:4101 - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [12:06:39] (03PS1) 10Kevin Bazira: ml-services: update tts isvc to fix captions drifting out of sync with audio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342239 (https://phabricator.wikimedia.org/T438151) [12:07:08] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342237 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [12:08:05] (03CR) 10Ozge: [C:03+1] ml-services: update tts isvc to fix captions drifting out of sync with audio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342239 (https://phabricator.wikimedia.org/T438151) (owner: 10Kevin Bazira) [12:08:21] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update tts isvc to fix captions drifting out of sync with audio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342239 (https://phabricator.wikimedia.org/T438151) (owner: 10Kevin Bazira) [12:09:44] !log cmooney@cumin1004 START - Cookbook sre.network.tls for network device ssw1-a1-eqiad [12:09:57] RESOLVED: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:10:57] FIRING: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:11:07] (03Merged) 10jenkins-bot: ml-services: update tts isvc to fix captions drifting out of sync with audio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342239 (https://phabricator.wikimedia.org/T438151) (owner: 10Kevin Bazira) [12:11:39] !log cmooney@cumin1004 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-a1-eqiad [12:11:57] FIRING: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:11:58] (03CR) 10Mszwarc: [C:03+1] SI: Unset all filters on links to cases [extensions/CheckUser] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342234 (https://phabricator.wikimedia.org/T434530) (owner: 10STran) [12:12:02] (03CR) 10Mszwarc: [C:03+1] SI: Unset all filters on links to cases [extensions/CheckUser] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342231 (https://phabricator.wikimedia.org/T434530) (owner: 10STran) [12:12:18] o/ [12:12:20] here [12:12:27] !incidents [12:12:27] 8346 (UNACKED) ProbeDown sre (10.2.2.47 ip4 wikifeeds:4101 probes/service http_wikifeeds_ip4 eqiad) [12:12:28] 8345 (RESOLVED) ProbeDown sre (10.2.2.47 ip4 wikifeeds:4101 probes/service http_wikifeeds_ip4 eqiad) [12:12:28] 8344 (RESOLVED) ProbeDown sre (2a02:ec80:600:ed1a::1 ip6 text-https:443 probes/service http_text-https_ip6 drmrs) [12:12:33] !ack 8346 [12:12:33] 8346 (ACKED) ProbeDown sre (10.2.2.47 ip4 wikifeeds:4101 probes/service http_wikifeeds_ip4 eqiad) [12:12:45] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [12:12:50] okay, let's look [12:13:13] one pod in crashloopbackoff [12:13:15] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5025.eqsin.wmnet with OS trixie [12:13:21] most pods unready [12:13:30] https://grafana.wikimedia.org/d/lxZAdAdMk/wikifeeds?from=now-1h&to=now&timezone=utc&var-dc=000000026&var-site=eqiad&var-prometheus=k8s&var-container_name=$__all&viewPanel=panel-10 [12:13:40] feels like scraping [12:14:18] general increase in timeouts, particularly for onthisday [12:14:22] yeah most likely [12:14:33] https://grafana.wikimedia.org/d/8169987e-2ef2-4bf2-ba85-eefad1edbefa/rest-gateway-per-service-breakdown?orgId=1&from=now-1h&to=now&timezone=utc&var-datasource=000000017&var-service=wikifeeds&var-route=$__all&refresh=1m [12:14:54] last deploy wwas around 12 hours ago [12:15:21] will you have a look at patterns? I will look at the pods and see if there's anything useful first [12:15:50] sounds good, I need to first even look up how the url looks like [12:15:51] RESOLVED: SwaggerProbeHasFailures: Not all openapi/swagger endpoints returned healthy - https://wikitech.wikimedia.org/wiki/Runbook#https://wikifeeds.svc.eqiad.wmnet:4101 - https://grafana.wikimedia.org/d/_77ik484k/openapi-swagger-endpoint-state?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DSwaggerProbeHasFailures [12:15:57] RESOLVED: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:16:57] RESOLVED: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:17:42] I'm not seeing increase in numbers but increase ttfb [12:17:56] https://w.wiki/UcJ8 [12:18:14] that actually looks like a bug or regression [12:19:15] yeah, there were some stuff with low browser score ten hours ago but not anymore [12:19:33] oh hum wikifeeds has been deployed yesterday and today [12:19:37] (03CR) 10Pmiazga: [C:04-1] "After moving MobileAppRedirect to MobileApp Extension, lets update the config variable here too. Sorry for the confusion" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339561 (https://phabricator.wikimedia.org/T434930) (owner: 10Milazg) [12:19:48] I'll get in touch with dbrant [12:19:57] FIRING: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:20:17] hnowlan: https://grafana.wikimedia.org/goto/swpmrl?orgId=default send this to him [12:21:00] it's sending seven times more requests to mobileapps service piling up [12:21:46] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [12:22:09] those spikes are really only happening today though [12:22:12] I've let him know [12:22:20] is mobileapps happy? [12:22:27] FIRING: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:22:35] yay [12:22:39] I think this is still externally generated though [12:23:22] the numbers don't show anything [12:23:26] pods are getting OOMkilled also [12:23:28] !incidents [12:23:28] 8347 (ACKED) ProbeDown sre (10.2.2.47 ip4 wikifeeds:4101 probes/service http_wikifeeds_ip4 eqiad) [12:23:28] 8346 (RESOLVED) ProbeDown sre (10.2.2.47 ip4 wikifeeds:4101 probes/service http_wikifeeds_ip4 eqiad) [12:23:28] 8345 (RESOLVED) ProbeDown sre (10.2.2.47 ip4 wikifeeds:4101 probes/service http_wikifeeds_ip4 eqiad) [12:23:29] 8344 (RESOLVED) ProbeDown sre (2a02:ec80:600:ed1a::1 ip6 text-https:443 probes/service http_text-https_ip6 drmrs) [12:24:53] it could be one of these widgets that refresh on hour, it started on 12 UTC [12:24:57] RESOLVED: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:26:03] it pag.ed earlier at 9:40 also [12:26:05] https://grafana.wikimedia.org/goto/sqs7bq?orgId=default [12:26:12] RESOLVED: ProbeDown: Service wikifeeds:4101 has failed probes (http_wikifeeds_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#wikifeeds:4101 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:26:18] that spike is showing up there too [12:26:35] (and around 9:40) [12:27:55] crazy increase in requests to mobileapps also [12:30:41] (03CR) 10Filippo Giunchedi: [C:03+1] Export a few stats about the magnum capi worker cluster (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1339182 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [12:33:08] !log sukhe@puppetserver1001 conftool action : set/pooled=no; selector: name=cp6002.drmrs.wmnet [12:34:04] !log sukhe@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cp6002.drmrs.wmnet with reason: reimage [12:34:49] !log sukhe@cumin1004 START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [12:36:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.48% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:38:02] 10SRE-SLO, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th): page_change SLO windows - https://phabricator.wikimedia.org/T438054#12327721 (10APizzata-WMF) Following further discussions with the team, we are designing an alternative approach for the Monthly Completeness SLO: on a monthly basis, all dai... [12:38:14] !log cmooney@cumin1004 START - Cookbook sre.network.tls for network device lsw1-a1-eqiad [12:38:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.28% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:38:55] !log cmooney@cumin1004 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a1-eqiad [12:40:59] !log dkertesz@cumin1004 START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: no reason specified, T438052] [12:41:03] T438052: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052 [12:41:14] !log dkertesz@cumin1004 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: no reason specified, T438052] [12:43:00] (03Abandoned) 10Tiziano Fogli: kafka-logging: add kafka-logging1003 with node ids 1003 [puppet] - 10https://gerrit.wikimedia.org/r/1329304 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [12:43:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:43:55] !log dkertesz@cumin1004 conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin,service=authdns-update [12:44:55] https://usercontent.irccloud-cdn.com/file/NDksCFUl/Screenshot%202026-09-16%20134422.png [12:45:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.8% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:45:23] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12327750 (10cmooney) @Jclark-ctr thanks for the help with the switches in rack A1 and the cabling. Both //lsw1-a1-eqiad// and //ssw1-a1-eqiad// are reachabl... [12:45:48] !log dkertesz@dns1004 START - running authdns-update [12:45:54] @hnowlan @dbrant @Amir1 this is what I'm seeing in Turnillo [12:47:59] !log dkertesz@dns1004 END - running authdns-update [12:48:56] RESOLVED: TransportLinksInUsageNoRedundancy: ulsfo inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=ulsfo - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [12:49:41] !log dkertesz@cumin1004 conftool action : set/pooled=yes; selector: cluster=dnsbox,dc=eqsin [12:50:38] (03CR) 10CWilliams: Revert "dbbackups: Pool db2250:s5 to replace db2201:s5 for backups" (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1342197 (owner: 10Marostegui) [12:50:54] (03PS1) 10Atsuko: airflow3: update image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342252 (https://phabricator.wikimedia.org/T437981) [12:51:36] (03CR) 10Marostegui: Revert "dbbackups: Pool db2250:s5 to replace db2201:s5 for backups" (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1342197 (owner: 10Marostegui) [12:52:07] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [12:52:44] !log sukhe@cumin1004 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet [12:53:32] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12327789 (10dkertesz) 05Open→03Resolved [12:54:04] (03CR) 10Brouberol: [C:03+1] airflow3: update image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342252 (https://phabricator.wikimedia.org/T437981) (owner: 10Atsuko) [12:54:28] (03CR) 10Atsuko: [C:03+2] airflow3: update image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342252 (https://phabricator.wikimedia.org/T437981) (owner: 10Atsuko) [12:56:11] (03CR) 10Elukey: [C:03+1] Switch Docker reporting from build2002 to build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1342237 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [12:56:54] (03Merged) 10jenkins-bot: airflow3: update image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342252 (https://phabricator.wikimedia.org/T437981) (owner: 10Atsuko) [12:59:43] !log dkertesz@cumin1004 conftool action : set/weight=1; selector: name=cp5025.eqsin.wmnet [13:00:05] urbanecm and TheresNoTime: May I have your attention please! UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1300) [13:00:06] Tran: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:17] o/ I'm here but I think there's a train blocker [13:00:44] !log dkertesz@cumin1004 conftool action : set/pooled=yes; selector: name=cp5025.eqsin.wmnet [13:02:16] (03PS2) 10Arnaudb: deployment_server/k8s: set kubeconfig files for aphlict [puppet] - 10https://gerrit.wikimedia.org/r/1341711 (https://phabricator.wikimedia.org/T436657) [13:02:41] PROBLEM - LDAP -writable server- on ldap-rw2001 is CRITICAL: Could not search/find objectclasses in dc=wikimedia,dc=org https://wikitech.wikimedia.org/wiki/LDAP%23Troubleshooting [13:03:04] (03PS3) 10Arnaudb: deployment_server/k8s: set kubeconfig files for aphlict [puppet] - 10https://gerrit.wikimedia.org/r/1341711 (https://phabricator.wikimedia.org/T436657) [13:03:58] (03CR) 10CWilliams: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1342197 (owner: 10Marostegui) [13:04:03] Tran: it's ok to go ahead, the blocker is not preventing backports [13:04:19] even to .20? [13:04:58] yeah, your .20 will be deployed to group0 only, that's the only effect. But I see yo have a .19 patch too, so you should see your change reflected everywhere [13:05:17] sounds good, I'm going to get started then. Thanks! [13:05:51] !log sukhe@cumin1004 END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts cp6002.drmrs.wmnet [13:05:54] !log sukhe@cumin1004 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp6002.drmrs.wmnet [13:06:53] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342231 (https://phabricator.wikimedia.org/T434530) (owner: 10STran) [13:06:54] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342234 (https://phabricator.wikimedia.org/T434530) (owner: 10STran) [13:07:28] 10SRE-SLO: Sloth: Add a CI test to verify that error queries involving specific metrics are defaulted to 0 - https://phabricator.wikimedia.org/T438159 (10tappof) 03NEW [13:08:28] 10SRE-SLO: Sloth: Add a CI test to verify that error queries involving specific metrics are defaulted to 0 - https://phabricator.wikimedia.org/T438159#12327877 (10tappof) https://gitlab.wikimedia.org/repos/sre/slothslos/-/merge_requests/52 [13:08:49] (03CR) 10Slyngshede: [C:03+1] idp: add the liftwing_studio service [puppet] - 10https://gerrit.wikimedia.org/r/1342135 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [13:08:59] sukhe@cumin1004 upgrade-firmware (PID 3004072) is awaiting input [13:09:00] (03Merged) 10jenkins-bot: SI: Unset all filters on links to cases [extensions/CheckUser] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342231 (https://phabricator.wikimedia.org/T434530) (owner: 10STran) [13:09:00] (03Merged) 10jenkins-bot: SI: Unset all filters on links to cases [extensions/CheckUser] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342234 (https://phabricator.wikimedia.org/T434530) (owner: 10STran) [13:09:23] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp6002.drmrs.wmnet [13:09:25] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1342231|SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234|SI: Unset all filters on links to cases (T434530)]] [13:09:29] T434530: Direct links to related SI cases are impacted by the filter - https://phabricator.wikimedia.org/T434530 [13:10:37] (03CR) 10Dpogorzelski: [C:03+2] idp: add the liftwing_studio service [puppet] - 10https://gerrit.wikimedia.org/r/1342135 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [13:10:38] (03CR) 10Ladsgroup: [C:03+2] thumbor: Enable cache expiry TTL on two containers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341952 (https://phabricator.wikimedia.org/T433964) (owner: 10Ladsgroup) [13:10:49] (03PS3) 10Arnaudb: modules: Prepare mesh.configuration minor version bump [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338959 (https://phabricator.wikimedia.org/T436657) [13:10:49] (03PS8) 10Arnaudb: mesh: add opt-in websocket support in configuration 1.17.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338751 (https://phabricator.wikimedia.org/T436657) [13:10:49] (03PS3) 10Arnaudb: mesh: document the idle timeout key the templates actually read [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338960 (https://phabricator.wikimedia.org/T436657) [13:10:50] (03PS8) 10Arnaudb: scaffold: point new services at mesh 1.17 and istio 1.5 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338752 (https://phabricator.wikimedia.org/T436657) [13:11:37] !log slyngshede@cumin1003 conftool action : set/pooled=no; selector: name=cp5026.eqsin.wmnet [13:12:46] (03CR) 10Slyngshede: [C:03+2] site.pp move cp5026 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342115 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [13:12:57] (03Merged) 10jenkins-bot: thumbor: Enable cache expiry TTL on two containers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341952 (https://phabricator.wikimedia.org/T433964) (owner: 10Ladsgroup) [13:12:58] (03PS2) 10Slyngshede: site.pp move cp5026 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342115 (https://phabricator.wikimedia.org/T436363) [13:13:44] !log stran@deploy1003 stran: Backport for [[gerrit:1342231|SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234|SI: Unset all filters on links to cases (T434530)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:14:59] (03PS4) 10Arnaudb: modules: Prepare mesh.configuration minor version bump [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338959 (https://phabricator.wikimedia.org/T436657) [13:14:59] (03PS9) 10Arnaudb: mesh: add opt-in websocket support in configuration 1.17.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338751 (https://phabricator.wikimedia.org/T436657) [13:14:59] (03PS4) 10Arnaudb: mesh: document the idle timeout key the templates actually read [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338960 (https://phabricator.wikimedia.org/T436657) [13:15:00] (03PS9) 10Arnaudb: scaffold: point new services at mesh 1.17 and istio 1.5 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338752 (https://phabricator.wikimedia.org/T436657) [13:15:06] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply [13:15:34] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply [13:16:18] testing now [13:16:30] (03PS8) 10Ayounsi: Add Arelion Prometheus integration [puppet] - 10https://gerrit.wikimedia.org/r/1341098 (https://phabricator.wikimedia.org/T311005) [13:17:54] (03PS4) 10Klausman: cookbooks/idm: Add user-cleanup cookbook for DPE SRE hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1342235 (https://phabricator.wikimedia.org/T437615) [13:17:59] lgtm, continuing [13:18:04] !log stran@deploy1003 stran: Continuing with deployment [13:18:36] (03CR) 10Slyngshede: [C:03+2] site.pp move cp5026 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342115 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [13:19:12] !log sukhe@cumin1004 START - Cookbook sre.hosts.provision for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [13:19:24] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [13:19:30] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [13:22:35] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342231|SI: Unset all filters on links to cases (T434530)]], [[gerrit:1342234|SI: Unset all filters on links to cases (T434530)]] (duration: 13m 10s) [13:22:38] T434530: Direct links to related SI cases are impacted by the filter - https://phabricator.wikimedia.org/T434530 [13:22:56] !log slyngshede@cumin1003 START - Cookbook sre.hosts.provision for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [13:24:05] !log btullis@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [13:24:35] !log btullis@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [13:25:17] done [13:25:30] !log bking@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on dse-k8s-etcd[1001-1003].eqiad.wmnet with reason: Maintenance to increase vCPUS T438084 [13:25:34] T438084: dse-k8s-etcd: Increase vCPU count - https://phabricator.wikimedia.org/T438084 [13:26:05] PROBLEM - Host cp5026 is DOWN: PING CRITICAL - Packet loss = 100% [13:26:30] 10ops-eqiad, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161 (10Jclark-ctr) 03NEW [13:27:01] 10ops-eqiad, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12327979 (10Jclark-ctr) [13:27:45] 10ops-eqiad, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12327983 (10Jclark-ctr) Server Came with no luggage tag or serial sticker on outside of chassis will have to open chassis to check serial and password on motherboard [13:29:18] (03CR) 10Ssingh: [C:03+1] "Looks good, one suggested edit, please feel free to ignore" [puppet] - 10https://gerrit.wikimedia.org/r/1328543 (https://phabricator.wikimedia.org/T435799) (owner: 10Fabfur) [13:29:45] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp6002.mgmt.drmrs.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [13:30:17] 10ops-eqiad, 06SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12327993 (10Jclark-ctr) [13:33:34] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5026.mgmt.eqsin.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [13:33:46] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [13:33:55] PROBLEM - Host cp6002 is DOWN: PING CRITICAL - Packet loss = 100% [13:34:12] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [13:34:18] !log slyngshede@cumin1003 START - Cookbook sre.hosts.reimage for host cp5026.eqsin.wmnet with OS trixie [13:35:36] (03CR) 10Ssingh: [C:03+2] site: Move cp6002 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1342040 (https://phabricator.wikimedia.org/T436363) (owner: 10BCornwall) [13:35:47] 06SRE, 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Requesting access to analytics-privatedata-users level 3 for derenrich - https://phabricator.wikimedia.org/T437668#12328019 (10HSwan-WMF) @hnowlan Approved. Thank you! [13:37:09] !log bking@cumin2003 START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet [13:37:15] !log bking@cumin2003 END (FAIL) - Cookbook sre.ganeti.reboot-vm (exit_code=99) for VM dse-k8s-etcd1003.eqiad.wmnet [13:37:26] !log bking@cumin2003 START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1003.eqiad.wmnet [13:38:06] !log sukhe@cumin1004 START - Cookbook sre.hosts.reimage for host cp6002.drmrs.wmnet with OS trixie [13:38:16] (03CR) 10Muehlenhoff: [C:03+2] Switch Docker reporting from build2002 to build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1342237 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [13:39:45] 10ops-eqiad, 06SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12328074 (10Jclark-ctr) [13:41:11] !log bking@cumin2003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1003.eqiad.wmnet [13:41:26] !log bking@cumin2003 START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1002.eqiad.wmnet [13:45:04] RECOVERY - Host cp6002 is UP: PING OK - Packet loss = 0%, RTA = 81.88 ms [13:45:16] !log bking@cumin2003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1002.eqiad.wmnet [13:46:13] !log bking@cumin2003 START - Cookbook sre.ganeti.reboot-vm for VM dse-k8s-etcd1001.eqiad.wmnet [13:46:25] RESOLVED: SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:46:55] !log bking@cumin2003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM dse-k8s-etcd1001.eqiad.wmnet [13:47:12] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:48:27] !log bking@cumin2003 START - Cookbook sre.hosts.remove-downtime for dse-k8s-etcd[1001-1003].eqiad.wmnet [13:48:30] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for dse-k8s-etcd[1001-1003].eqiad.wmnet [13:48:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [13:52:14] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [13:52:49] (03PS1) 10Urbanecm: JsonSchemaBuilder: Cache the root schema in the process [extensions/CommunityConfiguration] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342264 (https://phabricator.wikimedia.org/T437588) [13:53:03] (03PS1) 10Urbanecm: JsonSchemaBuilder: Cache the root schema in the process [extensions/CommunityConfiguration] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1342265 (https://phabricator.wikimedia.org/T437588) [13:53:31] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [13:54:39] !log sukhe@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage [13:54:46] (03PS1) 10Jforrester: wikifunctions: Upgrade evaluators from 2026-09-08-175644 to 2026-09-14-201648 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342267 (https://phabricator.wikimedia.org/T433459) [13:54:49] (03PS1) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-09-08-173612 to 2026-09-16-123542 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342268 (https://phabricator.wikimedia.org/T322056) [13:56:43] !log slyngshede@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage [13:57:32] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade evaluators from 2026-09-08-175644 to 2026-09-14-201648 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342267 (https://phabricator.wikimedia.org/T433459) (owner: 10Jforrester) [13:59:26] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6002.drmrs.wmnet with reason: host reimage [14:00:04] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1400) [14:00:14] (03Merged) 10jenkins-bot: wikifunctions: Upgrade evaluators from 2026-09-08-175644 to 2026-09-14-201648 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342267 (https://phabricator.wikimedia.org/T433459) (owner: 10Jforrester) [14:00:57] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [14:01:29] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:02:08] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [14:02:18] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:02:31] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:03:32] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5026.eqsin.wmnet with reason: host reimage [14:03:41] (03PS2) 10Jforrester: abstractwiki: Add three new articles per community advice to show off the feature [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342008 (https://phabricator.wikimedia.org/T434227) [14:04:54] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342008 (https://phabricator.wikimedia.org/T434227) (owner: 10Jforrester) [14:07:07] (03Merged) 10jenkins-bot: abstractwiki: Add three new articles per community advice to show off the feature [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342008 (https://phabricator.wikimedia.org/T434227) (owner: 10Jforrester) [14:07:32] !log jforrester@deploy1003 Started scap sync-world: Backport for [[gerrit:1342008|abstractwiki: Add three new articles per community advice to show off the feature (T434227)]] [14:07:36] T434227: Add some more articles to the weekly AW export so the generation metrics cover a wider range of content types/lengths - https://phabricator.wikimedia.org/T434227 [14:07:53] (03CR) 10Muehlenhoff: [C:03+2] Remove unused profile::openldap::hostname [puppet] - 10https://gerrit.wikimedia.org/r/1339732 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [14:08:23] !log dropped links tables on db1238 (T437278) [14:08:25] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:08:26] T437278: Drop unneeded tables from x4 and s4 - https://phabricator.wikimedia.org/T437278 [14:08:34] (03CR) 10JHathaway: [C:03+1] "looks good, thanks" [puppet] - 10https://gerrit.wikimedia.org/r/1342190 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [14:09:10] (03CR) 10JHathaway: [C:03+1] specs: change explicit test_on to defaults [puppet] - 10https://gerrit.wikimedia.org/r/1342189 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [14:09:47] !log dropped links tables on db2237 (T437278) [14:09:50] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:10:02] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync [14:10:06] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync [14:10:16] !log jforrester@deploy1003 sync-world failed: Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte [14:10:16] d/mediawiki-multiversion --singleversion-image-basename docker-registry.discovery.wmnet/restricted/mediawiki-singleversion --webserver-image-name docker-registry.discovery.wmnet/restricted/mediawiki-webserver --latest-tag latest --label vnd.wikimedia.builder.name=scap --label vnd.wikimedia.builder.version=4.289.0 --label vnd.wikimedia.scap.stage_dir=/srv/mediawiki-staging --label vnd.wikimedia.scap.build_state_dir=/srv/me [14:10:16] diawiki-staging/scap/image-build' returned non-zero exit status 1. (scap version: 4.289.0) (duration: 02m 44s) [14:10:32] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync [14:10:35] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync [14:11:14] I’m getting No Space Left On Device from the MW docker build in Spiderpig. [14:12:44] !log jforrester@deploy1003 Started scap sync-world: Backport for [[gerrit:1342008|abstractwiki: Add three new articles per community advice to show off the feature (T434227)]] [14:12:48] T434227: Add some more articles to the weekly AW export so the generation metrics cover a wider range of content types/lengths - https://phabricator.wikimedia.org/T434227 [14:12:49] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:13:00] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:13:50] oh no. that probably means i should NOT attempt a MW deploy James_F [14:13:59] urbanecm: Certainly not during my window. [14:14:09] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:14:19] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:14:22] among other reasons [14:14:24] !log jforrester@deploy1003 sync-world failed: Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.19,1.47.0-wmf.20,next --multiversion-image-basename docker-registry.discovery.wmnet/restricte [14:14:24] d/mediawiki-multiversion --singleversion-image-basename docker-registry.discovery.wmnet/restricted/mediawiki-singleversion --webserver-image-name docker-registry.discovery.wmnet/restricted/mediawiki-webserver --latest-tag latest --label vnd.wikimedia.builder.name=scap --label vnd.wikimedia.builder.version=4.289.0 --label vnd.wikimedia.scap.stage_dir=/srv/mediawiki-staging --label vnd.wikimedia.scap.build_state_dir=/srv/me [14:14:24] diawiki-staging/scap/image-build' returned non-zero exit status 1. (scap version: 4.289.0) (duration: 01m 39s) [14:14:32] Yeah, failed on the retry too. [14:14:42] Our config patch isn’t urgent, but not really worth reverting. [14:17:23] Looking, it says /srv is 94% used with 19GiB free, which maybe isn’t enough for 2x2x4GiB docker builds? (Two PHP versions as well as two deployment branches.) [14:24:13] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: sync [14:24:29] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:24:32] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:24:40] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: sync [14:25:15] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-09-08-173612 to 2026-09-16-123542 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342268 (https://phabricator.wikimedia.org/T322056) (owner: 10Jforrester) [14:26:30] (03CR) 10Muehlenhoff: [C:03+1] "Looks great, one comment inline." [puppet] - 10https://gerrit.wikimedia.org/r/1322963 (https://phabricator.wikimedia.org/T433601) (owner: 10JHathaway) [14:26:47] FIRING: HelmReleaseBadStatus: Helm release wikifunctions/javascript-evaluator on k8s@eqiad in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s&var-namespace=wikifunctions - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:27:30] (03CR) 10Jforrester: [C:03+1] wikifunctions: Upgrade orchestrator from 2026-09-08-173612 to 2026-09-16-123542 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342268 (https://phabricator.wikimedia.org/T322056) (owner: 10Jforrester) [14:28:31] (03CR) 10JHathaway: [C:03+2] jenkins: run specs against all default Debians [puppet] - 10https://gerrit.wikimedia.org/r/1342190 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [14:28:37] (03CR) 10JHathaway: [C:03+2] specs: change explicit test_on to defaults [puppet] - 10https://gerrit.wikimedia.org/r/1342189 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [14:28:48] (03CR) 10Elukey: [C:03+2] sre.hosts.provision: avoid user locked out for Supermicro BMCs [cookbooks] - 10https://gerrit.wikimedia.org/r/1334894 (https://phabricator.wikimedia.org/T419892) (owner: 10Elukey) [14:29:56] !log sukhe@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6002.drmrs.wmnet with OS trixie [14:30:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1400) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1430) [14:30:39] 10ops-eqiad, 06SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12328480 (10elukey) @Jclark-ctr even for this host let's make sure provision work, it shouldn't change much from what we added the last time in the cookbook for ml-serve1012-15 but better safe t... [14:31:47] (03PS2) 10Muehlenhoff: Remove obsolete nutcracker container image based on bullseye [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341665 (https://phabricator.wikimedia.org/T416452) [14:32:16] !log sukhe@puppetserver1001 conftool action : set/weight=1; selector: name=cp6002.drmrs.wmnet,service=cdn [14:32:20] !log sukhe@puppetserver1001 conftool action : set/weight=100; selector: name=cp6002.drmrs.wmnet,service=ats-be [14:33:46] !log sukhe@puppetserver1001 conftool action : set/pooled=yes; selector: name=cp6002.drmrs.wmnet [14:34:45] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:34:52] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12328544 (10Jclark-ctr) @eevans Dell is sending out a new Motherboard, backplane , and cables for this server due to availability it might be a little while for motherbo... [14:35:05] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:35:57] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:36:47] RESOLVED: HelmReleaseBadStatus: Helm release wikifunctions/javascript-evaluator on k8s@eqiad in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s&var-namespace=wikifunctions - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:38:14] (03PS1) 10Muehlenhoff: Create repository components and sync definitions for nodejs 26 [puppet] - 10https://gerrit.wikimedia.org/r/1342277 (https://phabricator.wikimedia.org/T437510) [14:42:04] !log cdobbins@cumin1004 START - Cookbook sre.hosts.reimage for host ncredir7003.magru.wmnet with OS trixie [14:44:21] !log installing python-filelock security updates [14:44:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:45:07] (03PS5) 10Krinkle: varnish: Move query string strip for upload.wm.o to pre-purge [puppet] - 10https://gerrit.wikimedia.org/r/1341937 (https://phabricator.wikimedia.org/T425216) [14:45:44] FIRING: KubernetesDeploymentUnavailableReplicas: ... [14:45:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [14:45:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [14:45:57] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5026.eqsin.wmnet with OS trixie [14:46:11] (03CR) 10Hashar: ci: set ci-build-images image and update creates (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1333796 (https://phabricator.wikimedia.org/T436775) (owner: 10Hashar) [14:46:22] (03PS4) 10Hashar: ci: set ci-build-images image and update creates [puppet] - 10https://gerrit.wikimedia.org/r/1333796 (https://phabricator.wikimedia.org/T436775) [14:46:34] (03PS1) 10Jforrester: Use parser output value instead of status [extensions/FlaggedRevs] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342279 (https://phabricator.wikimedia.org/T438154) [14:47:21] !log slyngshede@cumin1003 conftool action : set/weight=1; selector: name=cp5026.eqsin.wmnet [14:49:00] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342281 (https://phabricator.wikimedia.org/T430839) [14:49:03] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jnuche@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342281 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [14:49:11] (03CR) 10JHathaway: [C:03+1] Create repository components and sync definitions for nodejs 26 [puppet] - 10https://gerrit.wikimedia.org/r/1342277 (https://phabricator.wikimedia.org/T437510) (owner: 10Muehlenhoff) [14:49:30] !log slyngshede@cumin1003 conftool action : set/pooled=yes; selector: name=cp5026.eqsin.wmnet [14:50:33] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.20 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342281 (https://phabricator.wikimedia.org/T430839) (owner: 10TrainBranchBot) [14:50:44] RESOLVED: KubernetesDeploymentUnavailableReplicas: ... [14:50:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [14:50:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [14:50:49] !log installing apache2 security updates [14:50:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:51:46] James_F: JFYI T438148 [14:51:47] T438148: Image building is running out of space in `/srv` - https://phabricator.wikimedia.org/T438148 [14:51:58] jnuche: Aha, thanks. [14:52:35] jnuche: Clear the fix is to ram through the PHP 8.5 migration all in one go today, so we have half the number of MW images. [14:53:20] (03CR) 10Volans: "[this is not a review] just one comment inline" [cookbooks] - 10https://gerrit.wikimedia.org/r/1342235 (https://phabricator.wikimedia.org/T437615) (owner: 10Klausman) [14:53:48] James_F: hehehe [14:58:30] (03CR) 10Andrew Bogott: [C:03+2] Export a few stats about the magnum capi worker cluster [puppet] - 10https://gerrit.wikimedia.org/r/1339182 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [14:58:37] (03CR) 10Andrew Bogott: [C:03+2] magnum: move capi worker kubeconfig path to hiera [puppet] - 10https://gerrit.wikimedia.org/r/1339193 (owner: 10Andrew Bogott) [14:59:44] Traffic is now done migrating CP nodes from upload to text, two hosts per DC have been moved. Please see email and/or https://phabricator.wikimedia.org/T436363 [15:01:31] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jnuche@deploy1003 using scap backport" [extensions/FlaggedRevs] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342279 (https://phabricator.wikimedia.org/T438154) (owner: 10Jforrester) [15:01:44] FIRING: KubernetesDeploymentUnavailableReplicas: ... [15:01:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [15:01:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [15:02:59] * Raine just came from their therapy session to the "ram through the PHP 8.5 migration all in one go today" thought and now needs to go back in [15:02:59] (03Merged) 10jenkins-bot: Use parser output value instead of status [extensions/FlaggedRevs] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342279 (https://phabricator.wikimedia.org/T438154) (owner: 10Jforrester) [15:03:31] !log jnuche@deploy1003 Started scap sync-world: Backport for [[gerrit:1342279|Use parser output value instead of status (T438154)]] [15:03:35] T438154: TypeError: MediaWiki\Output\OutputPage::addPostProcessedParserOutput(): Argument #1 ($parserOutput) must be of type MediaWiki\Parser\ParserOutput, MediaWiki\Status\Status given, called in /srv/mediawiki/php-1.47.0-wmf.20/extens - https://phabricator.wikimedia.org/T438154 [15:06:44] RESOLVED: KubernetesDeploymentUnavailableReplicas: ... [15:06:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [15:06:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [15:07:36] (03CR) 10Hnowlan: [C:03+2] icinga: migrate cert check to prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1333749 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [15:07:56] 06SRE, 10SRE-Access-Requests, 06Data-Engineering, 06Data-Engineering-Radar, 13Patch-For-Review: Requesting access to analytics-privatedata-users level 3 for derenrich - https://phabricator.wikimedia.org/T437668#12328931 (10Ahoelzl) Approved. [15:08:07] 06SRE, 10SRE-Access-Requests, 06Data-Engineering, 06Data-Engineering-Radar, 13Patch-For-Review: Requesting access to analytics-privatedata-users level 3 for derenrich - https://phabricator.wikimedia.org/T437668#12328933 (10Ahoelzl) [15:08:38] (03PS1) 10Muehlenhoff: Switch firmware management host to cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1342284 (https://phabricator.wikimedia.org/T427897) [15:09:28] (03PS1) 10Sbisson: Keep Article Guidance on where it is on today [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342285 (https://phabricator.wikimedia.org/T433293) [15:09:51] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342284 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [15:10:52] (03CR) 10JHathaway: [C:03+1] Switch firmware management host to cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1342284 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [15:14:22] !log cdobbins@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage [15:16:02] (03CR) 10Elukey: [C:03+1] "LGTM to test it out!" [puppet] - 10https://gerrit.wikimedia.org/r/1341098 (https://phabricator.wikimedia.org/T311005) (owner: 10Ayounsi) [15:17:26] (03CR) 10Ladsgroup: "recheck" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [15:18:36] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7003.magru.wmnet with reason: host reimage [15:19:42] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie [15:19:56] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12329004 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS t... [15:20:03] (03CR) 10Giuseppe Lavagetto: [C:03+1] build-production-images: Remove buster from the rest of the base images [puppet] - 10https://gerrit.wikimedia.org/r/1342072 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [15:23:45] (03PS2) 10JHathaway: nftables: add ipencap support [puppet] - 10https://gerrit.wikimedia.org/r/1322963 (https://phabricator.wikimedia.org/T433601) [15:23:46] !log jnuche@deploy1003 jnuche, jforrester: Backport for [[gerrit:1342279|Use parser output value instead of status (T438154)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:23:50] T438154: TypeError: MediaWiki\Output\OutputPage::addPostProcessedParserOutput(): Argument #1 ($parserOutput) must be of type MediaWiki\Parser\ParserOutput, MediaWiki\Status\Status given, called in /srv/mediawiki/php-1.47.0-wmf.20/extens - https://phabricator.wikimedia.org/T438154 [15:23:54] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1322963 (https://phabricator.wikimedia.org/T433601) (owner: 10JHathaway) [15:24:21] (03CR) 10Giuseppe Lavagetto: [C:03+1] Remove seed_image config field, unused since 2023 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342074 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [15:24:21] !log jnuche@deploy1003 jnuche, jforrester: Continuing with deployment [15:25:09] (03CR) 10Giuseppe Lavagetto: [C:03+1] build-production-images: Remove seed_image field, unused since 2023 [puppet] - 10https://gerrit.wikimedia.org/r/1342073 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [15:28:57] (03PS3) 10JHathaway: nftables: add ipencap support [puppet] - 10https://gerrit.wikimedia.org/r/1322963 (https://phabricator.wikimedia.org/T433601) [15:30:07] (03CR) 10Blake: [C:03+2] mediawiki: stop using lamp.deployment. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338227 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:33:38] (03Merged) 10jenkins-bot: mediawiki: stop using lamp.deployment. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338227 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:35:26] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-cron: apply [15:35:39] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-cron: apply [15:36:53] !log jnuche@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342279|Use parser output value instead of status (T438154)]] (duration: 33m 21s) [15:36:57] T438154: TypeError: MediaWiki\Output\OutputPage::addPostProcessedParserOutput(): Argument #1 ($parserOutput) must be of type MediaWiki\Parser\ParserOutput, MediaWiki\Status\Status given, called in /srv/mediawiki/php-1.47.0-wmf.20/extens - https://phabricator.wikimedia.org/T438154 [15:37:59] (03CR) 10JHathaway: [C:03+2] nftables: add ipencap support (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1322963 (https://phabricator.wikimedia.org/T433601) (owner: 10JHathaway) [15:38:30] (03CR) 10JHathaway: "IPIP support is now merged, if you want to incorporate it into this patch." [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [15:39:10] (03PS1) 10Andrew Bogott: openstack_magnum_capi_worker_stats_exporter: run exporter as 'magnum' [puppet] - 10https://gerrit.wikimedia.org/r/1342297 (https://phabricator.wikimedia.org/T429557) [15:41:05] (03PS2) 10Samtar: CommonSettings-labs: Load ModeratorToolkit extension [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341286 (https://phabricator.wikimedia.org/T431000) [15:41:49] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-cron: apply [15:42:09] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7003.magru.wmnet with OS trixie [15:42:15] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply [15:42:25] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-cron: apply [15:42:30] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-cron: apply [15:43:00] (03CR) 10Andrew Bogott: [C:03+2] openstack_magnum_capi_worker_stats_exporter: run exporter as 'magnum' [puppet] - 10https://gerrit.wikimedia.org/r/1342297 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [15:45:36] (03PS1) 10Gmodena: wdqs: set QLEVER_NUM_QUERIES to 2x allocated CPU cores [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342298 (https://phabricator.wikimedia.org/T431395) [15:50:09] PROBLEM - MariaDB Replica Lag: x3 on clouddb1023 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 554.15 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:50:54] (03CR) 10Dillon: [C:03+1] "LGTM, thanks!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341285 (https://phabricator.wikimedia.org/T431000) (owner: 10Samtar) [15:51:09] PROBLEM - MariaDB Replica Lag: s3 on clouddb1023 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 588.06 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:51:44] FIRING: KubernetesDeploymentUnavailableReplicas: ... [15:51:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [15:51:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [15:52:26] (03PS1) 10Andrew Bogott: openstack_magnum_capi_worker_stats_exporter: run exporter as 'root' [puppet] - 10https://gerrit.wikimedia.org/r/1342300 (https://phabricator.wikimedia.org/T429557) [15:52:30] (03PS1) 10Ayounsi: Add netbox-bgp and update wheels [software/netbox-deploy] - 10https://gerrit.wikimedia.org/r/1342301 [15:52:33] (03CR) 10Dillon: [C:03+1] "LGTM, thanks!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341286 (https://phabricator.wikimedia.org/T431000) (owner: 10Samtar) [15:55:50] jouncebot: nowandnext [15:55:50] No deployments scheduled for the next 1 hour(s) and 4 minute(s) [15:55:50] In 1 hour(s) and 4 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1700) [15:56:47] (03CR) 10Andrew Bogott: [C:03+2] openstack_magnum_capi_worker_stats_exporter: run exporter as 'root' [puppet] - 10https://gerrit.wikimedia.org/r/1342300 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [15:57:51] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12329310 (10VRiley-WMF) Hey @Dzahn currently, right now I'm trying to update the firmware. However, if this continues to fail, admittedly I... [15:59:45] (03CR) 10Cathal Mooney: "Awesome, thank you!" [puppet] - 10https://gerrit.wikimedia.org/r/1319486 (https://phabricator.wikimedia.org/T433601) (owner: 10Cathal Mooney) [16:03:42] hi chat I going to deploy one or two beta-only patches in a moment fyi [16:04:59] (03CR) 10Krinkle: "I can't reproduce this error." [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [16:05:06] (03PS3) 10Samtar: IS-labs: Set wmgUseModeratorToolkit true for beta cluster enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341285 (https://phabricator.wikimedia.org/T431000) [16:06:28] (03CR) 10Hnowlan: [C:03+1] "Looks okay to me bar the config default." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [16:07:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341285 (https://phabricator.wikimedia.org/T431000) (owner: 10Samtar) [16:08:07] (03Merged) 10jenkins-bot: IS-labs: Set wmgUseModeratorToolkit true for beta cluster enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341285 (https://phabricator.wikimedia.org/T431000) (owner: 10Samtar) [16:08:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [16:09:36] (03PS3) 10Samtar: CommonSettings-labs: Load ModeratorToolkit extension [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341286 (https://phabricator.wikimedia.org/T431000) [16:12:12] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:14:09] RECOVERY - MariaDB Replica Lag: x3 on clouddb1023 is OK: OK slave_sql_lag Replication lag: 0.10 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:14:09] RECOVERY - MariaDB Replica Lag: s3 on clouddb1023 is OK: OK slave_sql_lag Replication lag: 37.12 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:14:19] (03CR) 10RLazarus: [C:03+2] build-production-images: Remove buster from the rest of the base images [puppet] - 10https://gerrit.wikimedia.org/r/1342072 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [16:14:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341286 (https://phabricator.wikimedia.org/T431000) (owner: 10Samtar) [16:14:58] (03PS26) 10Herron: sre.opensearch.roll-restart-reboot: include checklist items [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) [16:15:34] (03PS6) 10Ladsgroup: Set JPGs to be thumbnailed via VIPS if md5 starts with given config [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [16:15:40] (03CR) 10Ladsgroup: Set JPGs to be thumbnailed via VIPS if md5 starts with given config (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [16:15:49] (03Merged) 10jenkins-bot: CommonSettings-labs: Load ModeratorToolkit extension [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341286 (https://phabricator.wikimedia.org/T431000) (owner: 10Samtar) [16:16:19] 06SRE, 10LDAP-Access-Requests: Request for LDAP NDA Access for CentralNotice Metrics for Yahya - https://phabricator.wikimedia.org/T436294#12329488 (10NAramayo-WMF) @ssingh Confirming that the NDA has been completed and this process can proceed! [16:16:34] 10ops-eqiad, 06SRE, 06DC-Ops: krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12329492 (10VRiley-WMF) Okay, so supermicro came back and said the following... Hi Team, We have reviewed the boot logs provided for the system Based on our analysis, there are no explicit error m... [16:16:37] moritzm: ha, I was about to merge https://gerrit.wikimedia.org/r/1342072 and friends, and I see you sniped me overnight :D thanks [16:17:12] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:17:50] (03Abandoned) 10RLazarus: build-production-images: Remove buster from the rest of the base images [puppet] - 10https://gerrit.wikimedia.org/r/1342072 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [16:18:12] (03CR) 10Herron: sre.opensearch.roll-restart-reboot: include checklist items (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [16:18:33] (03Abandoned) 10RLazarus: build-production-images: Remove seed_image field, unused since 2023 [puppet] - 10https://gerrit.wikimedia.org/r/1342073 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [16:19:24] vriley@cumin1003 reimage (PID 357640) is awaiting input [16:19:30] (03CR) 10RLazarus: [V:03+2 C:03+2] Remove seed_image config field, unused since 2023 [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342074 (https://phabricator.wikimedia.org/T438101) (owner: 10RLazarus) [16:19:40] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie [16:19:55] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12329549 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host zuul1006.eqiad.wmnet with OS trixi... [16:20:15] 06SRE, 10docker-pkg, 06serviceops-deprecated, 13Patch-For-Review, 07Technical-Debt: Get rid of the concept of "seed image" in docker-pkg - https://phabricator.wikimedia.org/T272154#12329552 (10RLazarus) 05Open→03Resolved a:03Joe [16:20:20] 10ops-codfw, 10ops-eqiad, 06SRE, 06DC-Ops: Fix PXE miss-configurations - https://phabricator.wikimedia.org/T396717#12329554 (10VRiley-WMF) p:05High→03Medium [16:21:44] RESOLVED: KubernetesDeploymentUnavailableReplicas: ... [16:21:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [16:21:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [16:21:51] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [16:22:29] (03PS2) 10Blake: mediawiki: remove lamp.deployment. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342304 (https://phabricator.wikimedia.org/T417800) [16:25:21] (03CR) 10Hnowlan: [C:03+1] Set JPGs to be thumbnailed via VIPS if md5 starts with given config [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [16:26:26] (03CR) 10Blake: [C:03+1] Depool poolcounter[1006,2005] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342089 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [16:26:33] (03CR) 10Blake: [C:03+1] Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [16:26:40] (03CR) 10Blake: [C:03+1] Repool poolcounter[1007,2006] [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342091 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [16:29:40] (03CR) 10RLazarus: [C:03+1] "https://www.youtube.com/watch?v=Nix6tC3vvjs" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342304 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [16:29:56] (03CR) 10Ladsgroup: Set JPGs to be thumbnailed via VIPS if md5 starts with given config (032 comments) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [16:29:59] (03CR) 10Ladsgroup: [C:03+2] Set JPGs to be thumbnailed via VIPS if md5 starts with given config [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [16:30:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.8% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:30:39] 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops, 06Infrastructure-Foundations, 10provisioning-automation: Evaluate hw-raid controllers for Supermicro's Config J - https://phabricator.wikimedia.org/T378584#12329593 (10LSobanski) [16:31:23] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack: Add tracing support to Spicerack - https://phabricator.wikimedia.org/T418874#12329598 (10LSobanski) [16:31:29] 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops, 06Infrastructure-Foundations, 10provisioning-automation: Evaluate hw-raid controllers for Supermicro's Config J - https://phabricator.wikimedia.org/T378584#12329600 (10elukey) 05Open→03Resolved a:03elukey [16:32:29] (03CR) 10Blake: "yep that's definitely the saddest i've ever been about a lamp" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342304 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [16:35:39] (03CR) 10Jforrester: [C:03+1] "Thank you, looks great." [puppet] - 10https://gerrit.wikimedia.org/r/1342277 (https://phabricator.wikimedia.org/T437510) (owner: 10Muehlenhoff) [16:38:15] 06SRE, 06Infrastructure-Foundations: Create nodejs 26 production images - https://phabricator.wikimedia.org/T437509#12329638 (10Jdforrester-WMF) Patch from last time: https://gerrit.wikimedia.org/r/c/operations/docker-images/production-images/+/1248488 [16:39:36] (03Merged) 10jenkins-bot: Set JPGs to be thumbnailed via VIPS if md5 starts with given config [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1339064 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup-claude) [16:44:22] !log cdobbins@cumin1004 conftool action : set/pooled=yes; selector: name=ncredir7003 [16:46:31] !log cdobbins@cumin1004 conftool action : set/pooled=yes; selector: name=ncredir7003.magru.wmnet [16:49:29] (03PS2) 10DCausse: search: add semantic-highlighter to liftwing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341259 (https://phabricator.wikimedia.org/T433872) [16:51:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [16:52:33] !log jclark@cumin1003 START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie [16:52:53] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12329753 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jclark@cumin1003 for host zuul1006.eqiad.wmnet with OS t... [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1700) [17:05:51] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:05:52] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:06:06] (03PS1) 10Ladsgroup: thumbor: Use VIPS for large JPGs if the container starts with a [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342312 (https://phabricator.wikimedia.org/T220171) [17:07:14] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [17:07:17] !log jclark@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage [17:09:39] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:10:01] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:10:03] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:10:11] (03PS2) 10Ladsgroup: thumbor: Use VIPS for large JPGs if the container starts with a [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342312 (https://phabricator.wikimedia.org/T220171) [17:10:31] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:10:32] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:11:44] (03PS6) 10Jasmine: apache: Configure wikimedia.org redirect in apache, remove from redirects.dat [puppet] - 10https://gerrit.wikimedia.org/r/1333298 (https://phabricator.wikimedia.org/T427929) [17:13:21] !log jclark@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1006.eqiad.wmnet with reason: host reimage [17:14:18] !log jhancock@cumin2003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:15:58] (03CR) 10Ladsgroup: [C:03+2] thumbor: Use VIPS for large JPGs if the container starts with a [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342312 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup) [17:17:58] (03CR) 10Jasmine: apache: Configure wikimedia.org redirect in apache, remove from redirects.dat (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1333298 (https://phabricator.wikimedia.org/T427929) (owner: 10Jasmine) [17:18:27] (03Merged) 10jenkins-bot: thumbor: Use VIPS for large JPGs if the container starts with a [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342312 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup) [17:20:19] (03PS6) 10Bking: Kerberos: promote `krb1004` to Kerberos primary [puppet] - 10https://gerrit.wikimedia.org/r/1341330 (https://phabricator.wikimedia.org/T437937) [17:20:45] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [17:20:55] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [17:22:31] (03CR) 10Bking: [C:03+2] Kerberos: promote `krb1004` to Kerberos primary [puppet] - 10https://gerrit.wikimedia.org/r/1341330 (https://phabricator.wikimedia.org/T437937) (owner: 10Bking) [17:22:44] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install 10 new ceph nodes - https://phabricator.wikimedia.org/T438213 (10RobH) 03NEW [17:23:07] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:23:08] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:24:06] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install 10 new ceph nodes - https://phabricator.wikimedia.org/T438213#12329997 (10RobH) a:03BTullis @btullis, Please update this task description, checklists, and title with the new hostnames, as well as the racking details section. Pleas... [17:24:13] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install 10 new ceph nodes - https://phabricator.wikimedia.org/T438213#12330003 (10RobH) [17:27:32] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:27:33] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:29:49] !log jclark@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003" [17:30:35] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:30:36] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:32:53] jclark@cumin1003 reimage (PID 369694) is awaiting input [17:35:37] !log jclark@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1003" [17:35:39] !log jclark@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1006.eqiad.wmnet with OS trixie [17:35:53] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12330074 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jclark@cumin1003 for host zuul1006.eqiad.wmnet with OS trixi... [17:37:43] (03PS1) 10Ladsgroup: thumbor: Add jpg to vips [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342316 (https://phabricator.wikimedia.org/T220171) [17:38:21] PROBLEM - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:40:22] (03CR) 10Ladsgroup: [C:03+2] thumbor: Add jpg to vips [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342316 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup) [17:40:53] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12330084 (10Jclark-ctr) Cleared lvms and reimaged server [17:43:00] (03Merged) 10jenkins-bot: thumbor: Add jpg to vips [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342316 (https://phabricator.wikimedia.org/T220171) (owner: 10Ladsgroup) [17:44:07] vriley@cumin1003 provision (PID 375894) is awaiting input [17:44:33] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:44:34] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:45:20] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:45:21] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:46:08] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [17:46:43] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [17:46:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [17:51:16] !log vriley@cumin2003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:00:05] jnuche and dduvall: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T1800). [18:04:02] !log vriley@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:09:35] (03CR) 10Dzahn: [C:03+2] add attribution.wikimedia.org to miscweb tlsExtraSANs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341334 (https://phabricator.wikimedia.org/T437635) (owner: 10Dzahn) [18:10:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 8.609% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [18:11:30] (03PS1) 10CDanis: Fix known-client creation [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1342322 [18:11:58] (03CR) 10CDanis: [V:03+2 C:03+2] Fix known-client creation [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1342322 (owner: 10CDanis) [18:12:15] !log cdanis@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003" [18:12:17] !log cdanis@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003 [18:12:17] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12330271 (10VRiley-WMF) [18:13:11] !log cdanis@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: fix known-client creation - cdanis@cumin1003 [18:13:12] !log cdanis@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "fix known-client creation - cdanis@cumin1003" [18:16:29] !log dzahn@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [18:17:06] !log dzahn@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [18:17:24] !log dzahn@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [18:17:47] !log dzahn@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [18:17:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [18:18:21] !log k8s/miscweb: admin_ng deploy: creating namespace for attribution.wikimedia.org T437635 [18:18:24] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:18:24] T437635: Add attribution.wikimedia.org to miscweb - https://phabricator.wikimedia.org/T437635 [18:18:37] !log dzahn@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [18:20:36] !log dzahn@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [18:20:49] !log dzahn@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [18:21:25] !log robh@cumin1003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:21:26] !log robh@cumin1003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:21:45] !log dzahn@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [18:22:12] !log robh@cumin2003 START - Cookbook sre.hosts.provision for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:25:33] robh@cumin2003 provision (PID 2573440) is awaiting input [18:26:29] (03Abandoned) 10Ryan Kemper: Add sre.hadoop.reboot-coordinators cookbook [cookbooks] - 10https://gerrit.wikimedia.org/r/1261271 (https://phabricator.wikimedia.org/T421285) (owner: 10Ryan Kemper) [18:26:52] !log robh@cumin2003 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be1099.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:27:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [18:35:39] (03CR) 10Dzahn: [C:03+1] rsyslog: send php8.5-fpm logs to Logstash [puppet] - 10https://gerrit.wikimedia.org/r/1341351 (https://phabricator.wikimedia.org/T435393) (owner: 10Southparkfan) [18:35:40] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [18:37:35] 06SRE, 06Infrastructure-Foundations: Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T438229 (10bking) 03NEW [18:37:41] (03CR) 10Ayounsi: "From a quick look the redis-py update seems like a big jump with listed breaking changes, but no idea if they apply to us." [software/netbox-deploy] - 10https://gerrit.wikimedia.org/r/1342301 (owner: 10Ayounsi) [18:37:43] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [18:37:52] 06SRE, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T438229#12330384 (10bking) [18:38:21] RECOVERY - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:42:52] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [18:44:57] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [18:46:57] 10ops-eqiad, 06DC-Ops: Cookbook issues with cumin1003 - https://phabricator.wikimedia.org/T438230 (10VRiley-WMF) 03NEW [18:47:50] (03PS1) 10Dwisehaupt: wmnet: Adjust fundraisingdb-read for host reboot [dns] - 10https://gerrit.wikimedia.org/r/1342331 [18:48:53] 10ops-eqiad, 06DC-Ops: Cookbook issues with cumin1003 - https://phabricator.wikimedia.org/T438230#12330424 (10VRiley-WMF) [18:49:22] (03PS1) 10Bking: Kerberos: Prepare a VM for deployment as krb replica [puppet] - 10https://gerrit.wikimedia.org/r/1342333 (https://phabricator.wikimedia.org/T438229) [18:52:12] 10ops-eqiad, 06DC-Ops: Cookbook issues with cumin1003 - https://phabricator.wikimedia.org/T438230#12330440 (10VRiley-WMF) [18:53:46] (03PS1) 10Ahmon Dancy: deployment_server: run scap clean-images daily [puppet] - 10https://gerrit.wikimedia.org/r/1342337 (https://phabricator.wikimedia.org/T438148) [18:53:53] (03CR) 10Jgreen: [C:03+1] wmnet: Adjust fundraisingdb-read for host reboot [dns] - 10https://gerrit.wikimedia.org/r/1342331 (owner: 10Dwisehaupt) [18:54:05] 10ops-eqiad, 06DC-Ops: Cookbook issues with cumin1003 - https://phabricator.wikimedia.org/T438230#12330445 (10VRiley-WMF) [18:55:14] (03CR) 10Dwisehaupt: [C:03+2] wmnet: Adjust fundraisingdb-read for host reboot [dns] - 10https://gerrit.wikimedia.org/r/1342331 (owner: 10Dwisehaupt) [18:55:22] !log dwisehaupt@dns1005 START - running authdns-update [18:55:32] !log vriley@cumin2003 START - Cookbook sre.hosts.provision for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [18:56:47] (03PS1) 10Cklimas: Reapply "Tell VisualEditor about the app web edit tags", modified [extensions/MobileApp] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342339 [18:57:40] !log dwisehaupt@dns1005 END - running authdns-update [18:58:45] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 16 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [extensions/MobileApp] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342339 (owner: 10Cklimas) [18:58:57] 10SRE-swift-storage, 06Commons: HTTP 503 Backend fetch failed while editing Commons - https://phabricator.wikimedia.org/T307338#12330463 (10Aklapper) 05Open→03Declined It seems the issue was transient; closing this as part of regular task cleanup. Sorry, these ones are particularly hard to figure out,... [18:59:38] 10ops-eqiad, 06DC-Ops: Cookbook issues with cumin1003 - https://phabricator.wikimedia.org/T438230#12330479 (10VRiley-WMF) a:03elukey Assigning to Luca to take a look at this ticket. [19:00:18] (03PS2) 10Ahmon Dancy: deployment_server: run scap clean-images daily [puppet] - 10https://gerrit.wikimedia.org/r/1342337 (https://phabricator.wikimedia.org/T438148) [19:00:54] (03CR) 10Ahmon Dancy: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342337 (https://phabricator.wikimedia.org/T438148) (owner: 10Ahmon Dancy) [19:01:01] 10ops-eqiad, 06DC-Ops, 06Infrastructure-Foundations: Cookbook issues with cumin1003 - https://phabricator.wikimedia.org/T438230#12330494 (10RobH) [19:06:07] (03PS1) 10Andrew Bogott: openstack_magnum_capi_worker_stats_exporter: correct deployment in codfw1dev [puppet] - 10https://gerrit.wikimedia.org/r/1342344 (https://phabricator.wikimedia.org/T429557) [19:06:23] !log vriley@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host db1245.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [19:07:26] (03PS1) 10Dwisehaupt: wmnet: Set fundraisingdb-read back to frdb1008 [dns] - 10https://gerrit.wikimedia.org/r/1342348 [19:08:33] (03CR) 10Jgreen: [C:03+1] wmnet: Set fundraisingdb-read back to frdb1008 [dns] - 10https://gerrit.wikimedia.org/r/1342348 (owner: 10Dwisehaupt) [19:08:56] 10ops-eqiad, 06DC-Ops, 06Infrastructure-Foundations: Cookbook issues with cumin1003 - https://phabricator.wikimedia.org/T438230#12330531 (10VRiley-WMF) 05Open→03Invalid Closing ticket. 1004 is the new host. [19:09:44] FIRING: KubernetesDeploymentUnavailableReplicas: ... [19:09:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [19:09:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [19:13:46] 06SRE, 06Infrastructure-Foundations, 06Release-Engineering-Team (Radar): apt-get broken in docker-registry.wikimedia.org/bullseye:20260830 - https://phabricator.wikimedia.org/T437069#12330543 (10dancy) Any news on rebuilt bullseye base images? [19:15:23] (03CR) 10Andrew Bogott: [C:03+2] openstack_magnum_capi_worker_stats_exporter: correct deployment in codfw1dev [puppet] - 10https://gerrit.wikimedia.org/r/1342344 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [19:17:11] (03PS1) 10Andrew Bogott: Add a few alerts about magnum uptime [alerts] - 10https://gerrit.wikimedia.org/r/1342349 (https://phabricator.wikimedia.org/T429557) [19:17:50] (03CR) 10Dwisehaupt: [C:03+2] wmnet: Set fundraisingdb-read back to frdb1008 [dns] - 10https://gerrit.wikimedia.org/r/1342348 (owner: 10Dwisehaupt) [19:18:01] !log dwisehaupt@dns1005 START - running authdns-update [19:19:44] RESOLVED: KubernetesDeploymentUnavailableReplicas: ... [19:19:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [19:19:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [19:20:19] !log dwisehaupt@dns1005 END - running authdns-update [19:20:24] (03CR) 10CI reject: [V:04-1] Add a few alerts about magnum uptime [alerts] - 10https://gerrit.wikimedia.org/r/1342349 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [19:21:11] (03CR) 10Dzahn: [C:03+2] "deployed admin_ng" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341334 (https://phabricator.wikimedia.org/T437635) (owner: 10Dzahn) [19:22:13] (03PS2) 10Andrew Bogott: Add a few alerts about magnum uptime [alerts] - 10https://gerrit.wikimedia.org/r/1342349 (https://phabricator.wikimedia.org/T429557) [19:23:42] (03CR) 10Ryan Kemper: [C:03+1] Kerberos: Prepare a VM for deployment as krb replica [puppet] - 10https://gerrit.wikimedia.org/r/1342333 (https://phabricator.wikimedia.org/T438229) (owner: 10Bking) [19:24:22] (03CR) 10Dzahn: [C:03+2] ci: set ci-build-images image and update creates [puppet] - 10https://gerrit.wikimedia.org/r/1333796 (https://phabricator.wikimedia.org/T436775) (owner: 10Hashar) [19:24:45] (03CR) 10CI reject: [V:04-1] Add a few alerts about magnum uptime [alerts] - 10https://gerrit.wikimedia.org/r/1342349 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [19:25:38] (03CR) 10Dzahn: [C:03+1] zookeeper::server: Use default zookeeper CLASSPATH when using 3.4 [puppet] - 10https://gerrit.wikimedia.org/r/1322828 (https://phabricator.wikimedia.org/T424266) (owner: 10Scott French) [19:28:03] FIRING: [2x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:28:36] (03PS1) 10Dzahn: openstack: enable zookeeper logging in codfw1dev [puppet] - 10https://gerrit.wikimedia.org/r/1342354 [19:29:37] (03CR) 10Ryan Kemper: [C:03+1] "this is simple enough that PCC isn't strictly necessary, but I'll do a quick PCC run just for posterity's sake" [puppet] - 10https://gerrit.wikimedia.org/r/1342333 (https://phabricator.wikimedia.org/T438229) (owner: 10Bking) [19:29:59] (03PS2) 10Ryan Kemper: Kerberos: Prepare a VM for deployment as krb replica [puppet] - 10https://gerrit.wikimedia.org/r/1342333 (https://phabricator.wikimedia.org/T438229) (owner: 10Bking) [19:30:06] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1342333 (https://phabricator.wikimedia.org/T438229) (owner: 10Bking) [19:33:03] RESOLVED: [2x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#titan1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:33:20] (03PS3) 10Andrew Bogott: Add a few alerts about magnum uptime [alerts] - 10https://gerrit.wikimedia.org/r/1342349 (https://phabricator.wikimedia.org/T429557) [19:43:23] (03PS4) 10Andrew Bogott: Add a few alerts about magnum uptime [alerts] - 10https://gerrit.wikimedia.org/r/1342349 (https://phabricator.wikimedia.org/T429557) [19:48:53] (03PS1) 10Medelius: Make VE the default editor on enwiki desktop [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342362 (https://phabricator.wikimedia.org/T436574) [19:50:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:54:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [19:56:41] (03PS5) 10Andrew Bogott: Add a few alerts about magnum uptime [alerts] - 10https://gerrit.wikimedia.org/r/1342349 (https://phabricator.wikimedia.org/T429557) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor When your hammer is PHP, everything starts looking like a thumb. Rise for UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T2000). [20:00:05] cscott and cklimas/kemayo: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:02:29] o/ [20:02:45] hello! [20:04:08] If cscott isn't here yet, I don't mind getting out patch deployed. [20:04:36] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy1003 using scap backport" [extensions/MobileApp] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342339 (owner: 10Cklimas) [20:06:23] (03CR) 10Bking: [C:03+2] Kerberos: Prepare a VM for deployment as krb replica [puppet] - 10https://gerrit.wikimedia.org/r/1342333 (https://phabricator.wikimedia.org/T438229) (owner: 10Bking) [20:09:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [20:12:55] (03Merged) 10jenkins-bot: Reapply "Tell VisualEditor about the app web edit tags", modified [extensions/MobileApp] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342339 (owner: 10Cklimas) [20:13:21] !log kemayo@deploy1003 Started scap sync-world: Backport for [[gerrit:1342339|Reapply "Tell VisualEditor about the app web edit tags", modified]] [20:14:12] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12330829 (10Jclark-ctr) @cmooney Unplugged and reseated B7, and the link came back up. I also connected the links between the D8 and D1 spines with the A1 sp... [20:17:12] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:19:19] (03PS2) 10Subramanya Sastry: Parsoid Read Views: Enable on 61 wikiquote wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341949 (https://phabricator.wikimedia.org/T437917) [20:21:45] It's convenient that this is taking ages, because I got called aside to talk to a HVAC contractor. [20:21:46] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [20:24:46] (03CR) 10Thcipriani: [C:03+1] deployment_server: run scap clean-images daily [puppet] - 10https://gerrit.wikimedia.org/r/1342337 (https://phabricator.wikimedia.org/T438148) (owner: 10Ahmon Dancy) [20:31:35] (03CR) 10DLynch: [C:03+1] Make VE the default editor on enwiki desktop [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342362 (https://phabricator.wikimedia.org/T436574) (owner: 10Medelius) [20:33:09] !log kemayo@deploy1003 cklimas, kemayo: Backport for [[gerrit:1342339|Reapply "Tell VisualEditor about the app web edit tags", modified]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:35:10] (03PS2) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-09-08-173612 to 2026-09-16-192603 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342268 (https://phabricator.wikimedia.org/T322056) [20:35:48] I have the mwdebug extension active in my browser, but do I need to set a specific option next to the on switch in that? [20:37:09] !log kemayo@deploy1003 cklimas, kemayo: Continuing with deployment [20:37:50] cklimas: Sorry, didn't notice your question here. k8s-mwdebug (the default) would be fine. [20:38:18] cklimas: I tested it, and votewiki didn't melt and https://test2.wikipedia.org/w/index.php?diff=622427 has the tag applied after I manually set things up. [20:38:43] Kemayo: ah, ok. I tried on testwiki and though pages are loading fine, I couldn't get the tag to show up in edit history. [20:39:41] Kemayo: aha, the tag is also now working for me. [20:39:54] cklimas: If you've got the mwdebug extension set up correctly then on that diff I linked you'll see `app edit (web)` in the tags. Otherwise it'll just be `app web edit other`. [20:40:24] Kemayo: yep, I see it there. [20:40:35] Excellent. [20:42:38] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12330903 (10VRiley-WMF) Ran a reprovisioning on this server while we wait for Dell to come back (working out a warrenty issue with them on this server). Will monitor this and see if it crashe... [20:48:13] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12330933 (10VRiley-WMF) Thank you John, @Dzahn This should be ready for you at this point. [20:49:17] !log kemayo@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342339|Reapply "Tell VisualEditor about the app web edit tags", modified]] (duration: 35m 55s) [20:49:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [20:54:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.73% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:57:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:00:05] Deploy window Wikifunctions Services UTC Late (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T2100) [21:00:12] Excellent. [21:02:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:02:22] !log vriley@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be1098.eqiad.wmnet with OS trixie [21:02:35] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12330974 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin2003 for host ms-be1098.eqiad.wmnet wi... [21:04:04] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-09-08-173612 to 2026-09-16-192603 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342268 (https://phabricator.wikimedia.org/T322056) (owner: 10Jforrester) [21:06:37] (03Merged) 10jenkins-bot: wikifunctions: Upgrade orchestrator from 2026-09-08-173612 to 2026-09-16-192603 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342268 (https://phabricator.wikimedia.org/T322056) (owner: 10Jforrester) [21:08:14] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:08:38] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:09:14] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:09:50] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:09:57] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [21:10:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.38% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:10:32] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:13:31] James_F: no pressure, but if you should happen to finish with your window early, let me know, I'd like to bully poolcounter a little [21:14:12] rzl: Done ish now. We’re pondering one final evaluator upgrade, but that relies on GitLab CI going faster than treacle. [21:15:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.55% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:16:01] James_F: ah got you -- mind if I sneak out a config patch in between? happy to work around you obviously [21:16:07] rzl: Go for it. [21:16:09] rad, thanks [21:17:30] (03CR) 10TrainBranchBot: [C:03+2] "Approved by rzl@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342089 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [21:17:40] !log vriley@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage [21:18:29] (03Merged) 10jenkins-bot: Depool poolcounter[1006,2005] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342089 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [21:18:56] !log rzl@deploy1003 Started scap sync-world: Backport for [[gerrit:1342089|Depool poolcounter[1006,2005] for reboot (T435163)]] [21:22:59] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12331040 (10Dzahn) Thank you @VRiley-WMF and @Jclark-ctr . I confirm I can ssh to zuul1006 and puppet is running with our team-specific inse... [21:23:45] !log vriley@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1098.eqiad.wmnet with reason: host reimage [21:25:12] !log rzl@deploy1003 rzl: Backport for [[gerrit:1342089|Depool poolcounter[1006,2005] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:26:11] !log rzl@deploy1003 rzl: Continuing with deployment [21:28:06] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12331047 (10VRiley-WMF) 05Open→03Resolved [21:31:46] (03PS1) 10Jforrester: wikifunctions: Upgrade evaluators from 2026-09-14-201648 to 2026-09-16-212211 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342379 (https://phabricator.wikimedia.org/T433459) [21:32:24] rzl: OK for me to do a service deploy alongside you? Should be a no-op for real-world traffic, just book-keeping. [21:32:38] yep, go for it [21:32:44] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade evaluators from 2026-09-14-201648 to 2026-09-16-212211 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342379 (https://phabricator.wikimedia.org/T433459) (owner: 10Jforrester) [21:32:49] !log rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342089|Depool poolcounter[1006,2005] for reboot (T435163)]] (duration: 13m 53s) [21:35:07] (03Merged) 10jenkins-bot: wikifunctions: Upgrade evaluators from 2026-09-14-201648 to 2026-09-16-212211 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342379 (https://phabricator.wikimedia.org/T433459) (owner: 10Jforrester) [21:35:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:35:29] !log rzl@cumin2003 START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet [21:36:58] PROBLEM - Blazegraph Port for wdqs-blazegraph on wdqs1022 is CRITICAL: connect to address 127.0.0.1 and port 9999: Connection refused https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [21:37:47] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:37:58] RECOVERY - Blazegraph Port for wdqs-blazegraph on wdqs1022 is OK: TCP OK - 0.000 second response time on 127.0.0.1 port 9999 https://wikitech.wikimedia.org/wiki/Wikidata_query_service/Runbook [21:38:32] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:38:43] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:39:19] !log rzl@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet [21:39:38] !log rzl@cumin2003 START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet [21:40:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.04% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:40:17] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:40:28] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [21:40:55] !log vriley@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003" [21:41:14] !log vriley@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin2003" [21:41:15] !log vriley@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1098.eqiad.wmnet with OS trixie [21:41:29] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12331101 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin2003 for host ms-be1098.eqiad.wmnet with O... [21:42:05] !log vriley@cumin2003 START - Cookbook sre.hosts.reimage for host ms-be1099.eqiad.wmnet with OS trixie [21:42:17] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, and 2 others: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12331103 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin2003 for host ms-be1099.eqiad.wmnet wi... [21:43:27] !log rzl@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet [21:44:35] (03CR) 10TrainBranchBot: [C:03+2] "Approved by rzl@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [21:44:50] this backport should be done just in time for the 22:00 window, if it's in use [21:45:01] (03CR) 10CI reject: [V:04-1] Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [21:45:05] (03PS4) 10RLazarus: Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) [21:45:20] or at least it will be if I remember to rebase [21:45:43] (03CR) 10TrainBranchBot: "Approved by rzl@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [21:46:56] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:46:59] (03Merged) 10jenkins-bot: Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342090 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [21:47:19] !log rzl@deploy1003 Started scap sync-world: Backport for [[gerrit:1342090|Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] [21:49:59] 10ops-eqiad, 06DC-Ops, 06Machine-Learning-Team: Q#:rack/setup/install X - https://phabricator.wikimedia.org/T438253 (10RobH) 03NEW [21:50:04] (All done.) [21:50:14] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, September 17 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#dep" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1331830 (https://phabricator.wikimedia.org/T436398) (owner: 10C. Scott Ananian) [21:50:20] 10ops-eqiad, 06DC-Ops, 06Machine-Learning-Team: Q2:rack/setup/install ml-serve1017 - https://phabricator.wikimedia.org/T438253#12331137 (10RobH) [21:50:50] 10ops-eqiad, 06DC-Ops, 06Machine-Learning-Team: Q2:rack/setup/install ml-serve1017 - https://phabricator.wikimedia.org/T438253#12331139 (10RobH) [21:51:09] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, September 17 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#dep" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1340216 (https://phabricator.wikimedia.org/T328012) (owner: 10C. Scott Ananian) [21:51:38] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, September 17 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#dep" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341949 (https://phabricator.wikimedia.org/T437917) (owner: 10Subramanya Sastry) [21:51:43] 10ops-eqiad, 06DC-Ops, 06Machine-Learning-Team: Q2:rack/setup/install ml-serve1017 - https://phabricator.wikimedia.org/T438253#12331144 (10RobH) a:03isarantopoulos @isarantopoulos, Please double check the racking details for this, as I made assumptions based on previous gpu hosts. Please update the site.... [21:51:49] !log rzl@deploy1003 rzl: Backport for [[gerrit:1342090|Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:51:50] 10ops-eqiad, 06DC-Ops, 06Machine-Learning-Team: Q2:rack/setup/install ml-serve1017 - https://phabricator.wikimedia.org/T438253#12331148 (10RobH) [21:52:01] (03PS1) 10JHathaway: WIP: stdlib [puppet] - 10https://gerrit.wikimedia.org/r/1342382 [21:52:19] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342382 (owner: 10JHathaway) [21:52:33] !log rzl@deploy1003 rzl: Continuing with deployment [21:53:55] (03PS1) 10Jdlrobson: DonorIdentification: Confirm before unlinking donor status in preferences [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342383 (https://phabricator.wikimedia.org/T436698) [21:56:10] (03CR) 10CI reject: [V:04-1] WIP: stdlib [puppet] - 10https://gerrit.wikimedia.org/r/1342382 (owner: 10JHathaway) [21:56:58] !log rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342090|Repool poolcounter[1006,2005]; depool poolcounter[1007,2006] for reboot (T435163)]] (duration: 09m 39s) [21:58:08] done scapping for now; I'll do the other two reboots then put out another config patch to put things back to normal after the readers window is done [21:58:23] rzl: FYI they rarely use it. [21:58:40] General advice is “if they’re not here in the first five minutes, assume they’re not coming". [21:58:41] yeah I know, but it'll take me a few minutes anyway :D by then we'll know [21:58:45] :-) [21:58:50] (hoping to use the window in 2m - let me know when you are done) [21:59:07] Jdlrobson: I think it's all yours! [22:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260916T2200) [22:02:03] !log rzl@cumin2003 START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet [22:02:03] thanks rzl [22:02:20] (you'll see me rebooting some stuff but that's compatible) [22:02:39] as I say, whenever you're done I'd like to do one more config patch, but no urgency [22:02:49] ok! [22:06:02] !log rzl@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet [22:06:13] (03PS1) 10Jdlrobson: Make learn more link to new window [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342385 (https://phabricator.wikimedia.org/T438252) [22:06:37] !log rzl@cumin2003 START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet [22:06:46] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342385 (https://phabricator.wikimedia.org/T438252) (owner: 10Jdlrobson) [22:06:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342383 (https://phabricator.wikimedia.org/T436698) (owner: 10Jdlrobson) [22:09:33] (03CR) 10CI reject: [V:04-1] Make learn more link to new window [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342385 (https://phabricator.wikimedia.org/T438252) (owner: 10Jdlrobson) [22:10:04] (03Merged) 10jenkins-bot: Make learn more link to new window [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342385 (https://phabricator.wikimedia.org/T438252) (owner: 10Jdlrobson) [22:10:06] (03Merged) 10jenkins-bot: DonorIdentification: Confirm before unlinking donor status in preferences [extensions/WikimediaCustomizations] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1342383 (https://phabricator.wikimedia.org/T436698) (owner: 10Jdlrobson) [22:10:15] !log rzl@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet [22:10:46] (03PS2) 10Dduvall: buildx: New driver for building images via Buildx/BuildKit [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958) [22:12:27] !log jdlrobson@deploy1003 Started scap sync-world: Backport for [[gerrit:1342383|DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385|Make learn more link to new window (T438252)]] [22:12:32] T436698: Donor preference should have a confirm step - https://phabricator.wikimedia.org/T436698 [22:12:32] T438252: Learn more should open in a new window - https://phabricator.wikimedia.org/T438252 [22:15:18] (03CR) 10CI reject: [V:04-1] buildx: New driver for building images via Buildx/BuildKit [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958) (owner: 10Dduvall) [22:27:20] 06SRE, 10SRE-Access-Requests: Requesting access to releasers-mobile for LPetty-WMF - https://phabricator.wikimedia.org/T437662#12331284 (10LPetty-WMF) Hello, yesterday I tried uploading an apk file and received the following error message in the terminal **Permission denied (publickey). Connection closed... [22:29:04] 06SRE, 10SRE-swift-storage, 10MediaWiki-extensions-Score, 06Reader Experience Team: Add cache key information to metadata json - https://phabricator.wikimedia.org/T257093#12331285 (10Jonathanischoice) I've moved my patch to a new bug T438258 rather than clog up this one. [22:33:06] !log jdlrobson@deploy1003 jdlrobson: Backport for [[gerrit:1342383|DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385|Make learn more link to new window (T438252)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [22:33:11] T436698: Donor preference should have a confirm step - https://phabricator.wikimedia.org/T436698 [22:33:12] T438252: Learn more should open in a new window - https://phabricator.wikimedia.org/T438252 [22:33:48] vriley@cumin2003 reimage (PID 2614247) is awaiting input [22:35:30] !log jdlrobson@deploy1003 jdlrobson: Continuing with deployment [22:35:39] syncing now... then i'll be done rzl [22:36:45] rad [22:40:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:46:56] FIRING: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [22:47:48] !log jdlrobson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342383|DonorIdentification: Confirm before unlinking donor status in preferences (T436698)]], [[gerrit:1342385|Make learn more link to new window (T438252)]] (duration: 35m 21s) [22:47:53] T436698: Donor preference should have a confirm step - https://phabricator.wikimedia.org/T436698 [22:47:53] T438252: Learn more should open in a new window - https://phabricator.wikimedia.org/T438252 [22:48:48] all done rzl [22:49:02] thanks! going ahead [22:49:19] (03PS4) 10RLazarus: Repool poolcounter[1007,2006] [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342091 (https://phabricator.wikimedia.org/T435163) [22:49:30] (03CR) 10TrainBranchBot: [C:03+2] "Approved by rzl@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342091 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [22:50:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.55% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:50:21] (03Merged) 10jenkins-bot: Repool poolcounter[1007,2006] [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342091 (https://phabricator.wikimedia.org/T435163) (owner: 10RLazarus) [22:50:43] !log rzl@deploy1003 Started scap sync-world: Backport for [[gerrit:1342091|Repool poolcounter[1007,2006] (T435163)]] [22:53:54] 22:52:41 Non-zero exit status (1) from (cd /srv/deployment-charts/helmfile.d/dse-k8s-services/mediawiki-dumps-legacy && helm --kubeconfig /etc/kubernetes/mediawiki-dumps-legacy-deploy-dse-k8s-eqiad.config ls -a -o json --max=1000) [22:53:54] 22:52:41 stderr: Error: Kubernetes cluster unreachable: Get "https://dse-k8s-ctrl.svc.eqiad.wmnet:6443/version": dial tcp 10.2.2.73:6443: connect: connection refused [22:53:57] looking [22:55:46] hrm, dse-k8s-ctrl seems fine to me. let's see what I can retry [22:56:40] I'm guessing I can't just scap backport again because it's already merged, but if that fails I can try rerunning the scap sync-world --backport invocation [22:57:34] oh neat, I *can* just redo the scap backport, that rules [22:58:17] (logmsgbot is out to lunch but, started scap sync-world again at 22:57:12) [23:01:56] RESOLVED: TransportLinksInUsageNoRedundancy: eqsin inbound transports link usage is at redundancy capacity - https://wikitech.wikimedia.org/wiki/Network_monitoring#TransportLinksInUsageNoRedundancy - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/bandwidth-site-totals?var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DTransportLinksInUsageNoRedundancy [23:02:09] still no logmsgbot: rzl: Backport for [[gerrit:1342091|Repool poolcounter[1007,2006] (T435163)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there., and rzl: Continuing with deployment [23:03:52] oh I see, logmsgbot actually crashed because scap tried to !log a python traceback with newlines in it, and the irc client said no. so that's not great, and also, why didn't it come back? [23:10:05] !log rzl@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342091|Repool poolcounter[1007,2006] (T435163)]] (duration: 11m 09s) [23:10:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [23:10:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.86% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [23:11:29] correction, not a python traceback, but the output here had a newline in it. https://gitlab.wikimedia.org/repos/releng/scap/-/blob/master/scap/cli.py#L761 I'll file a bug, and in the meantime, through deploying [23:11:56] probably two bugs, scap shouldn't send that newline and logmsgbot shouldn't explode if it sees it [23:15:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.86% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [23:15:57] 06SRE: tcpircbot shouldn't crash if it receives a message with a newline - https://phabricator.wikimedia.org/T438269 (10RLazarus) 03NEW [23:23:58] 06SRE: tcpircbot shouldn't crash if it receives a message with a newline - https://phabricator.wikimedia.org/T438269#12331510 (10RLazarus) Also filed T438270 for the scap side. [23:25:41] (03PS3) 10Dduvall: buildx: New driver for building images via Buildx/BuildKit [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958) [23:30:36] (03CR) 10CI reject: [V:04-1] buildx: New driver for building images via Buildx/BuildKit [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958) (owner: 10Dduvall) [23:33:21] (03CR) 10Dduvall: "@glavagetto@wikimedia.org note I have not yet added a ton of tests for this new implementation but it should be good for a preliminary rev" [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958) (owner: 10Dduvall) [23:36:12] (03PS4) 10Dduvall: buildx: New driver for building images via Buildx/BuildKit [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958) [23:39:05] (03PS5) 10Dduvall: buildx: New driver for building images via Buildx/BuildKit [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958) [23:40:58] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1342389 [23:40:58] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1342389 (owner: 10TrainBranchBot) [23:49:31] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1342389 (owner: 10TrainBranchBot)