[00:23:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.79% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:28:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:41:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.79% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:46:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:47:38] (03PS1) 10Tryvix1509: throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333338 (https://phabricator.wikimedia.org/T436672) [00:48:16] 10ops-magru: magru: faulty link between cr1-magru and asw1-b3-magru - https://phabricator.wikimedia.org/T436675#12278457 (10RobH) a:03ayounsi remote hands update: > The fiber connections were cleaned on both the router and switch sides. > The QSFP module in switch asw1-b3-magru (port et-0/0/48) was replaced w... [00:49:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:59:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.41% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:00:02] (03PS1) 10Tryvix1509: InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333343 (https://phabricator.wikimedia.org/T436712) [01:00:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:02:20] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333338 (https://phabricator.wikimedia.org/T436672) (owner: 10Tryvix1509) [01:03:19] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333343 (https://phabricator.wikimedia.org/T436712) (owner: 10Tryvix1509) [01:06:52] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [01:10:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:11:12] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1333344 [01:11:12] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1333344 (owner: 10TrainBranchBot) [01:11:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.35% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:20:38] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1333344 (owner: 10TrainBranchBot) [01:21:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:24:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.83% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:29:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.42% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:50:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:55:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:00:45] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:08:28] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 43s) [02:08:38] FIRING: [2x] GnmiInterfaceCountersDrop: cr2-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [02:11:42] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:20:38] (03PS1) 10EggRoll97: Add two wmf groups to privileged status [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333390 (https://phabricator.wikimedia.org/T436734) [02:37:31] FIRING: Traffic bill over quota: Alert for device cr2-magru.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [02:55:18] (03PS2) 10EggRoll97: Add two wmf groups to privileged status [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333390 (https://phabricator.wikimedia.org/T436734) [02:56:18] (03CR) 10CI reject: [V:04-1] Add two wmf groups to privileged status [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333390 (https://phabricator.wikimedia.org/T436734) (owner: 10EggRoll97) [02:57:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 19.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:57:31] RESOLVED: Traffic bill over quota: Alert for device cr2-magru.wikimedia.org - Traffic bill over quota - https://alerts.wikimedia.org/?q=alertname%3DTraffic+bill+over+quota [02:59:03] (03PS3) 10EggRoll97: Add two wmf groups to privileged status [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333390 (https://phabricator.wikimedia.org/T436734) [03:07:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:09:32] (03PS1) 10Tim Starling: Set a short CC:max-age on cacheable REST responses [extensions/Produnto] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333423 [03:11:26] (03CR) 10CI reject: [V:04-1] Set a short CC:max-age on cacheable REST responses [extensions/Produnto] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333423 (owner: 10Tim Starling) [03:21:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:26:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:30:39] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [03:38:44] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [03:39:59] FIRING: CoreRouterInterfaceDown: Core router interface down - cr1-magru:et-0/0/1 (Core: asw1-b3-magru:et-0/0/48) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [03:42:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:42:54] FIRING: [4x] CoreBGPDown: Core BGP session down between asw1-b3-magru and cr1-magru (195.200.68.142) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [03:43:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [03:44:44] 10SRE-swift-storage, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): RdfStreamingUpdaterSpaceUsageTooHigh - https://phabricator.wikimedia.org/T431506#12278687 (10RKemper) 05In progress→03Resolved Following up here: the thanos swift object-expirer was enabled on July 15 in https://gerrit.wikimedia.org/r/... [03:47:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [03:49:24] (03CR) 10Samwilson: "This doesn't run for me — it seems not to be able to find the standard logging lib any more, and is finding thumbor-plugins `logging` firs" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333246 (https://phabricator.wikimedia.org/T435609) (owner: 10Ladsgroup-claude) [03:52:46] (03CR) 10EggRoll97: [C:03+1] "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333338 (https://phabricator.wikimedia.org/T436672) (owner: 10Tryvix1509) [03:53:07] (03CR) 10EggRoll97: [C:03+1] "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333343 (https://phabricator.wikimedia.org/T436712) (owner: 10Tryvix1509) [04:16:03] (03PS1) 10Tim Starling: Updater: Normalize MW_VERSION [extensions/Produnto] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333470 (https://phabricator.wikimedia.org/T436741) [04:17:05] (03PS2) 10Tim Starling: Set a short CC:max-age on cacheable REST responses [extensions/Produnto] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333423 [04:21:15] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [extensions/Produnto] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333470 (https://phabricator.wikimedia.org/T436741) (owner: 10Tim Starling) [04:21:16] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [extensions/Produnto] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333423 (owner: 10Tim Starling) [04:23:51] (03Merged) 10jenkins-bot: Updater: Normalize MW_VERSION [extensions/Produnto] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333470 (https://phabricator.wikimedia.org/T436741) (owner: 10Tim Starling) [04:23:57] (03Merged) 10jenkins-bot: Set a short CC:max-age on cacheable REST responses [extensions/Produnto] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333423 (owner: 10Tim Starling) [04:24:59] !log tstarling@deploy1003 Started scap sync-world: Backport for [[gerrit:1333470|Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423|Set a short CC:max-age on cacheable REST responses]] [04:25:03] T436741: semver has opinions (in Produnto) - https://phabricator.wikimedia.org/T436741 [04:29:39] !log tstarling@deploy1003 tstarling: Backport for [[gerrit:1333470|Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423|Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [04:44:03] having trouble verifying the bugfix [04:53:47] !log tstarling@deploy1003 tstarling: Continuing with deployment [05:00:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:01:14] !log tstarling@deploy1003 Scap cancelled without rolling back. [05:01:49] !log tstarling@deploy1003 Started scap sync-world: Backport for [[gerrit:1333470|Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423|Set a short CC:max-age on cacheable REST responses]] [05:01:52] T436741: semver has opinions (in Produnto) - https://phabricator.wikimedia.org/T436741 [05:03:24] !log tstarling@deploy1003 tstarling: Backport for [[gerrit:1333470|Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423|Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [05:04:26] !log tstarling@deploy1003 tstarling: Continuing with deployment [05:06:31] !log tstarling@deploy1003 Finished scap sync-world: Backport for [[gerrit:1333470|Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423|Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s) [05:06:52] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [05:10:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:39:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [05:44:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [05:49:44] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [05:53:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.86% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [05:58:52] (03PS1) 10Krinkle: varnish: Refactor 19-normalize-thumbnail-url.vtc [puppet] - 10https://gerrit.wikimedia.org/r/1333533 (https://phabricator.wikimedia.org/T425216) [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T0600) [06:03:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.66% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:05:16] 10ops-eqiad, 06DBA, 06DC-Ops: db1228 crashed again - https://phabricator.wikimedia.org/T436743 (10Marostegui) 03NEW [06:05:18] 10ops-eqiad, 06DBA, 06DC-Ops: db1228 crashed again - https://phabricator.wikimedia.org/T436743#12278763 (10Marostegui) 05Open→03Resolved a:03Marostegui [06:08:38] FIRING: [2x] GnmiInterfaceCountersDrop: cr2-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [06:12:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.39% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:13:44] !log re-enable magru cr1/asw1-b3 link - T436675 [06:13:46] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:13:47] T436675: magru: faulty link between cr1-magru and asw1-b3-magru - https://phabricator.wikimedia.org/T436675 [06:14:22] RECOVERY - OSPF status on cr2-magru is OK: OSPFv2: 3/3 UP : OSPFv3: 3/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:16:12] (03PS1) 10Marostegui: mariadb: Move db1228 to s4 [puppet] - 10https://gerrit.wikimedia.org/r/1333545 (https://phabricator.wikimedia.org/T435892) [06:17:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.86% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:17:17] (03CR) 10Marostegui: [C:03+2] mariadb: Move db1228 to s4 [puppet] - 10https://gerrit.wikimedia.org/r/1333545 (https://phabricator.wikimedia.org/T435892) (owner: 10Marostegui) [06:17:17] bjensen, hnowlan, ^ I've re-enabled the ex-faulty magru link after one optic was replaced. Hopefully it's all good now. If not magru might show some issues in the next hours max. In that case I need to re disable it and ask remote hands to swap the optic on the other side. [06:17:54] RESOLVED: [4x] CoreBGPDown: Core BGP session down between asw1-b3-magru and cr1-magru (195.200.68.142) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [06:18:17] !log marostegui@cumin1003 START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie [06:18:55] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr1-magru:et-0/0/1 (Core: asw1-b3-magru:et-0/0/48) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [06:20:11] (03CR) 10Slyngshede: [C:03+2] Geo-maps: Update Meta mapping for september 2026 [dns] - 10https://gerrit.wikimedia.org/r/1333134 (owner: 10Slyngshede) [06:20:21] !log slyngshede@dns1004 START - running authdns-update [06:21:07] !log Drop cu* tables from s3 bswiktionary T435965 [06:21:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:21:10] T435965: Drop CheckUser tables from bswiktionary - https://phabricator.wikimedia.org/T435965 [06:22:43] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12278778 (10Krinkle) >>! In T435283#12248179, @BBlack wrote: > […] IMHO, the right fix for this is to fix... [06:23:04] !log slyngshede@dns1004 END - running authdns-update [06:25:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:31:59] !log marostegui@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage [06:38:36] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12278798 (10Bawolff) The main thing i think is weird is the inconsistency of it. I think either choice is... [06:39:20] (03PS1) 10Slyngshede: data.yaml: Offboarding kindrobot [puppet] - 10https://gerrit.wikimedia.org/r/1333558 [06:39:24] !log marostegui@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage [06:39:53] (03PS1) 10Mszwarc: UserInfoCard: Send the source page with the api_request event [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333559 (https://phabricator.wikimedia.org/T435585) [06:40:14] (03PS1) 10Mszwarc: UserInfoCard: Send the source page with the api_request event [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333560 (https://phabricator.wikimedia.org/T435585) [06:40:41] (03PS1) 10Mszwarc: UserInfoCard: Send the place of the trigger with api_request [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333561 (https://phabricator.wikimedia.org/T435585) [06:41:05] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1333558 (owner: 10Slyngshede) [06:41:08] (03PS1) 10Mszwarc: UserInfoCard: Send the place of the trigger with api_request [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333562 (https://phabricator.wikimedia.org/T435585) [06:41:32] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333559 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [06:41:40] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333560 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [06:41:48] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333561 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [06:41:54] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333562 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [06:42:02] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332684 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [06:43:20] !log hashar@deploy1003 Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies [06:43:33] (03CR) 10Slyngshede: [C:03+2] data.yaml: Offboarding kindrobot [puppet] - 10https://gerrit.wikimedia.org/r/1333558 (owner: 10Slyngshede) [06:43:33] !log hashar@deploy1003 Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s) [06:44:26] (03CR) 10EarlyWarningBot: "[Failed command](https://integration.wikimedia.org/ci/job/quibble-vendor-mysql-php83/105970/consoleFull): `./node_modules/.bin/grunt qunit" [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333561 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [06:44:27] (03CR) 10EarlyWarningBot: "[Failed command](https://integration.wikimedia.org/ci/job/quibble-with-gated-extensions-vendor-mysql-php83/54489/consoleFull): `./node_mod" [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333561 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [06:44:36] (03CR) 10CI reject: [V:04-1] UserInfoCard: Send the place of the trigger with api_request [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333561 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [06:45:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.28% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:46:38] (03CR) 10Muehlenhoff: [C:03+2] ircstream: Mark the IRC port as intentionally open to the world [puppet] - 10https://gerrit.wikimedia.org/r/1286863 (https://phabricator.wikimedia.org/T149804) (owner: 10Muehlenhoff) [06:48:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:48:36] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Review of firewall services without srange - https://phabricator.wikimedia.org/T149804#12278817 (10MoritzMuehlenhoff) [06:49:23] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Review of firewall services without srange - https://phabricator.wikimedia.org/T149804#12278820 (10MoritzMuehlenhoff) [06:51:10] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Upgrade Cumin hosts to Trixie - https://phabricator.wikimedia.org/T427897#12278832 (10MoritzMuehlenhoff) [06:53:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:53:30] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333390 (https://phabricator.wikimedia.org/T436734) (owner: 10EggRoll97) [06:53:43] (03CR) 10Muehlenhoff: [C:03+2] Allow cumin1004 in alertmanager and IRC notifications [puppet] - 10https://gerrit.wikimedia.org/r/1333239 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [06:54:45] (03Abandoned) 10Muehlenhoff: profile::tcpircbot: Make process check check for python3 instead of python2 [puppet] - 10https://gerrit.wikimedia.org/r/673077 (owner: 10CRusnov) [06:58:21] (03CR) 10Muehlenhoff: [C:03+2] Allow cumin1004 for RAPI access [puppet] - 10https://gerrit.wikimedia.org/r/1333242 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [06:58:45] (03PS2) 10Mszwarc: UserInfoCard: Send the place of the trigger with api_request [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333561 (https://phabricator.wikimedia.org/T435585) [07:00:05] Amir1, urbanecm, and awight: #bothumor My software never has bugs. It just develops random features. Rise for UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T0700). [07:00:05] Hide_on_rosie and Msz2001: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:12] o/ [07:00:40] I'm ready. WikimediaDebug installed as usual. [07:02:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.42% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:02:24] Okay, I'll start deploying in a moment [07:04:13] 06SRE, 06Infrastructure-Foundations, 10netops: Missing series for BGP session_state from eqiad CRs since upgrade to 23.4R2-S8.7 - https://phabricator.wikimedia.org/T435909#12278845 (10cmooney) >>! In T435909#12275570, @ayounsi wrote: > Adding a 3rd netflow host in eqiad shuffled the targets around, it caused... [07:04:25] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333338 (https://phabricator.wikimedia.org/T436672) (owner: 10Tryvix1509) [07:04:25] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333343 (https://phabricator.wikimedia.org/T436712) (owner: 10Tryvix1509) [07:04:26] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333390 (https://phabricator.wikimedia.org/T436734) (owner: 10EggRoll97) [07:05:37] (03Merged) 10jenkins-bot: throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333338 (https://phabricator.wikimedia.org/T436672) (owner: 10Tryvix1509) [07:05:40] (03Merged) 10jenkins-bot: InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333343 (https://phabricator.wikimedia.org/T436712) (owner: 10Tryvix1509) [07:05:44] (03Merged) 10jenkins-bot: Add two wmf groups to privileged status [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333390 (https://phabricator.wikimedia.org/T436734) (owner: 10EggRoll97) [07:06:07] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1333338|throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343|InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390|Add two wmf groups to privileged status (T436734)]] [07:06:14] T436672: Lift IP cap for Mapudungun editathon on 2026-09-05 - https://phabricator.wikimedia.org/T436672 [07:06:15] T436712: Update $wgUploadNavigationUrl on ukwiki - https://phabricator.wikimedia.org/T436712 [07:06:15] T436734: Add wmf-legal and wmf-antiabuse-restricted to privileged global groups - https://phabricator.wikimedia.org/T436734 [07:07:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:08:01] (03PS2) 10Giuseppe Lavagetto: validating-admission-policies: add gVisor enforcement policy [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333210 (https://phabricator.wikimedia.org/T436655) [07:08:02] (03PS4) 10Giuseppe Lavagetto: shellbox: add gVisor support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333139 (https://phabricator.wikimedia.org/T436649) [07:08:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:08:24] Are 2 [07:08:44] (03PS1) 10Muehlenhoff: Temporarily depool puppetserver[12]002 for reboots [dns] - 10https://gerrit.wikimedia.org/r/1333573 [07:09:12] Msz2001: Are 3 patches (1333338 1333343 1333390) on production now? [07:09:32] Not yet, it's still in progress syncing to test servers [07:10:35] !log mszwarc@deploy1003 mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338|throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343|InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390|Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri [07:10:35] fied there. [07:11:01] Hide_on_rosie: Now you can verify the patches [07:12:09] !log marostegui@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie [07:12:44] How do I verify the "lift IP cap" [07:13:06] Oh, we don't usually verify those, I think it's impossible in practice [07:13:21] (03CR) 10Tiziano Fogli: team-sre/hardware: add BBU check (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1308071 (https://phabricator.wikimedia.org/T430149) (owner: 10Hnowlan) [07:14:50] (03CR) 10Tiziano Fogli: [C:03+1] kafka: migrate check_kafka_ssl to alertmanager [puppet] - 10https://gerrit.wikimedia.org/r/1307405 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [07:15:34] https://usercontent.irccloud-cdn.com/file/VVDJ6UZU/image.png [07:15:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:15:49] T436734 ok. [07:15:50] T436734: Add wmf-legal and wmf-antiabuse-restricted to privileged global groups - https://phabricator.wikimedia.org/T436734 [07:16:30] And the ukwiki one? [07:16:49] ukwiki is ok [07:17:03] Okay, proceeding [07:17:09] !log mszwarc@deploy1003 mszwarc, eggroll97, tryvix1509: Continuing with deployment [07:18:46] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in profile toolforge [puppet] - 10https://gerrit.wikimedia.org/r/1332500 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:18:52] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228 [07:19:19] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module acme_chief [puppet] - 10https://gerrit.wikimedia.org/r/1332504 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:19:55] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332505 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:20:22] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332506 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:20:35] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1242: Cloning db1228 [07:20:40] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332507 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:20:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:20:57] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228 [07:21:03] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332508 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:21:25] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332509 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:21:38] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module service [puppet] - 10https://gerrit.wikimedia.org/r/1332511 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:22:01] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332514 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:22:12] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1333338|throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343|InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390|Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s) [07:22:18] T436672: Lift IP cap for Mapudungun editathon on 2026-09-05 - https://phabricator.wikimedia.org/T436672 [07:22:19] T436712: Update $wgUploadNavigationUrl on ukwiki - https://phabricator.wikimedia.org/T436712 [07:22:19] T436734: Add wmf-legal and wmf-antiabuse-restricted to privileged global groups - https://phabricator.wikimedia.org/T436734 [07:22:42] Deployed [07:22:50] Many thanks! [07:22:54] Now, I'll proceed to my other patches [07:22:56] yw :) [07:23:28] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332517 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:23:48] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333559 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [07:23:48] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333560 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [07:23:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333561 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [07:23:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333562 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [07:24:28] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module swift [puppet] - 10https://gerrit.wikimedia.org/r/1332519 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:24:44] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module graphite [puppet] - 10https://gerrit.wikimedia.org/r/1332521 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:24:59] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module ceph [puppet] - 10https://gerrit.wikimedia.org/r/1332518 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:25:13] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module icinga [puppet] - 10https://gerrit.wikimedia.org/r/1332522 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:25:40] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1332513 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:26:42] (03Merged) 10jenkins-bot: UserInfoCard: Send the source page with the api_request event [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333559 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [07:26:45] (03Merged) 10jenkins-bot: UserInfoCard: Send the source page with the api_request event [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333560 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [07:27:20] (03Merged) 10jenkins-bot: UserInfoCard: Send the place of the trigger with api_request [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333562 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [07:27:22] (03Merged) 10jenkins-bot: UserInfoCard: Send the place of the trigger with api_request [extensions/CheckUser] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333561 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [07:27:29] (03CR) 10Muehlenhoff: [C:03+2] Temporarily depool puppetserver[12]002 for reboots [dns] - 10https://gerrit.wikimedia.org/r/1333573 (owner: 10Muehlenhoff) [07:27:34] !log jmm@dns1004 START - running authdns-update [07:27:52] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1333559|UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560|UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561|UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562|UserInfoCard: Send the place of the trigger with api_request (T435585)] [07:27:52] ] [07:27:55] T435585: Add page_id to the mediawiki_product_metrics_user_info_card_interaction - https://phabricator.wikimedia.org/T435585 [07:30:00] !log jmm@dns1004 END - running authdns-update [07:30:39] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [07:30:50] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet [07:36:41] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet [07:36:42] (03CR) 10Jelto: [C:03+1] "lgtm, thank you!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333211 (https://phabricator.wikimedia.org/T436380) (owner: 10JMeybohm) [07:36:46] (03PS1) 10Marostegui: mariadb: Change s1 candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333599 (https://phabricator.wikimedia.org/T436078) [07:37:15] (03CR) 10Marostegui: "Changed in dbctl too" [puppet] - 10https://gerrit.wikimedia.org/r/1333599 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [07:37:31] (03CR) 10Marostegui: [C:03+2] mariadb: Change s1 candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333599 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [07:37:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:38:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:38:18] (03PS2) 10Elukey: Puppet 8: Replace unscoped legacy facts in module dnsrecursor [puppet] - 10https://gerrit.wikimedia.org/r/1332505 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:38:29] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332505 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:39:28] FIRING: KeyholderUnarmed: 2 unarmed Keyholder key(s) on cumin1004:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [07:39:38] (03PS1) 10Marostegui: mariadb: Change s6 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333600 (https://phabricator.wikimedia.org/T436078) [07:39:50] (03PS3) 10Slyngshede: mw-web: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332725 (https://phabricator.wikimedia.org/T433363) [07:40:29] (03CR) 10Marostegui: "Changed in dbctl too." [puppet] - 10https://gerrit.wikimedia.org/r/1333600 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [07:40:36] (03CR) 10Marostegui: [C:03+2] mariadb: Change s6 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333600 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [07:41:29] (03PS1) 10Slyngshede: mw-api-ext: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333603 (https://phabricator.wikimedia.org/T433363) [07:43:25] (03CR) 10Slyngshede: mw-web: upsize for single-DC serving (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332725 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [07:44:28] RESOLVED: KeyholderUnarmed: 2 unarmed Keyholder key(s) on cumin1004:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [07:45:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [07:47:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [07:48:31] (03CR) 10Tiziano Fogli: Add alert on percentage of timeouts from POPs in blackbox pings (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [07:49:05] (03PS1) 10Marostegui: mariadb: Change x1 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333607 (https://phabricator.wikimedia.org/T436078) [07:49:09] !log mszwarc@deploy1003 mszwarc: Backport for [[gerrit:1333559|UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560|UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561|UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562|UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the [07:49:09] testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:49:12] T435585: Add page_id to the mediawiki_product_metrics_user_info_card_interaction - https://phabricator.wikimedia.org/T435585 [07:49:40] (03CR) 10Marostegui: "Changed in zarcillo" [puppet] - 10https://gerrit.wikimedia.org/r/1333607 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [07:49:46] !log mszwarc@deploy1003 mszwarc: Continuing with deployment [07:50:37] (03CR) 10Marostegui: [C:03+2] mariadb: Change x1 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333607 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [07:56:04] 10ops-magru: magru: faulty link between cr1-magru and asw1-b3-magru - https://phabricator.wikimedia.org/T436675#12278980 (10ayounsi) 05Open→03Resolved So far things has been stable! [07:56:39] PROBLEM - SSH on cp5022 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [07:56:39] PROBLEM - Ensure traffic_exporter for the backend instance binds on port 9122 and responds to HTTP requests on cp5022 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [07:56:39] PROBLEM - Ensure traffic_manager binds on 3128 and responds to HTTP requests on cp5022 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [07:57:14] hashar: I see that this week the UTC evening slot is used for train, right? Can I overrun the current backport window with one more config patch? [07:59:33] (03CR) 10Elukey: [C:03+1] "I needed to explicitly set Hosts because pcc wasn't able to find a target :)" [puppet] - 10https://gerrit.wikimedia.org/r/1332505 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:00:05] dancy and hashar: Deploy window MediaWiki train - Utc-7+Utc-0 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T0800) [08:01:04] (03PS2) 10Elukey: Puppet 8: Replace unscoped legacy facts in module varnish [puppet] - 10https://gerrit.wikimedia.org/r/1332506 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:01:10] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332506 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:02:22] (03PS2) 10Elukey: Puppet 8: Replace unscoped legacy facts in profile cache [puppet] - 10https://gerrit.wikimedia.org/r/1332507 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:02:28] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332507 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:03:04] (03PS2) 10Elukey: Puppet 8: Replace unscoped legacy facts in profile dns [puppet] - 10https://gerrit.wikimedia.org/r/1332508 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:03:10] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332508 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:03:23] RESOLVED: [2x] GnmiInterfaceCountersDrop: cr2-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [08:03:42] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1333559|UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560|UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561|UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562|UserInfoCard: Send the place of the trigger with api_request (T435585) [08:03:42] ]] (duration: 35m 49s) [08:03:43] !log jmm@cumin2003 START - Cookbook sre.puppet.disable-merges [08:03:44] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) [08:03:45] T435585: Add page_id to the mediawiki_product_metrics_user_info_card_interaction - https://phabricator.wikimedia.org/T435585 [08:04:13] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet [08:05:02] (03PS2) 10Elukey: Puppet 8: Replace unscoped legacy facts in module envoyproxy [puppet] - 10https://gerrit.wikimedia.org/r/1332509 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:05:06] I see both conductors are away, so I'll deploy the remaining config patch [08:05:08] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332509 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:05:33] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332684 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [08:06:11] (03PS1) 10Brouberol: turnilo: sort the datacubes dimensions by human readable title [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333670 [08:06:14] (03PS2) 10Elukey: Puppet 8: Replace unscoped legacy facts in module opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1332514 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:06:23] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332514 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:07:11] !log depooling and silencing cp5022 (T414411) [08:07:14] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:07:15] T414411: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411 [08:07:21] (03Merged) 10jenkins-bot: Update stream config for user_info_card_interaction [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332684 (https://phabricator.wikimedia.org/T435585) (owner: 10Mszwarc) [08:07:23] (03PS2) 10Elukey: Puppet 8: Replace unscoped legacy facts in module cassandra [puppet] - 10https://gerrit.wikimedia.org/r/1332517 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:07:32] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332517 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:07:47] !log slyngshede@cumin1003 conftool action : set/pooled=no; selector: name=cp5022.* [08:07:51] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1332684|Update stream config for user_info_card_interaction (T435585)]] [08:08:23] (03PS2) 10Elukey: Puppet 8: Replace unscoped legacy facts in module swift [puppet] - 10https://gerrit.wikimedia.org/r/1332519 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:08:24] !log fabfur@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating [08:08:31] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332519 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:08:44] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12279048 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=3a675be0-bc5c-44f8-8996-c3f40f00c25d) set by fabfur@cumin1003 for 4:00:00 on 1 host(s) and their... [08:08:51] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module varnish [puppet] - 10https://gerrit.wikimedia.org/r/1332506 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:09:18] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in profile cache [puppet] - 10https://gerrit.wikimedia.org/r/1332507 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:09:36] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in profile dns [puppet] - 10https://gerrit.wikimedia.org/r/1332508 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:09:54] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module envoyproxy [puppet] - 10https://gerrit.wikimedia.org/r/1332509 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:10:01] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet [08:10:16] (03PS2) 10Brouberol: turnilo: sort the datacubes dimensions by human readable title [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333670 [08:10:28] (03PS3) 10Elukey: Puppet 8: Replace unscoped legacy facts in module opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1332514 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:10:34] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332514 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:10:50] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module cassandra [puppet] - 10https://gerrit.wikimedia.org/r/1332517 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:10:59] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1332514 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:11:47] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet [08:11:52] (03CR) 10Jelto: "one question in-line, maybe you can try a helm template yourself and double check the action names" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333256 (https://phabricator.wikimedia.org/T436380) (owner: 10JMeybohm) [08:12:25] FIRING: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:12:26] fabfur: o/ cp5022 was reimaged yesterday, and we fixed its ipv6 AAAA record. What happened? IRRC it was already depooled [08:12:40] https://phabricator.wikimedia.org/T436386 [08:12:56] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations: attempts to reimage cp5022 are failing - https://phabricator.wikimedia.org/T436386#12279060 (10elukey) 05Open→03Resolved a:03elukey [08:14:11] !log mszwarc@deploy1003 mszwarc: Backport for [[gerrit:1332684|Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [08:14:19] T435585: Add page_id to the mediawiki_product_metrics_user_info_card_interaction - https://phabricator.wikimedia.org/T435585 [08:14:38] !log mszwarc@deploy1003 mszwarc: Continuing with deployment [08:18:23] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet [08:19:09] !log jmm@cumin2003 START - Cookbook sre.puppet.disable-merges [08:19:11] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0) [08:20:23] (03CR) 10Joal: [C:03+1] turnilo: sort the datacubes dimensions by human readable title [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333670 (owner: 10Brouberol) [08:20:31] (03CR) 10Tiziano Fogli: mailman: migrate mailman_hours_until_empty_outbound_queue nrpe check (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1311820 (https://phabricator.wikimedia.org/T370157) (owner: 10Hnowlan) [08:20:53] (03CR) 10Jelto: [C:03+2] scaffold: Fix helm release NOTES.txt text (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332556 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [08:22:28] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1332684|Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s) [08:22:31] T435585: Add page_id to the mediawiki_product_metrics_user_info_card_interaction - https://phabricator.wikimedia.org/T435585 [08:22:44] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:23:20] (03Merged) 10jenkins-bot: scaffold: Fix helm release NOTES.txt text [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332556 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [08:23:34] !log UTC morning backport window done [08:23:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:24:59] (03CR) 10Brouberol: [C:03+2] turnilo: sort the datacubes dimensions by human readable title [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333670 (owner: 10Brouberol) [08:26:02] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply [08:26:11] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply [08:27:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:28:31] RECOVERY - SSH on cp5022 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [08:29:31] RECOVERY - Ensure traffic_exporter for the backend instance binds on port 9122 and responds to HTTP requests on cp5022 is OK: HTTP OK: HTTP/1.0 200 OK - 36134 bytes in 0.728 second response time https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [08:29:31] RECOVERY - Ensure traffic_manager binds on 3128 and responds to HTTP requests on cp5022 is OK: HTTP OK: HTTP/1.1 200 OK - 48370 bytes in 0.895 second response time https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server [08:30:33] (03PS4) 10Tiziano Fogli: P:monitoring: Remove absented resources [puppet] - 10https://gerrit.wikimedia.org/r/1307745 (owner: 10Majavah) [08:31:48] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12279144 (10MoritzMuehlenhoff) [08:32:59] (03PS5) 10Tiziano Fogli: P:monitoring: Remove absented resources [puppet] - 10https://gerrit.wikimedia.org/r/1307745 (owner: 10Majavah) [08:36:05] !log installing libgraphite2 security updates [08:36:06] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:37:17] (03CR) 10Tiziano Fogli: [C:03+1] "I've just rebased and resolved some conflicts. Please double-check https://gerrit.wikimedia.org/r/c/operations/puppet/+/1307745/3..5." [puppet] - 10https://gerrit.wikimedia.org/r/1307745 (owner: 10Majavah) [08:38:17] (03PS1) 10Muehlenhoff: Add library hint for graphite2 [puppet] - 10https://gerrit.wikimedia.org/r/1333683 [08:39:06] !log ayounsi@cumin1003 START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie [08:41:11] (03CR) 10Muehlenhoff: [C:03+2] Add library hint for graphite2 [puppet] - 10https://gerrit.wikimedia.org/r/1333683 (owner: 10Muehlenhoff) [08:41:31] (03CR) 10Tiziano Fogli: [C:03+1] hadoop: add hdfs alert for HA status [alerts] - 10https://gerrit.wikimedia.org/r/1304769 (https://phabricator.wikimedia.org/T407138) (owner: 10Hnowlan) [08:43:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:46:22] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply [08:46:24] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply [08:46:30] !log bump space for prometheus k8s-dse in eqiad [08:46:31] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:48:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [08:51:42] FIRING: [3x] JobUnavailable: Reduced availability for job fastnetmon in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [08:56:53] FIRING: FNMNotReported: FastNetMon metrics not reported - https://wikitech.wikimedia.org/wiki/Fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DFNMNotReported [09:00:27] (03PS1) 10Jelto: service::catalog: Set ipip_encapsulation for citoid, cxserver in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1333691 (https://phabricator.wikimedia.org/T420436) [09:01:42] RESOLVED: [3x] JobUnavailable: Reduced availability for job fastnetmon in ops@magru - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [09:02:27] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip_encapsulation for citoid, cxserver in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1332741 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [09:02:31] (03CR) 10Hnowlan: [C:03+2] ncredir: migrate nrpe check to alertmanager [puppet] - 10https://gerrit.wikimedia.org/r/1307419 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [09:04:42] (03CR) 10Majavah: [C:03+2] P:monitoring: Remove absented resources [puppet] - 10https://gerrit.wikimedia.org/r/1307745 (owner: 10Majavah) [09:06:52] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [09:07:02] (03PS1) 10Elukey: sre.hosts.provision: make the FQDN check for SM dynamic [cookbooks] - 10https://gerrit.wikimedia.org/r/1333692 (https://phabricator.wikimedia.org/T419892) [09:08:56] !log elukey@cumin1003 START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART [09:09:35] !log ayounsi@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage [09:10:09] !log installing openjdk-8 security updates [09:10:10] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:11:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [09:12:25] RESOLVED: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:13:22] (03PS2) 10Hnowlan: ncredir: remove absented icinga check [puppet] - 10https://gerrit.wikimedia.org/r/1318724 (https://phabricator.wikimedia.org/T407117) [09:14:32] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage [09:14:49] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART [09:16:33] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply [09:16:58] (03PS1) 10Marostegui: mariadb: Change s7 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333698 (https://phabricator.wikimedia.org/T436078) [09:17:01] (03PS1) 10Jelto: fix helm command in all NOTES.txt [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333699 (https://phabricator.wikimedia.org/T433589) [09:17:25] (03PS2) 10Elukey: sre.hosts.provision: make the FQDN check for SM dynamic [cookbooks] - 10https://gerrit.wikimedia.org/r/1333692 (https://phabricator.wikimedia.org/T419892) [09:17:31] (03CR) 10Marostegui: "Changed in dbctl too" [puppet] - 10https://gerrit.wikimedia.org/r/1333698 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [09:17:39] (03CR) 10Marostegui: [C:03+2] mariadb: Change s7 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333698 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [09:17:42] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply [09:17:53] !log elukey@cumin1003 START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART [09:19:25] (03PS1) 10Majavah: P:wmcs::novaproxy: Clean up absented resources [puppet] - 10https://gerrit.wikimedia.org/r/1333701 [09:21:45] (03CR) 10Majavah: [C:03+2] Puppet 8: Replace unscoped legacy facts in profile toolforge [puppet] - 10https://gerrit.wikimedia.org/r/1332500 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [09:22:34] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:23:13] (03PS2) 10Jelto: fix helm command in all NOTES.txt [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333699 (https://phabricator.wikimedia.org/T433589) [09:23:39] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply [09:23:46] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply [09:24:21] (03PS2) 10Majavah: P:elasticsearch: Remove unused profiles [puppet] - 10https://gerrit.wikimedia.org/r/1332772 [09:24:21] (03PS2) 10Majavah: elasticsearch: Remove unused classes [puppet] - 10https://gerrit.wikimedia.org/r/1332773 [09:26:40] (03PS1) 10Marostegui: mariadb: Change s2 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333704 (https://phabricator.wikimedia.org/T436078) [09:27:15] (03CR) 10Marostegui: "Changed in dbctl too." [puppet] - 10https://gerrit.wikimedia.org/r/1333704 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [09:27:37] (03CR) 10Marostegui: [C:03+2] mariadb: Change s2 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333704 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [09:28:08] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART [09:29:22] (03CR) 10Hnowlan: [C:03+2] ncredir: remove absented icinga check [puppet] - 10https://gerrit.wikimedia.org/r/1318724 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [09:30:09] (03PS1) 10Majavah: P:wmcs::proxy: Remove static maps proxy profile [puppet] - 10https://gerrit.wikimedia.org/r/1333707 (https://phabricator.wikimedia.org/T431284) [09:30:43] !log installing openjdk-21 security updates [09:30:44] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:31:33] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:31:53] RESOLVED: FNMNotReported: FastNetMon metrics not reported - https://wikitech.wikimedia.org/wiki/Fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DFNMNotReported [09:33:49] (03CR) 10Hnowlan: [C:03+2] kafka: migrate check_kafka_ssl to alertmanager [puppet] - 10https://gerrit.wikimedia.org/r/1307405 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [09:34:13] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie [09:36:11] (03PS1) 10Hnowlan: kafka: remove absented check [puppet] - 10https://gerrit.wikimedia.org/r/1333710 (https://phabricator.wikimedia.org/T407117) [09:36:38] (03PS1) 10Brouberol: turnilo: always display the Time dimension first [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333711 [09:36:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [09:40:00] (03CR) 10Brouberol: [C:03+2] turnilo: always display the Time dimension first [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333711 (owner: 10Brouberol) [09:42:21] !log ayounsi@cumin1003 START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie [09:42:29] (03Merged) 10jenkins-bot: turnilo: always display the Time dimension first [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333711 (owner: 10Brouberol) [09:43:26] (03PS1) 10Marostegui: instances.yaml: Add db1228 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1333718 (https://phabricator.wikimedia.org/T435892) [09:44:02] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:44:49] (03CR) 10Hnowlan: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332767 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [09:44:51] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1228 to dbctl [puppet] - 10https://gerrit.wikimedia.org/r/1333718 (https://phabricator.wikimedia.org/T435892) (owner: 10Marostegui) [09:45:51] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply [09:45:58] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply [09:46:12] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply [09:46:21] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply [09:47:14] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1228 to dbctl T435892', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json [09:47:17] T435892: Move db1243 to m5 and db1228 to s4. - https://phabricator.wikimedia.org/T435892 [09:48:04] FIRING: [2x] ProbeDown: Service kafka-logging2004:9093 has failed probes (tcp_kafka_broker_tls_ip4) - https://wikitech.wikimedia.org/wiki/Kafka/Administration#Renew_TLS_certificate - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:48:16] (03PS1) 10Muehlenhoff: Revert "Temporarily depool puppetserver[12]002 for reboots" [dns] - 10https://gerrit.wikimedia.org/r/1333719 [09:48:24] ^ that might be me, investigating [09:51:05] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1242: Pool back db1242 [09:52:52] (03CR) 10Blake: [C:03+2] mediawiki: optionally use initContainers for sidecars. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325883 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [09:53:37] (03CR) 10Muehlenhoff: [C:03+2] Revert "Temporarily depool puppetserver[12]002 for reboots" [dns] - 10https://gerrit.wikimedia.org/r/1333719 (owner: 10Muehlenhoff) [09:53:42] !log jmm@dns1004 START - running authdns-update [09:54:39] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:55:02] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:55:05] (03PS1) 10Hnowlan: Revert "kafka: migrate check_kafka_ssl to alertmanager" [puppet] - 10https://gerrit.wikimedia.org/r/1333720 [09:55:09] yeah I broke something in that check, reverting. [09:56:05] (03PS2) 10JMeybohm: vap: Add support for overriding validationActions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333256 (https://phabricator.wikimedia.org/T436380) [09:56:05] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:56:07] !log jmm@dns1004 END - running authdns-update [09:56:27] (03Merged) 10jenkins-bot: mediawiki: optionally use initContainers for sidecars. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325883 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [09:57:07] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:57:22] (03CR) 10Hnowlan: [C:03+2] Revert "kafka: migrate check_kafka_ssl to alertmanager" [puppet] - 10https://gerrit.wikimedia.org/r/1333720 (owner: 10Hnowlan) [09:58:04] FIRING: [18x] ProbeDown: Service kafka-jumbo1013:9093 has failed probes (tcp_kafka_broker_tls_ip4) - https://wikitech.wikimedia.org/wiki/Kafka/Administration#Renew_TLS_certificate - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:58:33] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:58:51] !log marostegui@cumin1003 END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242 [09:59:13] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1242: Pool back db1242 [09:59:35] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [09:59:53] FIRING: FNMNotReported: FastNetMon metrics not reported - https://wikitech.wikimedia.org/wiki/Fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DFNMNotReported [10:00:02] !log marostegui@cumin1003 END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242 [10:00:04] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1000) [10:01:00] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1242: Pool back db1242 [10:03:38] !log blake@deploy1003 Started scap sync-world: non-build deployment for T417800 [10:03:41] T417800: Implement proper sidecar container support in mediawiki pods - https://phabricator.wikimedia.org/T417800 [10:04:51] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:05:23] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [10:05:42] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:06:07] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:06:29] (03CR) 10Muehlenhoff: "One issue with the tracking of the Kerberos principal. Otherwise looks good, when the manager approval is in, you can simply self-merge an" [puppet] - 10https://gerrit.wikimedia.org/r/1333280 (https://phabricator.wikimedia.org/T436703) (owner: 10Kamila Součková) [10:08:08] !log blake@deploy1003 Finished scap sync-world: non-build deployment for T417800 (duration: 05m 37s) [10:08:33] !log ayounsi@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage [10:08:38] jouncebot: Nowandnext [10:08:38] For the next 0 hour(s) and 51 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1000) [10:08:38] In 0 hour(s) and 51 minute(s): Services – Citoid / Zotero (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1100) [10:10:23] FIRING: [3x] GnmiInterfaceCountersDrop: asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [10:14:19] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage [10:18:04] FIRING: [18x] ProbeDown: Service kafka-jumbo1013:9093 has failed probes (tcp_kafka_broker_tls_ip4) - https://wikitech.wikimedia.org/wiki/Kafka/Administration#Renew_TLS_certificate - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:18:54] (03PS1) 10Marostegui: mariadb: Change s4 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333729 (https://phabricator.wikimedia.org/T436078) [10:19:47] (03CR) 10Marostegui: "Changed in dbctl." [puppet] - 10https://gerrit.wikimedia.org/r/1333729 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [10:20:24] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:20:38] (03CR) 10Marostegui: [C:03+2] mariadb: Change s4 codfw candidate master [puppet] - 10https://gerrit.wikimedia.org/r/1333729 (https://phabricator.wikimedia.org/T436078) (owner: 10Marostegui) [10:23:04] FIRING: [32x] ProbeDown: Service kafka-jumbo1010:9093 has failed probes (tcp_kafka_broker_tls_ip4) - https://wikitech.wikimedia.org/wiki/Kafka/Administration#Renew_TLS_certificate - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:24:39] (03PS1) 10Jforrester: wikifunctions: Set up the functionmaintainer right for the community [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333732 (https://phabricator.wikimedia.org/T435637) [10:25:24] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:26:16] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:26:38] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:29:37] (03PS1) 10Marostegui: instances.yaml: Remove db1172 [puppet] - 10https://gerrit.wikimedia.org/r/1333739 (https://phabricator.wikimedia.org/T436763) [10:29:53] RESOLVED: FNMNotReported: FastNetMon metrics not reported - https://wikitech.wikimedia.org/wiki/Fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DFNMNotReported [10:30:56] (03CR) 10Marostegui: [C:03+2] instances.yaml: Remove db1172 [puppet] - 10https://gerrit.wikimedia.org/r/1333739 (https://phabricator.wikimedia.org/T436763) (owner: 10Marostegui) [10:31:52] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove db1172 from dbctl T436763', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json [10:31:56] T436763: decommission db1172.eqiad.wmnet - https://phabricator.wikimedia.org/T436763 [10:32:15] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie [10:32:40] (03PS1) 10Marostegui: db1172: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1333742 (https://phabricator.wikimedia.org/T436763) [10:35:43] (03CR) 10Marostegui: [C:03+2] db1172: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1333742 (https://phabricator.wikimedia.org/T436763) (owner: 10Marostegui) [10:37:47] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333745 [10:40:27] FIRING: [6x] GnmiInterfaceCountersDrop: asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [10:42:28] (03PS1) 10Hnowlan: icinga: migrate cert check to prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1333749 (https://phabricator.wikimedia.org/T407117) [10:44:24] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:45:12] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:45:23] FIRING: [8x] GnmiInterfaceCountersDrop: asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [10:46:04] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242 [10:46:32] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:46:47] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:46:58] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:47:46] (03CR) 10Mvolz: [C:03+2] zotero: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333167 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [10:48:21] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:49:47] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:50:09] (03PS1) 10Marostegui: db1228: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1333757 (https://phabricator.wikimedia.org/T435892) [10:50:10] (03Merged) 10jenkins-bot: zotero: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333167 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [10:50:38] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:50:43] Mvolz: I'm around in case something goes wrong btw, but similar changes have already been deployed on other services and it should be fine [10:52:54] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [10:53:04] RESOLVED: [16x] ProbeDown: Service kafka-jumbo1010:9093 has failed probes (tcp_kafka_broker_tls_ip4) - https://wikitech.wikimedia.org/wiki/Kafka/Administration#Renew_TLS_certificate - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:53:45] (03CR) 10Marostegui: [C:03+2] db1228: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1333757 (https://phabricator.wikimedia.org/T435892) (owner: 10Marostegui) [10:54:05] (03PS2) 10Clément Goubert: toolhub: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333169 (https://phabricator.wikimedia.org/T429175) [10:54:12] (03PS2) 10Clément Goubert: wmf-navigator: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333171 (https://phabricator.wikimedia.org/T429175) [10:54:17] (03PS2) 10Clément Goubert: airflow: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333173 (https://phabricator.wikimedia.org/T429175) [10:54:24] (03PS2) 10Clément Goubert: cxserver: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333172 (https://phabricator.wikimedia.org/T429175) [10:54:34] (03PS2) 10Clément Goubert: wdqs_common: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333174 (https://phabricator.wikimedia.org/T429175) [10:54:39] (03PS2) 10Clément Goubert: push-notifications: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333170 (https://phabricator.wikimedia.org/T429175) [10:55:51] (03CR) 10Clément Goubert: [C:03+1] cxserver: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333172 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [10:55:58] (03CR) 10Clément Goubert: [C:03+1] "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333172 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [10:56:35] (03CR) 10Clément Goubert: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333170 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [10:57:48] (03PS1) 10Hnowlan: kafka: migrate tls check to prometheus, per node [puppet] - 10https://gerrit.wikimedia.org/r/1333763 (https://phabricator.wikimedia.org/T407117) [10:59:10] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/zotero: apply [10:59:29] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/zotero: apply [11:00:05] mvolz: How many deployers does it take to do Services – Citoid / Zotero deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1100). [11:00:12] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/zotero: apply [11:00:19] (03CR) 10Mvolz: [C:03+2] citoid: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333168 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [11:00:47] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/zotero: apply [11:02:42] (03Merged) 10jenkins-bot: citoid: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333168 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [11:03:00] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/zotero: apply [11:03:28] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/zotero: apply [11:05:01] !log slyngshede@cumin1003 START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet [11:05:02] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet [11:05:23] FIRING: [10x] GnmiInterfaceCountersDrop: asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [11:05:43] !log slyngshede@cumin1003 conftool action : set/pooled=yes; selector: name=cp5022.* [11:08:08] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [11:08:27] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [11:09:13] (03PS1) 10Gkyziridis: ml-services: Deploy latest version of liftwing openapi specs. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333774 (https://phabricator.wikimedia.org/T434682) [11:09:43] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/citoid: apply [11:10:10] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/citoid: apply [11:11:34] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/citoid: apply [11:12:01] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/citoid: apply [11:14:17] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333778 [11:14:23] (03PS1) 10Muehlenhoff: Record LDAP access for dkertesz [puppet] - 10https://gerrit.wikimedia.org/r/1333779 [11:15:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:16:08] (03CR) 10Muehlenhoff: [C:03+2] Record LDAP access for dkertesz [puppet] - 10https://gerrit.wikimedia.org/r/1333779 (owner: 10Muehlenhoff) [11:17:04] (03CR) 10Mvolz: [C:03+2] citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333778 (owner: 10PipelineBot) [11:17:08] (03CR) 10Muehlenhoff: "Perfect, thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1333201 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [11:19:23] (03Merged) 10jenkins-bot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333778 (owner: 10PipelineBot) [11:20:22] !log marostegui@cumin1003 START - Cookbook sre.mysql.decommission [11:20:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:20:47] !log marostegui@cumin1003 START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet [11:21:27] (03PS1) 10Marostegui: mariadb: Decommission db1172 [puppet] - 10https://gerrit.wikimedia.org/r/1333784 (https://phabricator.wikimedia.org/T436763) [11:22:27] (03CR) 10Marostegui: [C:03+2] mariadb: Decommission db1172 [puppet] - 10https://gerrit.wikimedia.org/r/1333784 (https://phabricator.wikimedia.org/T436763) (owner: 10Marostegui) [11:23:02] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [11:23:19] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [11:24:46] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/citoid: apply [11:25:14] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/citoid: apply [11:25:47] !log marostegui@cumin1003 START - Cookbook sre.dns.netbox [11:26:09] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/citoid: apply [11:26:35] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/citoid: apply [11:29:26] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333745 (owner: 10PipelineBot) [11:29:34] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332461 (owner: 10PipelineBot) [11:29:41] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332460 (owner: 10PipelineBot) [11:29:52] (03Abandoned) 10Mvolz: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1324582 (owner: 10PipelineBot) [11:30:09] (03CR) 10AikoChou: [C:03+1] ml-services: Deploy latest version of liftwing openapi specs. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333774 (https://phabricator.wikimedia.org/T434682) (owner: 10Gkyziridis) [11:30:44] !log marostegui@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" [11:30:54] (03CR) 10Btullis: [C:03+1] dse-k8s-eqiad: provision the airflow instance [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [11:31:16] !log marostegui@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003" [11:31:16] !log marostegui@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:31:17] !log marostegui@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet [11:31:32] !log marostegui@cumin1003 Removing db1172 from zarcillo T436763 [11:31:36] T436763: decommission db1172.eqiad.wmnet - https://phabricator.wikimedia.org/T436763 [11:31:37] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.decommission (exit_code=0) [11:31:40] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: decommission db1172.eqiad.wmnet - https://phabricator.wikimedia.org/T436763#12279964 (10ops-monitoring-bot) db1172 has been decommissioned by Data Persistence [11:31:42] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: decommission db1172.eqiad.wmnet - https://phabricator.wikimedia.org/T436763#12279965 (10ops-monitoring-bot) a:05Marostegui→03None This host is ready for DC-Ops to decommission [11:32:01] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy latest version of liftwing openapi specs. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333774 (https://phabricator.wikimedia.org/T434682) (owner: 10Gkyziridis) [11:34:19] (03Merged) 10jenkins-bot: ml-services: Deploy latest version of liftwing openapi specs. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333774 (https://phabricator.wikimedia.org/T434682) (owner: 10Gkyziridis) [11:34:31] PROBLEM - orchestrator resolve cache non-FQDNs on dborch1002 is CRITICAL: CRITICAL: 1 non-FQDN entries in orchestrator resolve cache: https://wikitech.wikimedia.org/wiki/Orchestrator [11:35:31] RECOVERY - orchestrator resolve cache non-FQDNs on dborch1002 is OK: OK: all orchestrator resolve cache entries are FQDNs https://wikitech.wikimedia.org/wiki/Orchestrator [11:37:34] (03CR) 10TrainBranchBot: [C:03+2] "Approved by taavi@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333301 (https://phabricator.wikimedia.org/T436464) (owner: 10Southparkfan) [11:39:10] (03Merged) 10jenkins-bot: LabsServices: adjust Arc Lamp and Excimer URIs to use service RRs [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333301 (https://phabricator.wikimedia.org/T436464) (owner: 10Southparkfan) [11:39:59] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#de" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328201 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [11:40:37] (03CR) 10Btullis: [C:03+2] Declare the webrequest.dumps.v1 stream in EventStreamConfig [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328201 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [11:41:38] (03Merged) 10jenkins-bot: Declare the webrequest.dumps.v1 stream in EventStreamConfig [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328201 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [11:42:51] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12280103 (10Verdy_p) Copy of a message I posted in Wikimedia Commons's Village Pump: == Severe bug in th... [11:43:33] (03CR) 10Mpostoronca: [C:03+2] EventStreamConfig: Register the abuse_review_interaction stream [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329360 (https://phabricator.wikimedia.org/T435517) (owner: 10Dreamy Jazz) [11:44:34] (03Merged) 10jenkins-bot: EventStreamConfig: Register the abuse_review_interaction stream [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329360 (https://phabricator.wikimedia.org/T435517) (owner: 10Dreamy Jazz) [11:45:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.42% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:48:05] (03PS1) 10Muehlenhoff: Add component/thumbor for trixie-wikimedia [puppet] - 10https://gerrit.wikimedia.org/r/1333802 (https://phabricator.wikimedia.org/T436505) [11:51:19] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . [11:51:43] !log gkyziridis@deploy1003 helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' . [11:53:06] (03PS1) 10Muehlenhoff: thumbor-plugins: Enable component/thumbor in Blubber config [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333803 (https://phabricator.wikimedia.org/T436505) [11:53:56] (03CR) 10Ladsgroup: [C:03+1] Add component/thumbor for trixie-wikimedia [puppet] - 10https://gerrit.wikimedia.org/r/1333802 (https://phabricator.wikimedia.org/T436505) (owner: 10Muehlenhoff) [11:55:20] Hello. urbanecm , TheresNoTime - I have scheduled a mediawiki-config change for the upcoming UTC afternoon backport window. The patch has +2 in gerrit. I haven't scheduled many of these, so I hope I've done it right. I have the WikimediaDebug extension installed and I'm ready to test the https://meta.wikimedia.org/w/api.php?action=streamconfigs&streams=webrequest.dumps.v1 with it. [11:56:27] jouncebot: nowandnext [11:56:27] For the next 0 hour(s) and 3 minute(s): Services – Citoid / Zotero (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1100) [11:56:27] In 1 hour(s) and 3 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1300) [11:57:03] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#de" [extensions/WikiLambda] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333264 (https://phabricator.wikimedia.org/T436579) (owner: 10Jforrester) [11:57:10] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#de" [extensions/WikiLambda] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333263 (https://phabricator.wikimedia.org/T436579) (owner: 10Jforrester) [11:57:14] btullis: you look all set then :) I should be around for that window but if not another deployer will be [11:57:15] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1228 crashed again - https://phabricator.wikimedia.org/T436743#12280171 (10Jclark-ctr) @Marostegui Although its to late to open a dell ticket since support ended . If you would like me to update firmwares you can reopen this ticket and assign to me and I can ta... [11:58:19] btullis: ah I see it was +2'd, that should have waited for the window. Gimme 10 minutes and I'll take a closer look (on mobile atm) [12:00:18] 10ops-eqiad, 06SRE, 06DBA, 06DC-Ops: db1228 crashed again - https://phabricator.wikimedia.org/T436743#12280176 (10Marostegui) Hey John! I think this host got everything updated at https://phabricator.wikimedia.org/T430934#12082613 which is what Dell told us to do, and after that it crashed again a few... [12:02:17] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12280178 (10Verdy_p) Note that I also found another problem that this bug has caused, notably a few battl... [12:03:20] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: decommission db1172.eqiad.wmnet - https://phabricator.wikimedia.org/T436763#12280180 (10Jclark-ctr) a:03Jclark-ctr [12:03:36] (03CR) 10Filippo Giunchedi: [C:03+1] P:wmcs::proxy: Remove static maps proxy profile [puppet] - 10https://gerrit.wikimedia.org/r/1333707 (https://phabricator.wikimedia.org/T431284) (owner: 10Majavah) [12:08:38] (03CR) 10Atsuko: [C:03+2] Grant sudo privileges for the analytics-experiment-users group [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:09:27] Maxim merged one of my config patches which needs to wait for a GitLab MR to be merged [12:09:44] I've asked them to review that one soon, but will revert the config patch if it gets in the way [12:10:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:10:29] TheresNoTime: Oh, I'm sorry that I gave it the +2 before the window. I should have asked first. [12:11:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.69% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:11:25] Maxim has approved the GitLab MR, going to merge [12:11:57] btullis: all good :) [12:12:24] Dreamy_Jazz: you're going to do a deploy now? [12:12:35] Once I've got this GitLab MR merged [12:13:18] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: decommission db1172.eqiad.wmnet - https://phabricator.wikimedia.org/T436763#12280213 (10Jclark-ctr) [12:13:20] 10ops-eqiad, 06DBA, 06DC-Ops, 10decommission-hardware: decommission db1172.eqiad.wmnet - https://phabricator.wikimedia.org/T436763#12280214 (10Jclark-ctr) 05Open→03Resolved [12:13:20] Happy to do both commits or let you do both as you want [12:13:26] (03CR) 10Muehlenhoff: [C:03+2] Add component/thumbor for trixie-wikimedia [puppet] - 10https://gerrit.wikimedia.org/r/1333802 (https://phabricator.wikimedia.org/T436505) (owner: 10Muehlenhoff) [12:13:27] *commits to mediawiki-config [12:13:36] ack - btullis if you're free fairly soon to test that patch of yours, it'll be live once Dreamy_Jazz does their deploy [12:14:01] I think I'd need btullis's change reverted if we do separately? [12:14:21] B/c both are merged in the mediawiki-config repo, and scap pulls the master branch [12:15:05] oh sorry I meant do them both yeah [12:15:27] Oh yeah, I see [12:15:37] I'll deploy both and ping when at testservers [12:16:51] jouncebot: nowandnext [12:16:51] No deployments scheduled for the next 0 hour(s) and 43 minute(s) [12:16:51] In 0 hour(s) and 43 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1300) [12:17:20] Yes, I'm around. [12:17:40] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1328201|Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360|EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] [12:17:47] T425087: Send JSON access logs for dumps.wikimedia.org to Kafka - https://phabricator.wikimedia.org/T425087 [12:17:48] T291645: Produce ECS formatted logstash logs to Event Platform, allowing them to be queried in the WMF Data Lake with SQL - https://phabricator.wikimedia.org/T291645 [12:17:48] T435517: AbuseReview: Server side instrumentation - https://phabricator.wikimedia.org/T435517 [12:20:36] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw [12:21:00] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip_encapsulation for citoid, cxserver in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1332741 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [12:21:30] (03CR) 10Nikerabbit: "I don't see a new deployment scheduled." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329314 (https://phabricator.wikimedia.org/T433483) (owner: 10Abijeet Patro) [12:22:13] !log dreamyjazz@deploy1003 dreamyjazz, btullis: Backport for [[gerrit:1328201|Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360|EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [12:22:32] Canary testing of my patch looks good to me: https://phabricator.wikimedia.org/T425087#12280316 [12:24:00] Thanks [12:24:08] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [12:24:59] !log dreamyjazz@deploy1003 dreamyjazz, btullis: Continuing with deployment [12:25:03] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [12:25:03] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw [12:27:52] (03PS1) 10Gmodena: rest-gateway: front wdqs v2 endpoints in shadow mode [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333815 (https://phabricator.wikimedia.org/T436778) [12:28:37] (03PS2) 10Jelto: service::catalog: Set ipip_encapsulation for citoid, cxserver in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1333691 (https://phabricator.wikimedia.org/T420436) [12:30:09] (03CR) 10CI reject: [V:04-1] rest-gateway: front wdqs v2 endpoints in shadow mode [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333815 (https://phabricator.wikimedia.org/T436778) (owner: 10Gmodena) [12:30:30] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328201|Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360|EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s) [12:30:37] T425087: Send JSON access logs for dumps.wikimedia.org to Kafka - https://phabricator.wikimedia.org/T425087 [12:30:38] T291645: Produce ECS formatted logstash logs to Event Platform, allowing them to be queried in the WMF Data Lake with SQL - https://phabricator.wikimedia.org/T291645 [12:30:38] T435517: AbuseReview: Server side instrumentation - https://phabricator.wikimedia.org/T435517 [12:30:53] Deploy finished [12:31:27] (03PS2) 10Gmodena: rest-gateway: front wdqs v2 endpoints in shadow mode [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333815 (https://phabricator.wikimedia.org/T436778) [12:31:29] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12280390 (10MoritzMuehlenhoff) [12:31:55] (03PS1) 10Matthias Mullie: Enable mobile MMV on all wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333816 (https://phabricator.wikimedia.org/T429970) [12:32:05] (03PS1) 10Atsuko: provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1333817 (https://phabricator.wikimedia.org/T416709) [12:33:51] (03CR) 10CI reject: [V:04-1] rest-gateway: front wdqs v2 endpoints in shadow mode [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333815 (https://phabricator.wikimedia.org/T436778) (owner: 10Gmodena) [12:34:14] !log jgiannelos@deploy1003 helmfile [eqiad] START helmfile.d/services/push-notifications: apply [12:34:17] (03PS1) 10Atsuko: trafficserver: enabling access to airflow-experiment-platform.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1333820 (https://phabricator.wikimedia.org/T416709) [12:34:20] !log jgiannelos@deploy1003 helmfile [staging] START helmfile.d/services/push-notifications: apply [12:34:24] !log jgiannelos@deploy1003 helmfile [staging] DONE helmfile.d/services/push-notifications: apply [12:34:29] !log jgiannelos@deploy1003 helmfile [eqiad] START helmfile.d/services/push-notifications: apply [12:34:32] !log jgiannelos@deploy1003 helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply [12:34:36] !log jgiannelos@deploy1003 helmfile [codfw] START helmfile.d/services/push-notifications: apply [12:34:41] !log jgiannelos@deploy1003 helmfile [codfw] DONE helmfile.d/services/push-notifications: apply [12:35:15] !log jgiannelos@deploy1003 helmfile [staging] START helmfile.d/services/push-notifications: apply [12:35:21] !log jgiannelos@deploy1003 helmfile [staging] DONE helmfile.d/services/push-notifications: apply [12:36:23] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 02 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#de" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333816 (https://phabricator.wikimedia.org/T429970) (owner: 10Matthias Mullie) [12:38:34] (03CR) 10Jelto: "codfw was uneventful (I88567cf04af428a577cb8385f839d77d006c6e7c), let's continue in eqiad" [puppet] - 10https://gerrit.wikimedia.org/r/1333691 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [12:38:38] (03CR) 10Btullis: [C:03+1] provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1333817 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:38:54] (03PS3) 10Jelto: service::catalog: Set ipip_encapsulation for citoid, cxserver in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1333691 (https://phabricator.wikimedia.org/T420436) [12:39:33] (03CR) 10Btullis: [C:03+1] trafficserver: enabling access to airflow-experiment-platform.w.o [puppet] - 10https://gerrit.wikimedia.org/r/1333820 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:39:37] (03CR) 10Atsuko: [C:03+2] provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1333817 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:39:47] (03CR) 10Brouberol: [C:03+1] provision the airflow-experiment-platform DNS records [dns] - 10https://gerrit.wikimedia.org/r/1333817 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:41:13] (03CR) 10Brouberol: [C:04-1] "This should only be merged when the following command succeeds (when the app has been deployed):" [puppet] - 10https://gerrit.wikimedia.org/r/1333820 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:41:39] !log atsuko@dns1004 START - running authdns-update [12:42:14] (03CR) 10Brouberol: dse-k8s-eqiad: provision the airflow instance (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:44:09] !log atsuko@dns1004 END - running authdns-update [12:47:08] (03CR) 10Ammarpad: "Hello @slyngshede@wikimedia.org @dzahn@wikimedia.org" [puppet] - 10https://gerrit.wikimedia.org/r/1325272 (owner: 10Ammarpad) [12:48:29] (03CR) 10Ammarpad: "Please let me know if there's still something needed. Thanks" [puppet] - 10https://gerrit.wikimedia.org/r/1325272 (owner: 10Ammarpad) [12:49:38] (03CR) 10Atsuko: [C:03+2] dse-k8s-eqiad: provision the airflow instance [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:51:52] (03Merged) 10jenkins-bot: dse-k8s-eqiad: provision the airflow instance [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:55:18] (03CR) 10Clément Goubert: "We have been talking over that rename for a while but just got to [filing the task](https://phabricator.wikimedia.org/T436785) and actuall" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [12:56:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:56:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.94% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [13:00:05] urbanecm and TheresNoTime: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1300). [13:00:05] Btullis, James_F, and matthiasmullie: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:15] o/ [13:00:28] (03PS1) 10Muehlenhoff: Add cumin1004 as additional firmware sync host [puppet] - 10https://gerrit.wikimedia.org/r/1333834 (https://phabricator.wikimedia.org/T427897) [13:00:42] FIRING: [3x] JobUnavailable: Reduced availability for job fastnetmon in ops@drmrs - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:00:42] (03PS3) 10Gmodena: rest-gateway: front wdqs v2 endpoints in shadow mode [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333815 (https://phabricator.wikimedia.org/T436778) [13:01:50] o/ [13:02:05] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply [13:02:08] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply [13:04:12] !log import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia T436505 [13:04:15] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:04:15] T436505: 400 Bad Request on File:Warsaw Pact in 1990 (orthographic projection).svg - https://phabricator.wikimedia.org/T436505 [13:05:53] FIRING: FNMNotReported: FastNetMon metrics not reported - https://wikitech.wikimedia.org/wiki/Fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DFNMNotReported [13:06:52] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [13:13:46] Is anyone around to deploy? [13:13:58] (03CR) 10Ladsgroup: [C:03+1] "Thank you!" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333803 (https://phabricator.wikimedia.org/T436505) (owner: 10Muehlenhoff) [13:15:27] FIRING: [10x] GnmiInterfaceCountersDrop: asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:21:33] Still need a deployer James_F ? [13:21:39] Yes, sorry. [13:21:45] Also matthiasmullie does. [13:21:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [13:21:52] np, can I deploy both of yours together? [13:21:57] Sure. [13:22:04] Yeah sure [13:22:48] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [extensions/WikiLambda] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333264 (https://phabricator.wikimedia.org/T436579) (owner: 10Jforrester) [13:22:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [extensions/WikiLambda] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333263 (https://phabricator.wikimedia.org/T436579) (owner: 10Jforrester) [13:22:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333816 (https://phabricator.wikimedia.org/T429970) (owner: 10Matthias Mullie) [13:23:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [13:23:25] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply [13:23:33] (03CR) 10Muehlenhoff: [C:03+2] thumbor-plugins: Enable component/thumbor in Blubber config [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333803 (https://phabricator.wikimedia.org/T436505) (owner: 10Muehlenhoff) [13:23:43] !log jgiannelos@deploy1003 helmfile [staging] START helmfile.d/services/push-notifications: apply [13:23:49] (03Merged) 10jenkins-bot: Enable mobile MMV on all wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333816 (https://phabricator.wikimedia.org/T429970) (owner: 10Matthias Mullie) [13:23:58] !log jgiannelos@deploy1003 helmfile [staging] DONE helmfile.d/services/push-notifications: apply [13:24:02] !log jgiannelos@deploy1003 helmfile [eqiad] START helmfile.d/services/push-notifications: apply [13:24:21] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply [13:24:35] !log jgiannelos@deploy1003 helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply [13:24:35] !log installing wireshark security updates [13:24:38] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:24:40] !log jgiannelos@deploy1003 helmfile [codfw] START helmfile.d/services/push-notifications: apply [13:25:16] !log jgiannelos@deploy1003 helmfile [codfw] DONE helmfile.d/services/push-notifications: apply [13:26:24] (03Merged) 10jenkins-bot: Add thumb.* to allowed hosts for commons images [extensions/WikiLambda] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1333264 (https://phabricator.wikimedia.org/T436579) (owner: 10Jforrester) [13:26:27] (03Merged) 10jenkins-bot: Add thumb.* to allowed hosts for commons images [extensions/WikiLambda] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333263 (https://phabricator.wikimedia.org/T436579) (owner: 10Jforrester) [13:26:54] !log samtar@deploy1003 Started scap sync-world: Backport for [[gerrit:1333264|Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263|Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816|Enable mobile MMV on all wikis (T429970)]] [13:26:59] T436579: many image fetches from AW failing - https://phabricator.wikimedia.org/T436579 [13:27:00] T429970: Enable beta mobile MMV on all wikis - https://phabricator.wikimedia.org/T429970 [13:28:02] (03CR) 10Atsuko: [C:03+2] "Merged I76351d03dd844b9c9792c4689170e5fb43e1dc02 and Id196e6fc9336bb5b582e2bba48cb3f3335e391fd, applied k8s:" [puppet] - 10https://gerrit.wikimedia.org/r/1333820 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [13:29:46] (03PS3) 10Blake: sidecars: Specify restartPolicy: Always. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) [13:30:23] TheresNoTime: Checked sneakily on mw-debug mid-scale out, now fixed. :-) [13:30:42] FIRING: [3x] JobUnavailable: Reduced availability for job fastnetmon in ops@drmrs - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:30:53] RESOLVED: FNMNotReported: FastNetMon metrics not reported - https://wikitech.wikimedia.org/wiki/Fastnetmon - https://w.wiki/8oU - https://alerts.wikimedia.org/?q=alertname%3DFNMNotReported [13:31:03] woo! [13:31:22] !log samtar@deploy1003 jforrester, samtar, mlitn: Backport for [[gerrit:1333264|Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263|Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816|Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:31:38] Also looking good on my end [13:31:52] !log samtar@deploy1003 jforrester, samtar, mlitn: Continuing with deployment [13:35:40] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs1020 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [13:36:39] (03CR) 10Blake: "Sorry, no need to review this just yet, I think this can wait until we've changed all the other 'sidecars' to accept a restartPolicy as we" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [13:36:47] !log samtar@deploy1003 Finished scap sync-world: Backport for [[gerrit:1333264|Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263|Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816|Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s) [13:36:53] T436579: many image fetches from AW failing - https://phabricator.wikimedia.org/T436579 [13:36:53] T429970: Enable beta mobile MMV on all wikis - https://phabricator.wikimedia.org/T429970 [13:37:03] James_F, matthiasmullie: done :) [13:37:10] thanks, TheresNoTime! [13:37:49] Thank you. [13:40:42] RESOLVED: [3x] JobUnavailable: Reduced availability for job fastnetmon in ops@drmrs - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:41:17] (03PS9) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) [13:43:40] (03CR) 10Jelto: helmfile.d: deploy etherpad to aux-k8s clusters (035 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [13:44:04] !log bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 T427897 [13:44:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:44:08] T427897: Upgrade Cumin hosts to Trixie - https://phabricator.wikimedia.org/T427897 [13:45:18] (03CR) 10Jelto: "websockets is a good topic. I think for the initial testing we can use long-polling. I'll look into websockets with our current ingress se" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [13:46:34] (03PS1) 10JavierMonton: stream: pageview-trending-relative-next [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333854 (https://phabricator.wikimedia.org/T431555) [13:49:17] (03CR) 10MVernon: [C:04-1] "Does the CDN take any notice of the Cache-Control: header set by swift? Maybe we can just drop ensure_max_age entirely..." [puppet] - 10https://gerrit.wikimedia.org/r/1327104 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [13:50:40] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs1019 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [13:51:31] (03PS1) 10Blake: mcrouter: update to 1.3.6 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333855 (https://phabricator.wikimedia.org/T417800) [13:52:16] (03PS1) 10Blake: mcrouter: Upgrade to 1.3.6. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333856 (https://phabricator.wikimedia.org/T417800) [13:52:33] (03PS1) 10Muehlenhoff: Record LDAP access for cklimas [puppet] - 10https://gerrit.wikimedia.org/r/1333857 [13:54:18] (03CR) 10Cathal Mooney: [C:03+1] "Ulsfo is rock solid on HE. Drmrs not quite perfect but been pretty good, so given where these are I think it's safe." [homer/public] - 10https://gerrit.wikimedia.org/r/1332723 (owner: 10Ayounsi) [13:54:34] (03PS1) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333858 (https://phabricator.wikimedia.org/T426337) [13:54:57] (03PS1) 10Jforrester: wikifunctions: Enable callbacks on production, not just staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333859 (https://phabricator.wikimedia.org/T427863) [13:55:17] (03CR) 10Muehlenhoff: [C:03+2] Record LDAP access for cklimas [puppet] - 10https://gerrit.wikimedia.org/r/1333857 (owner: 10Muehlenhoff) [13:56:38] (03PS2) 10Jforrester: wikifunctions: Enable callbacks on production, not just staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333859 (https://phabricator.wikimedia.org/T421848) [13:56:46] (03CR) 10MVernon: "[removing myself from reviewers, as the ceph modules aren't ones I've worked on - the apus ones are *cephadm* ]" [puppet] - 10https://gerrit.wikimedia.org/r/1332518 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [13:57:39] (03PS2) 10Ladsgroup: thumbor: Supported animated webp thumbnailing and use forked librsvg [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333198 (https://phabricator.wikimedia.org/T290345) [13:58:07] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333198 (https://phabricator.wikimedia.org/T290345) (owner: 10Ladsgroup) [13:58:11] (03PS2) 10Ladsgroup: swift: Also accept thumb.wikimedia.org in ensure_max_mage [puppet] - 10https://gerrit.wikimedia.org/r/1327104 (https://phabricator.wikimedia.org/T427465) [13:58:29] (03CR) 10MVernon: "Hi," [puppet] - 10https://gerrit.wikimedia.org/r/1332519 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [13:58:44] (03CR) 10Ladsgroup: "Aah, thanks. I looked at the code and thought the split is something in config value. Not simply the python's built-in split. Makes sense." [puppet] - 10https://gerrit.wikimedia.org/r/1327104 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [13:59:01] (03CR) 10Ladsgroup: [C:03+2] thumbor: Supported animated webp thumbnailing and use forked librsvg [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333198 (https://phabricator.wikimedia.org/T290345) (owner: 10Ladsgroup) [14:00:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1400) [14:01:44] (03Merged) 10jenkins-bot: thumbor: Supported animated webp thumbnailing and use forked librsvg [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333198 (https://phabricator.wikimedia.org/T290345) (owner: 10Ladsgroup) [14:02:43] (03CR) 10Cory Massaro: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333858 (https://phabricator.wikimedia.org/T426337) (owner: 10Jforrester) [14:04:23] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [14:04:33] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [14:05:19] (03Merged) 10jenkins-bot: wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333858 (https://phabricator.wikimedia.org/T426337) (owner: 10Jforrester) [14:06:03] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [14:06:07] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:06:20] !log apine@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:06:35] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [14:06:41] !log apine@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:07:36] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [14:08:10] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [14:08:16] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [14:08:42] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [14:08:43] !log apine@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:09:03] !log apine@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:09:37] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [14:09:49] !log apine@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:10:23] FIRING: [7x] GnmiInterfaceCountersDrop: asw1-b13-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [14:12:13] (03PS1) 10Thiemo Kreuz (WMDE): Remove revisionId from term fallback cache lines for properties [extensions/Wikibase] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333866 (https://phabricator.wikimedia.org/T434204) [14:12:38] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [14:13:03] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, September 03 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [extensions/Wikibase] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333866 (https://phabricator.wikimedia.org/T434204) (owner: 10Thiemo Kreuz (WMDE)) [14:14:34] (03CR) 10Elukey: [C:03+1] "@mvernon@wikimedia.org Hi! My read is that processor count doesn't carry a scope (like $:: etc.. as prefix) so in theory there could be a " [puppet] - 10https://gerrit.wikimedia.org/r/1332519 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:15:11] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12281013 (10MoritzMuehlenhoff) [14:15:19] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [14:17:11] (03PS1) 10Atsuko: airflow-test-k8s: remove growthbook-sync-related access [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333867 (https://phabricator.wikimedia.org/T416709) [14:17:39] (03CR) 10Atsuko: [C:03+2] dse-k8s-eqiad: provision the airflow instance (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [14:17:52] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12281030 (10RobH) > Dear customer, > > We acknowledge receipt of your case and will proceed to complete your requested task within the timeframe you specified. We will upd... [14:19:10] (03CR) 10Ozge: [C:03+1] ml-services: Initial deployment of Qwen3.8-27B-FP8 model on experimental ns. (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333204 (https://phabricator.wikimedia.org/T436644) (owner: 10Gkyziridis) [14:19:47] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12281035 (10MoritzMuehlenhoff) [14:19:51] (03CR) 10Clément Goubert: editcheck-headless: Add deployment for the technical pilot (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [14:20:07] !log apine@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:22:18] (03PS1) 10Tryvix1509: core-Permissions.php Allow English Wikiquote administrators to grant and remove the confirmed user permission [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333872 (https://phabricator.wikimedia.org/T436651) [14:22:37] (03CR) 10Majavah: [C:03+2] P:wmcs::proxy: Remove static maps proxy profile [puppet] - 10https://gerrit.wikimedia.org/r/1333707 (https://phabricator.wikimedia.org/T431284) (owner: 10Majavah) [14:22:59] (03PS1) 10Cory Massaro: Revert "wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333873 [14:23:08] (03PS2) 10Tryvix1509: core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333872 (https://phabricator.wikimedia.org/T436651) [14:23:16] (03CR) 10Jforrester: [C:03+2] Revert "wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333873 (owner: 10Cory Massaro) [14:23:23] FIRING: [2x] GnmiTargetDown: lsw1-d2-eqiad is unreachable through gNMI - https://wikitech.wikimedia.org/wiki/Network_telemetry#Troubleshooting - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic - https://alerts.wikimedia.org/?q=alertname%3DGnmiTargetDown [14:23:28] (03CR) 10Scott French: [C:03+2] Rakefile: Update mock services_proxy data for splits [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [14:25:24] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [14:25:36] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, September 03 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333872 (https://phabricator.wikimedia.org/T436651) (owner: 10Tryvix1509) [14:25:55] (03Merged) 10jenkins-bot: Revert "wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333873 (owner: 10Cory Massaro) [14:26:07] !log apine@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:26:20] !log apine@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:27:05] !log apine@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:27:13] (03CR) 10Elukey: "For Dell we have been setting balanced performance in the bios to all the hosts, and it is not 100% clear what we should do to all product" [cookbooks] - 10https://gerrit.wikimedia.org/r/1333146 (https://phabricator.wikimedia.org/T435537) (owner: 10Elukey) [14:27:22] !log apine@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:30:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1400) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1430) [14:32:21] (03PS1) 10Dpogorzelski: ml: images for Lift Wing Studio [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1333876 [14:32:48] !log installing pdns-recursor security updates [14:32:49] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:34:07] (03CR) 10Muehlenhoff: [C:03+2] Add cumin1004 as additional firmware sync host [puppet] - 10https://gerrit.wikimedia.org/r/1333834 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [14:38:23] RESOLVED: [2x] GnmiTargetDown: lsw1-d2-eqiad is unreachable through gNMI - https://wikitech.wikimedia.org/wiki/Network_telemetry#Troubleshooting - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic - https://alerts.wikimedia.org/?q=alertname%3DGnmiTargetDown [14:41:41] (03CR) 10Slyngshede: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1333275 (https://phabricator.wikimedia.org/T436691) (owner: 10Ssingh) [14:42:39] (03CR) 10Elukey: "I finally remembered where I saw something related, and indeed it was in the Java configs for keystores. We do have sslcert::x509_to_pkcs8" [puppet] - 10https://gerrit.wikimedia.org/r/1331768 (https://phabricator.wikimedia.org/T436393) (owner: 10Cwhite) [14:44:01] (03Merged) 10jenkins-bot: Rakefile: Update mock services_proxy data for splits [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [14:44:40] (03CR) 10Elukey: [C:03+1] "Didn't test it but it seems easy enough!" [cookbooks] - 10https://gerrit.wikimedia.org/r/1329523 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [14:44:45] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru [14:45:15] (03CR) 10Ssingh: [C:03+2] sretest2013: set role to cache::text [puppet] - 10https://gerrit.wikimedia.org/r/1333275 (https://phabricator.wikimedia.org/T436691) (owner: 10Ssingh) [14:46:46] (03PS8) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [14:47:10] (03Abandoned) 10Scott French: mesh.configuration: Fix split listener stream idle timeout in 1.15.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325996 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [14:47:14] (03Abandoned) 10Scott French: mesh.configuration: Add support for split host_regex in 1.15.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305514 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [14:47:31] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12281200 (10ssingh) @Jhancock.wm: this is ready on the software side for the provisioning when the server gets delivered. We will take care of the reimage and... [14:49:26] (03CR) 10CI reject: [V:04-1] Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [14:50:23] FIRING: [7x] GnmiInterfaceCountersDrop: asw1-b13-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [14:51:21] (03PS9) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) [14:52:38] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12281254 (10elukey) 05In progress→03Resolved a:03elukey The host seems stable, I am inclined to close this and... [14:55:23] FIRING: [7x] GnmiInterfaceCountersDrop: asw1-b13-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [14:55:23] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921#12281303 (10VRiley-WMF) 05Open→03In progress Thanks, I am proceeding with this now [14:56:59] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru [14:57:38] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921#12281323 (10VRiley-WMF) adding this for record keepin on steps Physically relocate the host now. Update Netbox Update Netbox device page with new rack location De... [15:00:23] FIRING: [7x] GnmiInterfaceCountersDrop: asw1-b13-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [15:03:20] (03CR) 10Cathal Mooney: "Thanks for the review, I've updated the query and changed the time to 15s as advised." [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [15:03:29] (03CR) 10Cathal Mooney: Add alert on percentage of timeouts from POPs in blackbox pings (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [15:10:24] RESOLVED: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d2-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [15:11:00] !log import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia [15:11:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:14:30] (03CR) 10CI reject: [V:04-1] Remove revisionId from term fallback cache lines for properties [extensions/Wikibase] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333866 (https://phabricator.wikimedia.org/T434204) (owner: 10Thiemo Kreuz (WMDE)) [15:15:30] (03CR) 10Clément Goubert: [C:03+1] mcrouter: update to 1.3.6 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333855 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:15:42] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814 (10cmooney) 03NEW p:05Triage→03Medium [15:15:46] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru [15:15:53] reminder we'll be starting the dc live test momentarily, ping us if anything (!) [15:17:15] (03CR) 10Clément Goubert: "Can you document `cache.mcrouter.sidecar` in the module's values.yaml please :)" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333856 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:17:37] !log sukhe@lvs1020:~$ sudo systemctl restart pybal.service [15:17:37] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:17:55] (03CR) 10Clément Goubert: "It automarked as resolved for some reason" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333856 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:20:16] (03CR) 10Brouberol: [C:03+1] airflow-test-k8s: remove growthbook-sync-related access [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333867 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [15:20:22] (03PS3) 10Scott French: mesh: Support Host splits in configuration 1.16.0, networkpolicy 1.3.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331858 (https://phabricator.wikimedia.org/T427666) [15:21:08] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad [15:21:16] RECOVERY - Check if Pybal has been restarted after pybal.conf was changed on lvs1020 is OK: OK: pybal.service was restarted after /etc/pybal/pybal.conf was changed. https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [15:21:25] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad [15:22:07] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad [15:22:09] !log root@deploy1003 Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - T436781 [15:22:10] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad [15:22:12] T436781: Live-Test - September 2nd 2026 - https://phabricator.wikimedia.org/T436781 [15:22:46] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad [15:23:58] (03CR) 10Tryvix1509: wmf-config/core-Permissions.php: sort keys alphabetically (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1167927 (owner: 10LD) [15:24:07] (03CR) 10Clément Goubert: sidecars: Specify restartPolicy: Always. (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:24:34] (03PS1) 10Ladsgroup: turnilo: Order entries alphabetically [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333897 (https://phabricator.wikimedia.org/T435634) [15:27:13] !log cmooney@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003 [15:27:48] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru [15:28:14] !log cmooney@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003 [15:29:42] (03CR) 10Atsuko: [C:03+2] airflow-test-k8s: remove growthbook-sync-related access [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333867 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [15:30:23] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b13-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [15:31:46] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin [15:32:11] (03Merged) 10jenkins-bot: airflow-test-k8s: remove growthbook-sync-related access [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333867 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [15:34:07] (03CR) 10MVernon: [C:03+1] "I still think the question of "is anything paying any attention to max-age?" should be asked :)" [puppet] - 10https://gerrit.wikimedia.org/r/1327104 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [15:34:13] (03CR) 10Eevans: [C:03+1] Puppet 8: Replace unscoped legacy facts in module cassandra [puppet] - 10https://gerrit.wikimedia.org/r/1332517 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:35:23] RESOLVED: [2x] GnmiInterfaceCountersDrop: asw1-b13-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [15:36:16] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814#12281602 (10Jclark-ctr) @cmooney can these be pulled at anytime or do you want to be online when pulled? [15:38:02] (03PS4) 10Andrew Bogott: Add cloudvps-tenant-usage.py [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) [15:38:23] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply [15:39:11] !log atsuko@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply [15:39:24] (03PS5) 10Andrew Bogott: Add cloudvps-tenant-usage.py [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) [15:40:53] FIRING: [3x] GnmiInterfaceCountersDrop: asw1-b13-drmrs is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [15:41:18] (03CR) 10Btullis: [C:03+2] logstash: Consume the webrequest.dumps.v1 stream from Kafka [puppet] - 10https://gerrit.wikimedia.org/r/1328205 (https://phabricator.wikimedia.org/T425087) (owner: 10Btullis) [15:42:21] !log sukhe@lvs1019:~$ sudo systemctl restart pybal.service [15:42:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:42:42] RECOVERY - Check if Pybal has been restarted after pybal.conf was changed on lvs1019 is OK: OK: pybal.service was restarted after /etc/pybal/pybal.conf was changed. https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [15:44:09] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin [15:45:58] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin [15:46:04] (03PS2) 10Blake: mcrouter: Upgrade to 1.3.6. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333856 (https://phabricator.wikimedia.org/T417800) [15:46:39] !log slyngshede@cumin1003 END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad [15:47:36] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad [15:51:24] (03CR) 10Blake: "Please let me know if the comment here is sufficient, thanks!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333856 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:53:21] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad [15:53:29] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad [15:53:31] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad [15:53:47] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad [15:53:47] !log slyngshede@cumin1003 [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918 [15:54:03] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad [15:54:04] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921#12281684 (10VRiley-WMF) Starting the first three cloudvirt1067 D5 - U09 CableID: 5311 Port: 7 cloudvirt1066 D5 - U08 CableID: 5312 Port: 8 cloudvirt1065 D5 - U0... [15:54:16] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad [15:55:06] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad [15:55:26] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad [15:55:40] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad [15:55:46] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad [15:55:52] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad [15:56:11] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad [15:56:13] !log slyngshede@cumin1003 [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320 [15:56:15] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad [15:56:26] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad [15:56:27] !log root@deploy1003 helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync [15:56:50] !log root@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync [15:56:52] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad [15:56:59] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad [15:57:00] !log root@deploy1003 helmfile [codfw] START helmfile.d/services/mw-cron: apply [15:57:04] !log root@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-cron: apply [15:57:05] !log root@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-cron: apply [15:57:12] !log root@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply [15:57:14] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad [15:57:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?panelId=18&fullscreen&orgId=1&var-datasource=codfw%20prometheus/ops - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [15:57:29] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad [15:57:37] <_joe_> uhm high erros from the api [15:58:10] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad [15:58:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 25% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:58:21] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin [15:58:35] (03CR) 10Brouberol: [C:03+1] "Good call thanks! I'll deploy this along with a change to https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/blob/main/wmf" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333173 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [15:58:35] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad [15:58:43] (03CR) 10Brouberol: [C:03+2] airflow: Use urldownloader LVS endpoint [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333173 (https://phabricator.wikimedia.org/T429175) (owner: 10Clément Goubert) [15:58:45] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, and 2 others: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12281711 (10elukey) I did a bit more digging to refresh my memory, and IIUC in prod we use the class `cpufrequtils` to set... [15:59:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.28% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:59:21] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin [16:02:15] RESOLVED: [4x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [16:05:24] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12281720 (10FCeratto-WMF) Jose confirmed the SSH key over Slack and they match. [16:05:41] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12281721 (10FCeratto-WMF) [16:09:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:10:07] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad [16:10:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:10:16] !log slyngshede@cumin1003 START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad [16:10:18] !log root@deploy1003 Forcefully removing global lock: Datacenter switchover from codfw to eqiad - T436781 [16:10:18] !log root@deploy1003 Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - T436781 (duration: 48m 09s) [16:10:19] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad [16:10:25] T436781: Live-Test - September 2nd 2026 - https://phabricator.wikimedia.org/T436781 [16:11:32] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin [16:11:42] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:14:13] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs [16:15:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.9% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:24:28] (03PS3) 10Scott French: modules: Prepare mesh.configuration / networkpolicy minor version bump [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328738 (https://phabricator.wikimedia.org/T427666) [16:24:46] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs [16:26:17] (03PS6) 10Scott French: api-gateway: Basic cluster specifier support and Lua plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) [16:26:17] (03PS8) 10Scott French: api-gateway: Support x-wikimedia-debug routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311964 (https://phabricator.wikimedia.org/T433752) [16:26:17] (03PS5) 10Scott French: api-gateway: Support Host-based diversion in the mw-api plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322140 (https://phabricator.wikimedia.org/T433752) [16:28:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 7.817% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [16:29:25] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs [16:33:49] 06SRE, 10SRE-Access-Requests, 06Data-Engineering, 13Patch-For-Review: Requesting Kerberos access for kamila - https://phabricator.wikimedia.org/T436703#12281847 (10Ahoelzl) Approved. [16:34:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 914.5ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [16:34:23] !log vriley@cumin1003 START - Cookbook sre.dns.netbox [16:38:25] FIRING: SystemdUnitFailed: check_netbox_uncommitted_dns_changes.service on netbox1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:38:55] !log vriley@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003" [16:39:00] !log vriley@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003" [16:39:00] !log vriley@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:40:16] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12281895 (10taavi) 05Open→03Declined You need to sign up for an account at lists.wikimedia.org if you do not have one already. [16:41:27] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs [16:41:30] !log vriley@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065 [16:41:58] !log vriley@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065 [16:43:25] RESOLVED: SystemdUnitFailed: check_netbox_uncommitted_dns_changes.service on netbox1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:46:34] !log vriley@cumin1003 START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [16:48:30] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin [16:49:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 811.8ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [16:49:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:51:46] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814#12282002 (10cmooney) >>! In T436814#12281602, @Jclark-ctr wrote: > @cmooney can these be pulled at anytime or do you want to be online when pulled? T... [16:59:33] !log vriley@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1700) [17:00:38] !log vriley@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie [17:00:51] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin [17:00:52] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921#12282130 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1065.eqiad.wmnet with OS trixie [17:06:52] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [17:20:02] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin [17:27:01] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12282373 (10Romaine) 05Declined→03Open Having signed up at lists.wikimedia.org does not solve the described issue. [17:29:34] jouncebot: nowandnext [17:29:35] For the next 0 hour(s) and 30 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1700) [17:29:35] In 0 hour(s) and 30 minute(s): MediaWiki train - Utc-7+Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1800) [17:30:45] a bit later in the infra window than planned, but I will be merging some noop (i.e., new feature is disabled) rest-gateway changes shortly [17:31:05] (03CR) 10Scott French: "Thanks for the review!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:31:08] (03CR) 10Scott French: [C:03+2] api-gateway: Basic cluster specifier support and Lua plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:33:49] (03Merged) 10jenkins-bot: api-gateway: Basic cluster specifier support and Lua plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:35:29] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [17:35:47] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [17:38:00] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921#12282448 (10VRiley-WMF) @fgiunchedi while trying to image the first server, I think it may be getting stuck on the RAID portion. Will these cloudvirts need a speci... [17:38:04] !log swfrench@deploy1003 helmfile [codfw] START helmfile.d/services/rest-gateway: apply [17:38:28] !log swfrench@deploy1003 helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply [17:45:13] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin [17:46:03] (03PS2) 10Jforrester: [WikiLambda] Log the …Orchestrator and …AbstractClient channels too [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324368 [17:47:44] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12282505 (10Aklapper) After signing in, what does happen / what is shown when you go to https://lists.wikimedia.org/postorius/lists/wikimediabe-l.lists.wikimedia.org/members/owner/ ? [17:47:58] !log swfrench@deploy1003 helmfile [eqiad] START helmfile.d/services/rest-gateway: apply [17:48:27] !log swfrench@deploy1003 helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply [17:49:25] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad [17:52:12] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921#12282514 (10VRiley-WMF) cloudvirt1072 D5 - U10 CableID: 5309 Port: 9 cloudvirt1073 D5 - U11 CableID: 5299 Port: 38 cloudvirt1074 D5 - U13 CableID: 5301 Port: 40 [17:53:35] I'm done with production deployments for now. you'll see some additional helmfile applies while I perform some tests in staging, but that's all. [17:55:01] (03CR) 10DLynch: "Sounds helpful for avoiding the mixup in the future, thanks! I'm happy to rename it now if that's happening immediately, or just get swept" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [17:55:09] (03PS3) 10DLynch: editcheck-headless: Add deployment for the technical pilot [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) [17:58:19] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [17:58:43] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [18:00:05] dancy and hashar: Deploy window MediaWiki train - Utc-7+Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T1800) [18:00:21] o/\ [18:00:47] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.18 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333978 (https://phabricator.wikimedia.org/T430837) [18:00:50] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by dancy@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333978 (https://phabricator.wikimedia.org/T430837) (owner: 10TrainBranchBot) [18:01:47] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.18 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333978 (https://phabricator.wikimedia.org/T430837) (owner: 10TrainBranchBot) [18:02:46] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [18:02:57] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [18:05:15] (03CR) 10CI reject: [V:04-1] editcheck-headless: Add deployment for the technical pilot [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [18:06:09] dancy: ouch, is your arm okay? [18:06:18] haha [18:06:20] mangled! [18:06:29] Crazy train accident [18:09:37] (03CR) 10DLynch: "CI failure looks unrelated. It's a 503 fetching `https://helm-charts.wikimedia.org/stable/charts/kserve-resources-0.1.1.tgz`." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [18:10:11] dancy, does that make /you/ the train blocker? [18:12:37] clearly we should require wearing a high-vis vest when working near trains, that'll make things safer [18:13:27] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1332767 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [18:14:00] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [18:14:15] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad [18:14:42] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [18:16:28] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [18:16:29] !log vriley@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie [18:16:40] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [18:16:49] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921#12282585 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1065.eqiad.wmnet with OS trixie executed with er... [18:17:02] !log dancy@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs T430837 [18:17:05] T430837: 1.47.0-wmf.18 deployment blockers - https://phabricator.wikimedia.org/T430837 [18:19:09] (03CR) 10Ladsgroup: "To my understand, that's how CDN caching sets the TTL, it removes it from header but stores it for that long internally. I can double chec" [puppet] - 10https://gerrit.wikimedia.org/r/1327104 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [18:19:28] Speaking of unsafe train operations, we have a new train blocker: https://phabricator.wikimedia.org/T436851 [18:20:09] (03PS1) 10Cwhite: opensearch: use chained certificate and supply root to cluster [puppet] - 10https://gerrit.wikimedia.org/r/1333991 (https://phabricator.wikimedia.org/T436394) [18:20:11] (03PS1) 10Cwhite: opensearch: use pkcs8-formatted key [puppet] - 10https://gerrit.wikimedia.org/r/1333992 (https://phabricator.wikimedia.org/T436393) [18:20:11] ... I jinxed it didn't I [18:20:14] (03PS1) 10Cwhite: profile: convert admin key to pkcs8 [puppet] - 10https://gerrit.wikimedia.org/r/1333993 (https://phabricator.wikimedia.org/T436393) [18:20:17] (03PS1) 10Cwhite: opensearch: curator to use root ca path [puppet] - 10https://gerrit.wikimedia.org/r/1333994 (https://phabricator.wikimedia.org/T436394) [18:20:19] (03PS1) 10Cwhite: opensearch-dashboards: use root ca to validate connections to OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1333995 (https://phabricator.wikimedia.org/T436394) [18:20:21] (03PS1) 10Cwhite: prometheus: es-exporter use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333996 (https://phabricator.wikimedia.org/T436394) [18:20:23] (03PS1) 10Cwhite: profile: logstash opensearch outputs use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333997 (https://phabricator.wikimedia.org/T436394) [18:20:26] (03PS1) 10Cwhite: profile: prometheus elasticsearch exporter use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333998 (https://phabricator.wikimedia.org/T436394) [18:21:59] (03CR) 10CI reject: [V:04-1] profile: convert admin key to pkcs8 [puppet] - 10https://gerrit.wikimedia.org/r/1333993 (https://phabricator.wikimedia.org/T436393) (owner: 10Cwhite) [18:23:58] (03PS2) 10Cwhite: profile: convert admin key to pkcs8 [puppet] - 10https://gerrit.wikimedia.org/r/1333993 (https://phabricator.wikimedia.org/T436393) [18:24:41] (03PS1) 10Ssingh: Revert^3 "wikimedia.org: add TXT record for BIMI" [dns] - 10https://gerrit.wikimedia.org/r/1334000 [18:25:18] (03PS1) 10Ladsgroup: cache: Reduce the webp threshold to 70 [puppet] - 10https://gerrit.wikimedia.org/r/1334001 (https://phabricator.wikimedia.org/T431150) [18:26:05] (03PS2) 10Cwhite: opensearch: use chained certificate and supply root to cluster [puppet] - 10https://gerrit.wikimedia.org/r/1333991 (https://phabricator.wikimedia.org/T436394) [18:26:32] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [18:26:43] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [18:28:31] (03CR) 10Ssingh: "Turning it off, as asked by Fundraising." [dns] - 10https://gerrit.wikimedia.org/r/1334000 (owner: 10Ssingh) [18:28:33] (03CR) 10Ssingh: [C:03+2] Revert^3 "wikimedia.org: add TXT record for BIMI" [dns] - 10https://gerrit.wikimedia.org/r/1334000 (owner: 10Ssingh) [18:28:36] !log sukhe@dns1004 START - running authdns-update [18:31:02] !log sukhe@dns1004 END - running authdns-update [18:35:20] dancy: I think the UW change should just be reverted for now. [18:38:49] (03PS3) 10Ladsgroup-claude: Package the project and stop using pkg_resources [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333246 [18:38:49] (03PS2) 10Ladsgroup-claude: Route all logging through thumbor's logger [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333247 [18:38:49] (03PS2) 10Ladsgroup-claude: Tidy up resource and rounding leftovers [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333248 [18:38:50] (03PS2) 10Ladsgroup-claude: Replace python-memcached with pymemcache [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333249 [18:38:51] (03PS2) 10Ladsgroup-claude: Call ImageMagick 7's magick instead of the deprecated convert [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333250 [18:38:52] (03PS2) 10Ladsgroup-claude: Clean up requirements and add a dependency bump path [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333251 [18:38:56] (03PS2) 10Ladsgroup-claude: Add a unit test layer that runs without the container [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333252 [18:39:00] (03PS2) 10Ladsgroup-claude: Replace flake8 with ruff and apply its fixes [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333253 [18:39:04] (03PS2) 10Ladsgroup-claude: Apply ruff format [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333254 [18:41:33] James_F: Ok. I see that you're already on it. Thank you [18:44:06] (03PS1) 10Jforrester: Revert "Remove fallback to Special:Upload and redirect user to alternative form" [extensions/UploadWizard] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1334006 (https://phabricator.wikimedia.org/T436851) [18:44:19] dancy: Let’s deploy the revert to wmf.18 at least? [18:45:34] (03CR) 10Ladsgroup: "Sorry. I assumed tests would build the thumbor so something like this should be caught. I ran it locally and the build works now." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333246 (owner: 10Ladsgroup-claude) [18:45:44] I'm afk at the moment. Will be back in 15. Feel free to backport before then [18:45:56] Sure. [18:46:11] ty [18:47:04] Eurgh, it’s an i18n-touching commit. Deploy will be slow. [18:47:11] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy1003 using scap backport" [extensions/UploadWizard] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1334006 (https://phabricator.wikimedia.org/T436851) (owner: 10Jforrester) [18:48:06] (03CR) 10Jforrester: "PS2: Dropped the i18n update so this is faster to deploy; it's relatively trivial." [extensions/UploadWizard] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1334006 (https://phabricator.wikimedia.org/T436851) (owner: 10Jforrester) [18:48:09] (03PS2) 10Cwhite: opensearch: use pkcs8-formatted key [puppet] - 10https://gerrit.wikimedia.org/r/1333992 (https://phabricator.wikimedia.org/T436393) [18:48:16] (03PS2) 10Jforrester: Revert "Remove fallback to Special:Upload and redirect user to alternative form" [extensions/UploadWizard] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1334006 (https://phabricator.wikimedia.org/T436851) [18:48:24] (03PS3) 10Cwhite: profile: convert admin key to pkcs8 [puppet] - 10https://gerrit.wikimedia.org/r/1333993 (https://phabricator.wikimedia.org/T436393) [18:48:32] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy1003 using scap backport" [extensions/UploadWizard] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1334006 (https://phabricator.wikimedia.org/T436851) (owner: 10Jforrester) [18:48:41] (03PS2) 10Cwhite: opensearch: curator to use root ca path [puppet] - 10https://gerrit.wikimedia.org/r/1333994 (https://phabricator.wikimedia.org/T436394) [18:48:47] (03PS2) 10Cwhite: opensearch-dashboards: use root ca to validate connections to OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1333995 (https://phabricator.wikimedia.org/T436394) [18:48:54] (03PS2) 10Cwhite: prometheus: es-exporter use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333996 (https://phabricator.wikimedia.org/T436394) [18:48:59] (03PS2) 10Cwhite: profile: logstash opensearch outputs use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333997 (https://phabricator.wikimedia.org/T436394) [18:49:06] (03PS2) 10Cwhite: profile: prometheus elasticsearch exporter use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333998 (https://phabricator.wikimedia.org/T436394) [18:52:29] (03CR) 10Ladsgroup: [V:03+2 C:03+2] cache: Reduce the webp threshold to 70 [puppet] - 10https://gerrit.wikimedia.org/r/1334001 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [18:52:36] (03Merged) 10jenkins-bot: Revert "Remove fallback to Special:Upload and redirect user to alternative form" [extensions/UploadWizard] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1334006 (https://phabricator.wikimedia.org/T436851) (owner: 10Jforrester) [18:52:59] (03PS3) 10Ladsgroup: swift: Also accept thumb.wikimedia.org in ensure_max_mage [puppet] - 10https://gerrit.wikimedia.org/r/1327104 (https://phabricator.wikimedia.org/T427465) [18:53:01] !log jforrester@deploy1003 Started scap sync-world: Backport for [[gerrit:1334006|Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] [18:53:04] T436851: InvalidArgumentException: $text must be a string. - https://phabricator.wikimedia.org/T436851 [18:53:04] (03CR) 10Ladsgroup: [V:03+2 C:03+2] swift: Also accept thumb.wikimedia.org in ensure_max_mage [puppet] - 10https://gerrit.wikimedia.org/r/1327104 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [18:55:33] !log deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency [18:55:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:57:13] !log jforrester@deploy1003 jforrester: Backport for [[gerrit:1334006|Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:57:58] !log jforrester@deploy1003 jforrester: Continuing with deployment [18:59:35] (03CR) 10Ladsgroup: [V:03+2 C:03+2] swift: Also accept thumb.wikimedia.org in ensure_max_mage (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1327104 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [19:01:57] (03PS3) 10Cwhite: opensearch: use chained certificate and supply root to cluster [puppet] - 10https://gerrit.wikimedia.org/r/1333991 (https://phabricator.wikimedia.org/T436394) [19:01:57] (03PS3) 10Cwhite: opensearch: use pkcs8-formatted key [puppet] - 10https://gerrit.wikimedia.org/r/1333992 (https://phabricator.wikimedia.org/T436393) [19:01:57] (03PS4) 10Cwhite: profile: convert admin key to pkcs8 [puppet] - 10https://gerrit.wikimedia.org/r/1333993 (https://phabricator.wikimedia.org/T436393) [19:01:58] (03PS3) 10Cwhite: opensearch: curator to use root ca path [puppet] - 10https://gerrit.wikimedia.org/r/1333994 (https://phabricator.wikimedia.org/T436394) [19:01:59] (03PS3) 10Cwhite: opensearch-dashboards: use root ca to validate connections to OpenSearch [puppet] - 10https://gerrit.wikimedia.org/r/1333995 (https://phabricator.wikimedia.org/T436394) [19:02:00] (03PS3) 10Cwhite: prometheus: es-exporter use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333996 (https://phabricator.wikimedia.org/T436394) [19:02:04] (03PS3) 10Cwhite: profile: logstash opensearch outputs use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333997 (https://phabricator.wikimedia.org/T436394) [19:02:08] (03PS3) 10Cwhite: profile: prometheus elasticsearch exporter use root ca [puppet] - 10https://gerrit.wikimedia.org/r/1333998 (https://phabricator.wikimedia.org/T436394) [19:08:44] FIRING: KubernetesDeploymentUnavailableReplicas: ... [19:08:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [19:08:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [19:09:57] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332513 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:10:29] (03PS2) 10JHathaway: Puppet 8: Replace unscoped legacy facts in module elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1332513 (https://phabricator.wikimedia.org/T435225) [19:10:31] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332513 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:12:31] (03CR) 10Ladsgroup: "This touches swift, I think Matthew should be aware." [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [19:16:18] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module icinga [puppet] - 10https://gerrit.wikimedia.org/r/1332522 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:16:54] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module ceph [puppet] - 10https://gerrit.wikimedia.org/r/1332518 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:17:36] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module graphite [puppet] - 10https://gerrit.wikimedia.org/r/1332521 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:22:51] (03CR) 10JHathaway: [C:03+2] "@ltoscano@wikimedia.org that is exactly correct, apologies if that was not more clear in the commit message @mvernon@wikimedia.org" [puppet] - 10https://gerrit.wikimedia.org/r/1332519 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:23:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 846.2ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [19:24:04] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module cassandra [puppet] - 10https://gerrit.wikimedia.org/r/1332517 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:24:35] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module opensearch [puppet] - 10https://gerrit.wikimedia.org/r/1332514 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:27:50] (03PS1) 10Blake: httpd: Add version 1.0.3 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334012 (https://phabricator.wikimedia.org/T417800) [19:28:01] (03PS1) 10Blake: httpd: Add switch to enable Kubernetes native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334013 (https://phabricator.wikimedia.org/T417800) [19:28:09] (03PS1) 10Blake: php-fpm: Add version 1.0.2 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334015 (https://phabricator.wikimedia.org/T417800) [19:28:17] (03PS1) 10Blake: php-fpm: Make php-fpm-exporter a sidecar. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334016 (https://phabricator.wikimedia.org/T417800) [19:28:54] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module service [puppet] - 10https://gerrit.wikimedia.org/r/1332511 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:29:37] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module envoyproxy [puppet] - 10https://gerrit.wikimedia.org/r/1332509 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:30:19] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in profile dns [puppet] - 10https://gerrit.wikimedia.org/r/1332508 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:30:59] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in profile cache [puppet] - 10https://gerrit.wikimedia.org/r/1332507 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:31:39] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module varnish [puppet] - 10https://gerrit.wikimedia.org/r/1332506 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:32:01] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module dnsrecursor [puppet] - 10https://gerrit.wikimedia.org/r/1332505 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:32:43] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12282889 (10Jhancock.wm) [19:34:03] (03PS2) 10JHathaway: Puppet 8: Replace unscoped legacy facts in module acme_chief [puppet] - 10https://gerrit.wikimedia.org/r/1332504 (https://phabricator.wikimedia.org/T435225) [19:34:18] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332504 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:37:02] (03CR) 10LD: "recheck" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1167927 (owner: 10LD) [19:38:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 848.4ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [19:41:08] FIRING: GnmiInterfaceCountersDrop: ... [19:41:08] cr2-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=cr2-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [19:43:27] James_F: Sorry for the extra-slow deployment. I'm working on a fix for that. [19:46:35] Ha, yes, no worries. [19:53:31] (03CR) 10DLynch: "recheck" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [19:55:34] !log bking@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade T431608 [19:55:38] T431608: Upgrade Matomo from 4.16.1 to 5.12.0 (latest stable) - https://phabricator.wikimedia.org/T431608 [19:57:28] !log jforrester@deploy1003 Finished scap sync-world: Backport for [[gerrit:1334006|Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s) [19:57:32] T436851: InvalidArgumentException: $text must be a string. - https://phabricator.wikimedia.org/T436851 [20:00:04] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: That opportune time for a UTC late backport window deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T2000). [20:00:04] No Gerrit patches in the queue for this window AFAICS. [20:12:09] (03PS1) 10Krinkle: Remove the last vestiges of $wgVirtualRestConfig [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334032 (https://phabricator.wikimedia.org/T436054) [20:13:44] RESOLVED: KubernetesDeploymentUnavailableReplicas: ... [20:13:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [20:13:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [20:14:12] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12282999 (10Nux) Interestingly thumbs work. So when the image you upload is big (like 400x400) the thumb... [20:14:14] FIRING: KubernetesDeploymentUnavailableReplicas: ... [20:14:14] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [20:14:14] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [20:16:12] !log dancy@deploy1003 Installing scap version "4.288.0" for 3 host(s) [20:17:26] (03PS2) 10Kamila Součková: admin: add kamila to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1333280 (https://phabricator.wikimedia.org/T436703) [20:18:06] (03PS1) 10Ottomata: kafka: run MirrorMaker in docker for non k8s environments [puppet] - 10https://gerrit.wikimedia.org/r/1334033 (https://phabricator.wikimedia.org/T432089) [20:18:07] !log dancy@deploy1003 Installation of scap version "4.288.0" completed for 3 hosts [20:18:37] !log dancy@deploy1003 Started scap sync-world: testing [20:19:14] RESOLVED: KubernetesDeploymentUnavailableReplicas: ... [20:19:14] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [20:19:14] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [20:27:58] !log dancy@deploy1003 Finished scap sync-world: testing (duration: 09m 21s) [20:28:30] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.1% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:31:00] I don't love that PHPFPMTooBusy fires pretty regularly when we deploy MW these days [20:31:22] I think it's just that we're typically cruising closer to the red line than we used to [20:32:44] rzl: I noticed that the progress report would go up and down during the tail end of the deployment.. [20:33:04] rzl: (implying that pods were repeatedly going in and out of readiness over several minutes) [20:33:18] mm. I don't love that either [20:33:24] rzl: Example: https://spiderpig.wikimedia.org/jobs/2734 [20:33:57] Unfortunately I don't know what services are the stragglers [20:34:38] I'm also zoomed out to the last week and it looks like something got worse pretty sharply at 2026-09-01 04:00 or so, both in PHP active workers and in the latency histogram: [20:34:39] https://grafana.wikimedia.org/d/35WSHOjVk/application-servers-red-k8s?orgId=1&refresh=1m&from=now-7d&to=now&timezone=utc&var-site=eqiad&var-deployment=mw-web&var-method=GET&var-code=200&var-service=mediawiki [20:34:56] (meaning we have less headroom to play with during a deploy, before that alert fires) [20:36:15] basically I don't want deployers to get in the habit of ignoring that :) that can be the first sign that the new release has some major problem and we're about to have an outage [20:36:34] but right now it doesn't signal that reliably, because we dance over that threshold and back all the time [20:37:00] Nod. I see PHPFPMTooBusy too often [20:49:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:50:54] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [20:51:23] ^ this is "-l chart=function-evaluator", only has an envoy version bump [21:00:05] Deploy window Wikifunctions Services UTC Late (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T2100) [21:01:06] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:01:21] (timed out again, still digging) [21:01:45] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:01:55] (same again, evaluator only) [21:03:28] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:04:01] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [21:04:49] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:06:20] rzl: Did that work that time? [21:06:22] James_F: I'll update on task, but the evaluators are now updated in both DCs (cleanly, after I kicked the stuck pod in eqiad) and I'm hands-off -- over to you for the window, if you want to update the orchestrator [21:06:29] Aha, perfect. [21:06:42] I’ll see what breaks with the orchestrator this time. :-) [21:06:46] glhf! [21:06:52] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [21:07:00] (it'll still have the envoy diff, which is still good to deploy) [21:07:08] Ack. [21:07:38] (03PS1) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256 (try 2) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334052 (https://phabricator.wikimedia.org/T426337) [21:08:16] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256 (try 2) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334052 (https://phabricator.wikimedia.org/T426337) (owner: 10Jforrester) [21:10:37] (03Merged) 10jenkins-bot: wikifunctions: Upgrade orchestrator from 2026-08-19-165650 to 2026-09-01-120256 (try 2) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334052 (https://phabricator.wikimedia.org/T426337) (owner: 10Jforrester) [21:11:11] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:11:20] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:11:33] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:11:36] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module acme_chief [puppet] - 10https://gerrit.wikimedia.org/r/1332504 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [21:11:41] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:11:51] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [21:12:34] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:12:39] (03PS1) 10Eevans: install_server: preseed partman-basicfilesystems/no_swap [puppet] - 10https://gerrit.wikimedia.org/r/1334054 (https://phabricator.wikimedia.org/T436396) [21:13:37] (03PS3) 10JHathaway: Puppet 8: Replace unscoped legacy facts in module elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1332513 (https://phabricator.wikimedia.org/T435225) [21:13:41] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332513 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [21:14:22] (03CR) 10Eevans: [C:03+2] install_server: preseed partman-basicfilesystems/no_swap [puppet] - 10https://gerrit.wikimedia.org/r/1334054 (https://phabricator.wikimedia.org/T436396) (owner: 10Eevans) [21:15:27] (03PS3) 10Jforrester: wikifunctions: Enable callbacks on production, not just staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333859 (https://phabricator.wikimedia.org/T421848) [21:16:11] (03CR) 10Jforrester: [C:03+2] wikifunctions: Enable callbacks on production, not just staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333859 (https://phabricator.wikimedia.org/T421848) (owner: 10Jforrester) [21:17:32] (03CR) 10Cwhite: [C:03+1] "🎉" [puppet] - 10https://gerrit.wikimedia.org/r/1332767 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [21:18:43] (03Merged) 10jenkins-bot: wikifunctions: Enable callbacks on production, not just staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333859 (https://phabricator.wikimedia.org/T421848) (owner: 10Jforrester) [21:18:50] (03CR) 10Scott French: "This is part 1 of a 2-patch series to introduce support for `splits` in the relevant mesh modules." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328738 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [21:18:56] (03CR) 10Scott French: "This is part 2 of the series, and actually introduces the new feature to the relevant modules. Thanks!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331858 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [21:19:27] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [21:19:34] (03CR) 10Cwhite: [C:03+1] "LGTM! Adding Brian from search for awareness." [puppet] - 10https://gerrit.wikimedia.org/r/1332772 (owner: 10Majavah) [21:19:35] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [21:19:45] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [21:21:22] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module elasticsearch [puppet] - 10https://gerrit.wikimedia.org/r/1332513 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [21:21:29] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm [21:22:21] (03CR) 10Cwhite: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1332773 (owner: 10Majavah) [21:23:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.21% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:23:25] FIRING: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:24:08] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [21:24:20] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [21:25:25] FIRING: [3x] ProbeDown: Service aqs1017-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:28:25] RESOLVED: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:28:31] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324368 (owner: 10Jforrester) [21:28:31] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jforrester@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333732 (https://phabricator.wikimedia.org/T435637) (owner: 10Jforrester) [21:28:44] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [21:29:22] (03Merged) 10jenkins-bot: [WikiLambda] Log the …Orchestrator and …AbstractClient channels too [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324368 (owner: 10Jforrester) [21:29:33] (03CR) 10Kamila Součková: "Will do, thank you!" [puppet] - 10https://gerrit.wikimedia.org/r/1333280 (https://phabricator.wikimedia.org/T436703) (owner: 10Kamila Součková) [21:29:51] (03Merged) 10jenkins-bot: wikifunctions: Set up the functionmaintainer right for the community [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1333732 (https://phabricator.wikimedia.org/T435637) (owner: 10Jforrester) [21:30:14] !log jforrester@deploy1003 Started scap sync-world: Backport for [[gerrit:1324368|[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732|wikifunctions: Set up the functionmaintainer right for the community (T435637)]] [21:30:17] T435637: Prepare the Maintainer role for community management of Types - https://phabricator.wikimedia.org/T435637 [21:30:25] FIRING: [4x] ProbeDown: Service aqs1017-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:30:37] (03CR) 10Cwhite: [C:03+2] opensearch: use chained certificate and supply root to cluster [puppet] - 10https://gerrit.wikimedia.org/r/1333991 (https://phabricator.wikimedia.org/T436394) (owner: 10Cwhite) [21:33:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:34:18] !log bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes T431608 [21:34:21] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:34:22] T431608: Upgrade Matomo from 4.16.1 to 5.12.0 (latest stable) - https://phabricator.wikimedia.org/T431608 [21:34:57] !log jforrester@deploy1003 jforrester: Backport for [[gerrit:1324368|[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732|wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:35:27] !log jforrester@deploy1003 jforrester: Continuing with deployment [21:38:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.45% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [21:39:05] (03PS1) 10Krinkle: Extract MathJax DOM filter into a separate file [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1334068 (https://phabricator.wikimedia.org/T435274) [21:39:24] (03PS1) 10Krinkle: Respect contextual binomial sizing [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1334070 (https://phabricator.wikimedia.org/T434477) [21:40:03] !log jforrester@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324368|[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732|wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s) [21:40:10] T435637: Prepare the Maintainer role for community management of Types - https://phabricator.wikimedia.org/T435637 [21:40:58] (03CR) 10Jforrester: "<3" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334032 (https://phabricator.wikimedia.org/T436054) (owner: 10Krinkle) [21:42:06] James_F: let me know if I can deploy after you in a few minutes. [21:44:49] !log bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes T431608 [21:44:52] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:44:53] T431608: Upgrade Matomo from 4.16.1 to 5.12.0 (latest stable) - https://phabricator.wikimedia.org/T431608 [21:47:36] Krinkle: Go for it. [21:48:29] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1334068 (https://phabricator.wikimedia.org/T435274) (owner: 10Krinkle) [21:48:30] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1334070 (https://phabricator.wikimedia.org/T434477) (owner: 10Krinkle) [21:48:31] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334032 (https://phabricator.wikimedia.org/T436054) (owner: 10Krinkle) [21:49:29] (03Merged) 10jenkins-bot: Remove the last vestiges of $wgVirtualRestConfig [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334032 (https://phabricator.wikimedia.org/T436054) (owner: 10Krinkle) [21:50:11] (03Merged) 10jenkins-bot: Extract MathJax DOM filter into a separate file [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1334068 (https://phabricator.wikimedia.org/T435274) (owner: 10Krinkle) [21:50:13] (03Merged) 10jenkins-bot: Respect contextual binomial sizing [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1334070 (https://phabricator.wikimedia.org/T434477) (owner: 10Krinkle) [21:50:25] RESOLVED: [4x] ProbeDown: Service aqs1017-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:50:39] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1334068|Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070|Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032|Remove the last vestiges of $wgVirtualRestConfig (T436054)]] [21:50:50] T435274: Missing table grid borders in Client side MathJax (check polyfills for SVG support) - https://phabricator.wikimedia.org/T435274 [21:50:51] T434477: \frac 1 {\binom{n}{k}} font size too large - https://phabricator.wikimedia.org/T434477 [21:50:51] T418144: MathJax rendering of binomial coefficients not correct - https://phabricator.wikimedia.org/T418144 [21:50:51] T401718: \textstyle not applied to \binom in MathML - https://phabricator.wikimedia.org/T401718 [21:50:52] T436054: Remove $wgVirtualRestConfig from MediaWiki - https://phabricator.wikimedia.org/T436054 [21:54:57] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1334068|Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070|Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032|Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:59:53] !log krinkle@deploy1003 krinkle: Continuing with deployment [22:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260902T2200) [22:01:22] hey Krinkle let me know when you are done! I have one patch to backport. [22:02:28] (03PS1) 10Cwhite: opensearch: use pki hiera setting for pki_root_ca_cn [puppet] - 10https://gerrit.wikimedia.org/r/1334080 [22:04:05] 06SRE, 10LDAP-Access-Requests: Grant Access to wmf, for WMF staff/contractors nda group for gsduser - https://phabricator.wikimedia.org/T435852#12283425 (10GSduser) 05Declined→03Open Is this going to give me access to the web request tables and other data cubes in turnilo and superset? If so, great, if no... [22:04:53] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1334068|Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070|Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032|Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s) [22:04:54] (03CR) 10Cwhite: [C:03+2] opensearch: use pki hiera setting for pki_root_ca_cn [puppet] - 10https://gerrit.wikimedia.org/r/1334080 (owner: 10Cwhite) [22:05:02] T435274: Missing table grid borders in Client side MathJax (check polyfills for SVG support) - https://phabricator.wikimedia.org/T435274 [22:05:03] T434477: \frac 1 {\binom{n}{k}} font size too large - https://phabricator.wikimedia.org/T434477 [22:05:03] T418144: MathJax rendering of binomial coefficients not correct - https://phabricator.wikimedia.org/T418144 [22:05:04] T401718: \textstyle not applied to \binom in MathML - https://phabricator.wikimedia.org/T401718 [22:05:04] T436054: Remove $wgVirtualRestConfig from MediaWiki - https://phabricator.wikimedia.org/T436054 [22:11:07] (03PS3) 10Ladsgroup: Tidy up resource and rounding leftovers [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333248 (owner: 10Ladsgroup-claude) [22:12:00] (03PS4) 10Ladsgroup: Tidy up resource and rounding leftovers [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333248 (owner: 10Ladsgroup-claude) [22:17:52] (03PS3) 10Ladsgroup: Call ImageMagick 7's magick instead of the deprecated convert [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333250 (owner: 10Ladsgroup-claude) [22:18:14] (03PS4) 10Ladsgroup: Call ImageMagick 7's magick instead of the deprecated convert [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333250 (owner: 10Ladsgroup-claude) [22:20:43] eevans@cumin1003 reimage (PID 2324620) is awaiting input [22:21:22] Krinkle: Are you done? [22:21:56] yes [22:22:07] Cool. [22:22:13] Jdlrobson: Over to you. [22:22:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.31% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:22:24] thanks Krinkle James_F [22:22:38] (03PS9) 10Scott French: api-gateway: Support x-wikimedia-debug routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311964 (https://phabricator.wikimedia.org/T433752) [22:22:38] (03PS6) 10Scott French: api-gateway: Support Host-based diversion in the mw-api plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322140 (https://phabricator.wikimedia.org/T433752) [22:22:52] (03PS3) 10Ladsgroup: Clean up requirements and add a dependency bump path [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333251 (owner: 10Ladsgroup-claude) [22:23:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [extensions/Campaigns] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333307 (https://phabricator.wikimedia.org/T436681) (owner: 10Jdlrobson) [22:24:04] (03PS3) 10Ladsgroup: Add a unit test layer that runs without the container [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333252 (owner: 10Ladsgroup-claude) [22:25:30] (03PS4) 10Ladsgroup: Add a unit test layer that runs without the container [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333252 (owner: 10Ladsgroup-claude) [22:26:02] (03Merged) 10jenkins-bot: Campaigns should not override existing campaign query strings [extensions/Campaigns] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1333307 (https://phabricator.wikimedia.org/T436681) (owner: 10Jdlrobson) [22:26:25] !log jdlrobson@deploy1003 Started scap sync-world: Backport for [[gerrit:1333307|Campaigns should not override existing campaign query strings (T436681)]] [22:26:28] T436681: Campaign extension overwrites existing campaign on create account link - https://phabricator.wikimedia.org/T436681 [22:27:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.93% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [22:27:18] (03PS5) 10Ladsgroup: Add a unit test layer that runs without the container [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333252 (owner: 10Ladsgroup-claude) [22:30:18] (03PS3) 10Ladsgroup: Replace flake8 with ruff and apply its fixes [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333253 (owner: 10Ladsgroup-claude) [22:30:42] !log jdlrobson@deploy1003 jdlrobson: Backport for [[gerrit:1333307|Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [22:32:53] !log jdlrobson@deploy1003 jdlrobson: Continuing with deployment [22:35:08] (03PS3) 10Ladsgroup: Replace python-memcached with pymemcache [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333249 (owner: 10Ladsgroup-claude) [22:37:28] !log jdlrobson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1333307|Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s) [22:37:31] T436681: Campaign extension overwrites existing campaign on create account link - https://phabricator.wikimedia.org/T436681 [22:39:26] (03PS1) 10Cwhite: opensearch: roll back certificate changes on beta-logs [puppet] - 10https://gerrit.wikimedia.org/r/1334099 (https://phabricator.wikimedia.org/T436394) [22:44:37] (03CR) 10Cwhite: [C:03+2] opensearch: roll back certificate changes on beta-logs [puppet] - 10https://gerrit.wikimedia.org/r/1334099 (https://phabricator.wikimedia.org/T436394) (owner: 10Cwhite) [22:44:48] (03CR) 10Ladsgroup: "Running it locally without docker was a bit of an effort because you need to install some pip dependencies but not all (some require binar" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1333252 (owner: 10Ladsgroup-claude) [22:47:33] done [23:06:05] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12283631 (10ecarg) thanks @RLazarus ! I have a draft I wanted to push up to `slothslos` but it seems I need permissions, would you... [23:14:39] !log eevans@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm [23:14:55] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm [23:20:46] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12283708 (10RLazarus) You should be able to fork the repo, push the changes to your fork, and open a merge request on the original... [23:27:25] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm [23:27:39] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm [23:38:44] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm [23:38:58] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm [23:40:52] jouncebot: nowandnext [23:40:53] No deployments scheduled for the next 6 hour(s) and 19 minute(s) [23:40:53] In 6 hour(s) and 19 minute(s): MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260903T0600) [23:40:53] In 6 hour(s) and 19 minute(s): Primary database switchover (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260903T0600) [23:41:08] FIRING: GnmiInterfaceCountersDrop: ... [23:41:08] cr2-eqiad is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=cr2-eqiad:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [23:41:15] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1334140 [23:41:15] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1334140 (owner: 10TrainBranchBot) [23:43:43] (03PS1) 10Dreamy Jazz: private/readme.php: Remove now removed secrets [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334141 (https://phabricator.wikimedia.org/T436880) [23:44:02] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334141 (https://phabricator.wikimedia.org/T436880) (owner: 10Dreamy Jazz) [23:44:52] (03Merged) 10jenkins-bot: private/readme.php: Remove now removed secrets [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334141 (https://phabricator.wikimedia.org/T436880) (owner: 10Dreamy Jazz) [23:45:11] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1334141|private/readme.php: Remove now removed secrets (T436880)]] [23:45:14] T436880: Archive the SimilarEditors extension - https://phabricator.wikimedia.org/T436880 [23:47:08] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1334140 (owner: 10TrainBranchBot) [23:49:33] !log dreamyjazz@deploy1003 dreamyjazz: Backport for [[gerrit:1334141|private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [23:50:57] !log dreamyjazz@deploy1003 dreamyjazz: Continuing with deployment [23:51:40] !log eevans@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage [23:55:32] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1334141|private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s) [23:55:35] T436880: Archive the SimilarEditors extension - https://phabricator.wikimedia.org/T436880 [23:57:36] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage [23:59:58] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown