[00:00:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [00:56:22] (03PS1) 10Tim Starling: Produnto: use wgCopyUploadProxy to contact GitLab [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332525 (https://phabricator.wikimedia.org/T421436) [00:58:41] (03CR) 10Tim Starling: "I used `mw-debug-repl` in production to confirm that it's possible to download zip files from GitLab this way." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332525 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [01:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [01:11:08] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1332526 [01:11:08] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1332526 (owner: 10TrainBranchBot) [01:18:49] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1332526 (owner: 10TrainBranchBot) [01:57:29] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops, and 2 others: db2207 down - https://phabricator.wikimedia.org/T435271#12269484 (10Jhancock.wm) 05In progress→03Resolved [01:58:10] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Degraded RAID on maps-test2001 - https://phabricator.wikimedia.org/T435311#12269487 (10Jhancock.wm) 05Open→03Declined [02:00:25] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:08:40] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 08m 15s) [02:10:50] (03PS1) 10Raymond Ndibe: toolhub: Bump container and enable Toolhub Evolved banner [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332528 (https://phabricator.wikimedia.org/T434556) [02:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:53:41] (03CR) 10Subramanya Sastry: [C:03+1] Produnto: use wgCopyUploadProxy to contact GitLab [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332525 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [03:09:26] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332525 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [03:10:17] (03Merged) 10jenkins-bot: Produnto: use wgCopyUploadProxy to contact GitLab [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332525 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [03:11:04] !log tstarling@deploy1003 Started scap sync-world: Backport for [[gerrit:1332525|Produnto: use wgCopyUploadProxy to contact GitLab (T421436)]] [03:11:08] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [03:29:26] !log tstarling@deploy1003 tstarling: Backport for [[gerrit:1332525|Produnto: use wgCopyUploadProxy to contact GitLab (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [03:29:29] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [03:31:11] !log tstarling@deploy1003 tstarling: Continuing with deployment [03:44:24] !log tstarling@deploy1003 Finished scap sync-world: Backport for [[gerrit:1332525|Produnto: use wgCopyUploadProxy to contact GitLab (T421436)]] (duration: 33m 20s) [03:44:27] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [04:00:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [04:26:14] (03PS1) 10Tim Starling: mediawiki: Add zip extension [puppet] - 10https://gerrit.wikimedia.org/r/1332532 (https://phabricator.wikimedia.org/T436488) [04:48:10] FIRING: BFDdown: BFD session down between cr2-eqdfw and fe80::b6f9:5dff:fe30:e538 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqdfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:53:10] FIRING: [3x] BFDdown: BFD session down between cr2-eqdfw and fe80::b6f9:5dff:fe30:e538 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:58:10] RESOLVED: [3x] BFDdown: BFD session down between cr2-eqdfw and fe80::b6f9:5dff:fe30:e538 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [06:02:22] (03PS1) 10Giuseppe Lavagetto: php: add php-zip extension [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1332538 (https://phabricator.wikimedia.org/T436488) [06:18:56] (03PS7) 10Jelto: sre.hosts.reboot-multiple: add new cookbook for unattended reboots [cookbooks] - 10https://gerrit.wikimedia.org/r/1305707 (https://phabricator.wikimedia.org/T395411) [06:29:11] (03CR) 10KartikMistry: [C:03+1] ArticleGuidance: Configure feedback links to local talk pages [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329314 (https://phabricator.wikimedia.org/T433483) (owner: 10Abijeet Patro) [06:37:52] (03CR) 10Arnaudb: [C:03+1] "thanks for the fix!" [cookbooks] - 10https://gerrit.wikimedia.org/r/1331653 (https://phabricator.wikimedia.org/T436361) (owner: 10Jelto) [07:00:05] Amir1, urbanecm, and awight: It is that lovely time of the day again! You are hereby commanded to deploy UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T0700). [07:00:05] Peterxy, hamishcz, revi, and abijeet: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:09] hohohoho [07:00:12] hohoho [07:00:33] hi revi [07:00:35] :) [07:00:41] I'm around to deploy abijeet's patch as well [07:00:41] hoi hoi :D [07:00:47] o/ [07:03:22] anyone deploying this slot? [07:04:11] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [07:04:15] (03CR) 10Slyngshede: [C:03+2] switchdc: remove parsoid [cookbooks] - 10https://gerrit.wikimedia.org/r/1331446 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [07:04:20] I can deploy if no one around. [07:04:46] let's see :P [07:05:04] urbanecm: around? [07:05:06] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [07:05:38] kart_: mobile only and ooo :D [07:05:45] oh OK [07:05:45] Feel free to deploy if you can! [07:05:50] hi and bye [07:05:50] Sure [07:05:59] revi: hey! Long time no see :) [07:06:03] hihi [07:06:06] Let me start with abijeet's change. [07:06:14] except I kept existing at the secret cabal [07:07:22] abijeet: I think we missed dependency: [07:07:22] Error for Change '1329314', project: 'operations/mediawiki-config', branch: 'master': [07:07:22] Change '1329314' has dependency '1327237' targeting the master branch [07:07:22] of MediaWiki code project 'mediawiki/extensions/ArticleGuidance', but [07:07:23] the dependency is not present in live train branch: wmf/1.47.0-wmf.17 [07:07:31] (03Merged) 10jenkins-bot: switchdc: remove parsoid [cookbooks] - 10https://gerrit.wikimedia.org/r/1331446 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [07:08:38] hmm, i think we will have to wait for the train to move through. [07:08:48] Lets revert for now. I'll reschedule for later. [07:08:56] yes. Let's move it to later. [07:09:02] Patch isn't merge. [07:09:55] ok, thanks [07:12:30] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1332516 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:12:36] revi: Should I deploy your patch? [07:12:44] sure, unless you have reason not to :D [07:13:21] then can i go next? [07:13:57] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kartik@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332474 (https://phabricator.wikimedia.org/T436449) (owner: 10Revi) [07:14:14] hamishcz: sure [07:14:21] note that I think I am unsure how or if I can test the variable… since the APCOND is not visible onwiki [07:14:38] revi: I'll ping for testing. [07:14:40] oh. [07:14:46] can try find a user [07:14:51] sure [07:14:54] (03Merged) 10jenkins-bot: kowiki: convert extendedconfirmed calculation to begin from first edit [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332474 (https://phabricator.wikimedia.org/T436449) (owner: 10Revi) [07:15:13] !log kartik@deploy1003 Started scap sync-world: Backport for [[gerrit:1332474|kowiki: convert extendedconfirmed calculation to begin from first edit (T436449)]] [07:15:16] T436449: Adjust extendedconfirmed calculation to first-edit on kowiki - https://phabricator.wikimedia.org/T436449 [07:19:38] !log kartik@deploy1003 kartik, revi: Backport for [[gerrit:1332474|kowiki: convert extendedconfirmed calculation to begin from first edit (T436449)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:20:12] revi: see if you can test it.. [07:20:31] let me try.....… [07:20:37] !log jelto@deploy1003 helmfile [staging] START helmfile.d/services/miscweb: apply [07:20:51] but I doubt unless I have account with 500 edits and 29 days since registration I think I can't really test it myself :P [07:21:06] neither me :) [07:21:23] I can in theory desysop myself and see if I become EC'ed [07:21:28] but I don't want to risk my sysop lol [07:21:33] :) [07:21:37] !log jelto@deploy1003 helmfile [staging] DONE helmfile.d/services/miscweb: apply [07:21:54] so I guess we shall proceed as is, as enwiki did the same change last week AFAIK [07:22:19] Sure. Let's do it. [07:22:27] !log kartik@deploy1003 kartik, revi: Continuing with deployment [07:24:29] (03CR) 10Dpogorzelski: [C:03+2] kserve: 0.20 upstream alignment [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327557 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [07:25:22] !og installing expat security updates [07:26:44] (03PS1) 10Brouberol: airflow: add permissions to manage configmaps/pvcs to all instances [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332545 (https://phabricator.wikimedia.org/T436259) [07:27:48] (03CR) 10Dpogorzelski: [C:03+2] kserve: add LLMInferenceService support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329506 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [07:28:15] (03PS1) 10Brouberol: global_config: register the config-master-wikimedia external service [puppet] - 10https://gerrit.wikimedia.org/r/1332546 (https://phabricator.wikimedia.org/T436259) [07:29:09] (03CR) 10CI reject: [V:04-1] airflow: add permissions to manage configmaps/pvcs to all instances [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332545 (https://phabricator.wikimedia.org/T436259) (owner: 10Brouberol) [07:30:13] !log kartik@deploy1003 Finished scap sync-world: Backport for [[gerrit:1332474|kowiki: convert extendedconfirmed calculation to begin from first edit (T436449)]] (duration: 15m 00s) [07:30:16] T436449: Adjust extendedconfirmed calculation to first-edit on kowiki - https://phabricator.wikimedia.org/T436449 [07:30:25] revi: done [07:30:37] thanks :D [07:30:43] hamishcz: let's deploy your change now. [07:30:55] can can :o [07:30:55] hamishcz: will you able to test it? [07:31:01] yes sure [07:31:05] cool [07:31:08] (03CR) 10Dpogorzelski: [C:03+2] kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [07:31:17] (03PS2) 10Brouberol: airflow: add permissions to manage configmaps/pvcs to all instances [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332545 (https://phabricator.wikimedia.org/T436259) [07:31:18] (03PS1) 10Brouberol: Allow airflow-dumps/test-k8s to egress to config-master.discovery.wmnet [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332547 (https://phabricator.wikimedia.org/T436259) [07:31:19] (03CR) 10CI reject: [V:04-1] kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [07:31:21] (03CR) 10Dpogorzelski: [V:03+2 C:03+2] kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [07:31:24] (03CR) 10Brouberol: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9350/co" [puppet] - 10https://gerrit.wikimedia.org/r/1332546 (https://phabricator.wikimedia.org/T436259) (owner: 10Brouberol) [07:31:33] (03CR) 10CI reject: [V:04-1] kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [07:31:42] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module rancid [puppet] - 10https://gerrit.wikimedia.org/r/1332501 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:31:44] !log installing openssl security updates [07:31:45] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:32:10] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in module rsync [puppet] - 10https://gerrit.wikimedia.org/r/1332502 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:32:17] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kartik@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332181 (https://phabricator.wikimedia.org/T436426) (owner: 10Hamish) [07:32:58] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace unscoped legacy facts in profile puppet_compiler [puppet] - 10https://gerrit.wikimedia.org/r/1332503 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [07:33:14] (03Merged) 10jenkins-bot: thwikibooks: update wordmark and tagline [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332181 (https://phabricator.wikimedia.org/T436426) (owner: 10Hamish) [07:33:25] !log kartik@deploy1003 Started scap sync-world: Backport for [[gerrit:1332181|thwikibooks: update wordmark and tagline (T436426)]] [07:33:28] T436426: Requesting logo change for th.wikibooks.org - https://phabricator.wikimedia.org/T436426 [07:37:20] !log kartik@deploy1003 hamishz, kartik: Backport for [[gerrit:1332181|thwikibooks: update wordmark and tagline (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:38:01] Deny the patch please. [07:38:05] hamishcz: you can test. [07:38:13] hamishcz: not working? [07:38:17] Seems that I forgot something. [07:38:24] Yeah it's not working [07:38:26] oh. OK. [07:38:31] !log kartik@deploy1003 hamishz, kartik: Rolling back deployment [07:41:06] !log kartik@deploy1003 Finished scap sync-world: Backport for [[gerrit:1332181|thwikibooks: update wordmark and tagline (T436426)]] (duration: 07m 41s) [07:41:09] T436426: Requesting logo change for th.wikibooks.org - https://phabricator.wikimedia.org/T436426 [07:41:56] (03PS1) 10KartikMistry: Revert "thwikibooks: update wordmark and tagline" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332555 [07:43:13] ah, I should have merge rolledback change first. [07:43:18] (03PS4) 10Dpogorzelski: kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) [07:43:31] I'm so sorry for the inconvenience :) [07:43:36] (03CR) 10KartikMistry: [C:03+2] Revert "thwikibooks: update wordmark and tagline" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332555 (owner: 10KartikMistry) [07:43:56] hamishcz: no issue. it happens! [07:45:00] (03Merged) 10jenkins-bot: Revert "thwikibooks: update wordmark and tagline" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332555 (owner: 10KartikMistry) [07:48:41] (03CR) 10Dpogorzelski: kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [07:48:51] (03CR) 10Dpogorzelski: [V:03+2 C:03+2] kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [07:50:03] I'm going to dig out why they are not working as it is really LGTM... [07:50:07] Thanks kart_! [07:56:07] 10ops-ulsfo, 06DC-Ops: ulsfo: OOB migration from copper to fiber - https://phabricator.wikimedia.org/T436499 (10ayounsi) 03NEW p:05Triage→03High [08:00:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [08:01:56] !log installing Linux 6.12.107 on Trixie hosts [08:01:57] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:03:23] hamishcz: sure! [08:04:33] (03CR) 10Trueg: [C:03+1] Update default Eventgate URL value in the WDQS chart. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331764 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [08:13:01] (03PS1) 10Jelto: scaffold: Fix helm release NOTES.txt text [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332556 (https://phabricator.wikimedia.org/T433589) [08:14:59] (03CR) 10Jelto: "Do you think this `NOTES.txt` makes more sense? Or should helfile be used instead of helm? If we agree on a proper NOTES text I can update" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332556 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [08:21:02] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1332502 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:21:45] (03CR) 10Arnaudb: "comments are inline" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [08:39:31] !log mvernon@cumin2003 START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe [08:39:55] !log mvernon@cumin2003 START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe [08:41:12] (03CR) 10Marostegui: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332520 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [08:42:49] !log elukey@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts pki1002.eqiad.wmnet [08:43:53] !log mvernon@cumin2003 END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe [08:44:30] (03CR) 10Marostegui: [C:03+1] "This can go anytime, but I've created a task to add the grants in production: https://phabricator.wikimedia.org/T436502" [puppet] - 10https://gerrit.wikimedia.org/r/1331666 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [08:45:13] (03CR) 10Marostegui: [C:03+1] "Jaime is OOO so let's wait for him to be back but I think this is fine." [puppet] - 10https://gerrit.wikimedia.org/r/1331494 (owner: 10Muehlenhoff) [08:45:56] (03CR) 10Marostegui: "I'd leave it commented, so we can re-use it if needed." [puppet] - 10https://gerrit.wikimedia.org/r/1329594 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [08:49:32] !log installing emacs security updates [08:49:33] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:51:00] (03CR) 10Arnaudb: "something else popped in the review:" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327572 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [08:51:43] (03PS1) 10Marostegui: mariadb: Productionize db1270 [puppet] - 10https://gerrit.wikimedia.org/r/1332670 (https://phabricator.wikimedia.org/T407942) [08:52:49] (03CR) 10Marostegui: [C:03+2] mariadb: Productionize db1270 [puppet] - 10https://gerrit.wikimedia.org/r/1332670 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [08:56:18] (03CR) 10Federico Ceratto: "We have the same content at line 213 (few lines above) for the new servers 'db[12]90[1-3]' so I would delete the chunk for db-test" [puppet] - 10https://gerrit.wikimedia.org/r/1329594 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [08:56:36] !log mvernon@cumin2003 END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe [08:57:07] !log Upgrading CI Jenkins 2.555.3 to 2.568.2 [08:57:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:57:15] elukey@cumin1003 upgrade-firmware (PID 1422054) is awaiting input [08:57:20] (03PS1) 10Marostegui: site.pp: Remove old hosts roles. [puppet] - 10https://gerrit.wikimedia.org/r/1332671 [08:58:15] !log Stop mariadb on sanitarium s2,s4,s6,s7 there will be lag on wikireplicas for those sections T407942 [08:58:17] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 22 hosts with reason: Cloning [08:58:18] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:58:18] T407942: Productionize db12[65-90] - https://phabricator.wikimedia.org/T407942 [09:01:10] (03CR) 10Marostegui: [C:03+2] site.pp: Remove old hosts roles. [puppet] - 10https://gerrit.wikimedia.org/r/1332671 (owner: 10Marostegui) [09:01:31] (03CR) 10Clément Goubert: [C:03+1] php: add php-zip extension [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1332538 (https://phabricator.wikimedia.org/T436488) (owner: 10Giuseppe Lavagetto) [09:06:35] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host pki1002.eqiad.wmnet [09:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [09:16:38] 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management: 400 Bad Request on File:Warsaw Pact in 1990 (orthographic projection).svg - https://phabricator.wikimedia.org/T436505 (10RhinosF1) 03NEW [09:16:41] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki1002.eqiad.wmnet [09:16:42] !log elukey@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts pki1002.eqiad.wmnet [09:16:47] 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management: 400 Bad Request on File:Warsaw Pact in 1990 (orthographic projection).svg - https://phabricator.wikimedia.org/T436505#12270163 (10RhinosF1) [09:17:54] (03CR) 10Trueg: [C:03+2] Update default Eventgate URL value in the WDQS chart. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331764 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [09:17:57] 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management: 400 Bad Request on File:Warsaw Pact in 1990 (orthographic projection).svg - https://phabricator.wikimedia.org/T436505#12270165 (10MatthewVernon) https://upload.wikimedia.org/wikipedia/commons/a/a8/Warsaw_Pact_in_1990_%28orthographic_projection%29.s... [09:18:25] FIRING: [24x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki1002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:20:24] (03Merged) 10jenkins-bot: Update default Eventgate URL value in the WDQS chart. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331764 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [09:21:21] !log elukey@cumin1003 START - Cookbook sre.hosts.provision for host pki1002.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [09:21:52] 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management: 400 Bad Request on File:Warsaw Pact in 1990 (orthographic projection).svg - https://phabricator.wikimedia.org/T436505#12270181 (10MatthewVernon) I'm not sure why https://en.wikipedia.org/wiki/File:Warsaw_Pact_in_1990_(orthographic_projection).svg i... [09:24:05] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:aux-master-eqiad [09:24:08] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet [09:24:09] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet [09:24:47] PROBLEM - Host pki1002 is DOWN: PING CRITICAL - Packet loss = 100% [09:26:39] this is me, already depooled --^ [09:28:25] RESOLVED: [24x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki1002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:29:10] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host pki1002.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [09:29:11] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet [09:29:12] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet [09:29:15] RECOVERY - Host pki1002 is UP: PING OK - Packet loss = 0%, RTA = 0.52 ms [09:29:17] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet [09:29:18] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet [09:29:54] FIRING: [44x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:29:56] FIRING: [44x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:30:15] !log elukey@puppetserver1001 conftool action : set/pooled=true; selector: dnsdisc=pki,name=eqiad [09:31:11] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12270197 (10elukey) @VRiley-WMF o/ since the host was already depooled I took the liberty to upgrade Idrac and Bios firmw... [09:33:40] FIRING: [24x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki1002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:33:44] (03PS9) 10Blake: memcached: add alerts for x509 certificate expiration. [alerts] - 10https://gerrit.wikimedia.org/r/1329591 (https://phabricator.wikimedia.org/T353511) [09:34:12] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet [09:34:13] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet [09:34:14] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:aux-master-eqiad [09:34:54] RESOLVED: [44x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:37:11] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:aux-master-codfw [09:37:15] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet [09:37:16] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet [09:42:03] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet [09:42:05] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet [09:42:11] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet [09:42:12] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet [09:43:19] (03CR) 10Clément Goubert: [C:03+1] memcached: add alerts for x509 certificate expiration. [alerts] - 10https://gerrit.wikimedia.org/r/1329591 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [09:44:50] 10SRE-swift-storage, 06Commons, 10MediaWiki-File-management: 400 Bad Request on File:Warsaw Pact in 1990 (orthographic projection).svg - https://phabricator.wikimedia.org/T436505#12270224 (10Samwilson) [09:46:14] (03CR) 10JMeybohm: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332510 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [09:46:22] (03CR) 10Blake: [C:03+2] memcached: add alerts for x509 certificate expiration. [alerts] - 10https://gerrit.wikimedia.org/r/1329591 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [09:46:39] (03PS1) 10Ozge: linkrecommendation: Update image to 2026-08-28-113209-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332676 (https://phabricator.wikimedia.org/T434259) [09:46:55] (03PS2) 10Ozge: linkrecommendation: Update image to 2026-08-28-113209-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332676 (https://phabricator.wikimedia.org/T434259) [09:47:18] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet [09:47:20] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet [09:47:20] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:aux-master-codfw [09:48:32] (03Merged) 10jenkins-bot: memcached: add alerts for x509 certificate expiration. [alerts] - 10https://gerrit.wikimedia.org/r/1329591 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [09:51:08] (03CR) 10Clément Goubert: [C:03+1] api-gateway: Basic cluster specifier support and Lua plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [09:55:04] (03CR) 10Ozge: [C:03+2] linkrecommendation: Update image to 2026-08-28-113209-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332676 (https://phabricator.wikimedia.org/T434259) (owner: 10Ozge) [09:57:47] (03Merged) 10jenkins-bot: linkrecommendation: Update image to 2026-08-28-113209-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332676 (https://phabricator.wikimedia.org/T434259) (owner: 10Ozge) [09:59:40] !log disabled puppet on A:cp to gradually apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1330314, starting from cp6016 [09:59:40] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1000) [10:00:47] (03CR) 10Fabfur: [C:03+2] cache::haproxy: fix logging of normalized Host header [puppet] - 10https://gerrit.wikimedia.org/r/1330314 (https://phabricator.wikimedia.org/T434766) (owner: 10Fabfur) [10:01:02] !log ozge@deploy1003 helmfile [staging] START helmfile.d/services/linkrecommendation: apply [10:02:09] !log ozge@deploy1003 helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply [10:04:23] 10SRE-Access-Requests, 06Infrastructure-Foundations, 10LDAP-Access-Requests, 13Patch-For-Review: Superset data access request for abibendall - https://phabricator.wikimedia.org/T430938#12270289 (10FCeratto-WMF) @ABendall-WMF I tried reaching out to you on Slack with no success. Can you please follow up htt... [10:04:33] (03PS2) 10JMeybohm: Puppet 8: Replace unscoped legacy facts in module k8s [puppet] - 10https://gerrit.wikimedia.org/r/1332510 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [10:04:36] (03CR) 10JMeybohm: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332510 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [10:05:00] (03PS1) 10Ladsgroup: varnish: Push webp on Safari except 14 and 15 [puppet] - 10https://gerrit.wikimedia.org/r/1332678 (https://phabricator.wikimedia.org/T431150) [10:07:03] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{aux-k8s-worker[1006-1009].eqiad.wmnet} and (A:aux-master-eqiad or A:aux-worker-eqiad) [10:07:06] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet [10:07:27] !log enable puppet on A:cp to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1330314 (T434766) [10:07:29] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:07:30] T434766: Normalize URI host in Turnilo's webrequest_sampled_live - https://phabricator.wikimedia.org/T434766 [10:07:42] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet [10:07:47] (03CR) 10Marostegui: [C:03+1] "Probably just a file that got started to get changed and not the other. In any case, remember that this file is for tracking and it doesn'" [puppet] - 10https://gerrit.wikimedia.org/r/1331411 (https://phabricator.wikimedia.org/T427884) (owner: 10Muehlenhoff) [10:10:16] !log ozge@deploy1003 helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply [10:12:24] !log ozge@deploy1003 helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply [10:12:50] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet [10:12:51] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet [10:12:57] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet [10:13:33] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet [10:14:07] !log ozge@deploy1003 helmfile [codfw] START helmfile.d/services/linkrecommendation: apply [10:15:11] (03CR) 10Giuseppe Lavagetto: [C:03+2] drivers: Move all docker client calls to driver [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [10:15:52] !log ozge@deploy1003 helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply [10:17:38] (03CR) 10Giuseppe Lavagetto: [C:03+2] drivers: Driver registration and factory [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1326398 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [10:18:41] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet [10:18:42] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet [10:18:48] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet [10:19:23] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet [10:20:44] (03CR) 10Clément Goubert: "From what I can tell, this could probably be a deployment of the (badly named) `python-webapp` chart, which is a generic minimal mesh + ap" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326332 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [10:20:54] (03Merged) 10jenkins-bot: drivers: Move all docker client calls to driver [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1325982 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [10:22:08] (03CR) 10JMeybohm: [C:03+1] "We should stick to helm only since these messages are also shown when installing charts in local clusters where helmfile is not used." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332556 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [10:23:01] (03Merged) 10jenkins-bot: drivers: Driver registration and factory [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1326398 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [10:24:40] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet [10:24:41] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet [10:24:47] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet [10:25:22] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet [10:29:55] RESOLVED: [24x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki1002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:30:34] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet [10:30:35] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet [10:30:35] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{aux-k8s-worker[1006-1009].eqiad.wmnet} and (A:aux-master-eqiad or A:aux-worker-eqiad) [10:31:37] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) (owner: 10BryanDavis) [10:35:09] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332532 (https://phabricator.wikimedia.org/T436488) (owner: 10Tim Starling) [10:35:23] (03CR) 10Marostegui: [C:03+1] "Ah ok! Thanks" [puppet] - 10https://gerrit.wikimedia.org/r/1329594 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [10:36:08] (03CR) 10Kamila Součková: [C:03+2] php: add php-zip extension [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1332538 (https://phabricator.wikimedia.org/T436488) (owner: 10Giuseppe Lavagetto) [10:36:16] (03CR) 10Kamila Součková: [V:03+2 C:03+2] php: add php-zip extension [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1332538 (https://phabricator.wikimedia.org/T436488) (owner: 10Giuseppe Lavagetto) [10:36:53] (03CR) 10JMeybohm: [C:03+2] Puppet 8: Replace unscoped legacy facts in module k8s [puppet] - 10https://gerrit.wikimedia.org/r/1332510 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [10:39:43] (03CR) 10FNegri: [C:03+1] "@rolisaemeka-ctr@wikimedia.org LGTM, do you have the permissions to merge and deploy? Otherwise I can do it. Instructions are at https://w" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332528 (https://phabricator.wikimedia.org/T434556) (owner: 10Raymond Ndibe) [10:46:21] (03CR) 10Marostegui: "Can we have a PCC with some hosts to make sure it is all good?" [puppet] - 10https://gerrit.wikimedia.org/r/1332520 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [10:47:53] 10ops-codfw, 06SRE, 10SRE-swift-storage, 10Ceph, and 2 others: Q1:rack/setup/install apus-be200[7-9] - https://phabricator.wikimedia.org/T436180#12270371 (10MatthewVernon) [these are a new type of h/w so will need some preseed work I suspect] [10:48:16] jouncebot: nowandnext [10:48:16] For the next 0 hour(s) and 11 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1000) [10:48:16] In 2 hour(s) and 11 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1300) [10:48:53] I'm going to rebuild MW images so the deploy window builds don't take forever [10:50:40] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 10Ceph, and 2 others: Q1:rack/setup/install apus-be100[7-9] - https://phabricator.wikimedia.org/T436181#12270374 (10MatthewVernon) [10:53:20] (03PS1) 10Hnowlan: wmnet: remove graphite, statsd [dns] - 10https://gerrit.wikimedia.org/r/1332687 (https://phabricator.wikimedia.org/T435340) [10:53:44] (03CR) 10Lucas Werkmeister (WMDE): scaffold: Fix helm release NOTES.txt text (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332556 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [10:54:05] !log kamila@deploy1003 Started scap sync-world: rebuild for T436488 [10:54:08] T436488: Add PHP zip extension to MediaWiki images - https://phabricator.wikimedia.org/T436488 [10:54:44] (03PS1) 10Ozge: editing-suggestions: Update model to exclude NPOV suggestions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332690 (https://phabricator.wikimedia.org/T435358) [10:57:47] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [10:57:47] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [10:57:47] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [11:01:34] (03CR) 10Marostegui: [C:03+1] "Looks good: https://puppet-compiler.wmflabs.org/output/1329665/9351/" [puppet] - 10https://gerrit.wikimedia.org/r/1329665 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [11:01:35] (03CR) 10Federico Ceratto: [C:03+2] preseed.yaml: Remove db-test entry [puppet] - 10https://gerrit.wikimedia.org/r/1329594 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [11:05:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:06:31] (03PS2) 10Kamila Součková: mediawiki: Add zip extension [puppet] - 10https://gerrit.wikimedia.org/r/1332532 (https://phabricator.wikimedia.org/T436488) (owner: 10Tim Starling) [11:06:35] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332532 (https://phabricator.wikimedia.org/T436488) (owner: 10Tim Starling) [11:11:55] !log powercycle netmon2002, unresponsive [11:11:56] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:12:58] (03PS6) 10Kamila Součková: P:mediawiki::php: add 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) (owner: 10BryanDavis) [11:13:02] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) (owner: 10BryanDavis) [11:13:41] (03CR) 10Ozge: [C:03+2] editing-suggestions: Update model to exclude NPOV suggestions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332690 (https://phabricator.wikimedia.org/T435358) (owner: 10Ozge) [11:14:37] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [11:14:37] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [11:14:37] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.132 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [11:15:18] (03CR) 10Marostegui: [C:03+1] "https://puppet-compiler.wmflabs.org/output/1332520/9352/ looks good." [puppet] - 10https://gerrit.wikimedia.org/r/1332520 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [11:16:14] (03Merged) 10jenkins-bot: editing-suggestions: Update model to exclude NPOV suggestions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332690 (https://phabricator.wikimedia.org/T435358) (owner: 10Ozge) [11:17:44] FIRING: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:18:42] (03CR) 10Kamila Součková: [C:03+2] mediawiki: Add zip extension [puppet] - 10https://gerrit.wikimedia.org/r/1332532 (https://phabricator.wikimedia.org/T436488) (owner: 10Tim Starling) [11:20:51] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . [11:21:07] (03CR) 10Filippo Giunchedi: [C:04-2] "Please implement the following:" [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) (owner: 10Andrew Bogott) [11:21:08] !log ozge@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [11:24:53] (03CR) 10Kamila Součková: [C:03+2] P:mediawiki::php: add 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) (owner: 10BryanDavis) [11:25:07] (03CR) 10Kamila Součková: [C:03+2] "LGTM, thank you!" [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) (owner: 10BryanDavis) [11:25:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [11:28:34] !log kamila@deploy1003 Finished scap sync-world: rebuild for T436488 (duration: 34m 43s) [11:29:01] T436488: Add PHP zip extension to MediaWiki images - https://phabricator.wikimedia.org/T436488 [11:29:45] (03CR) 10Kamila Součková: [C:03+2] "Acknowledged" [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) (owner: 10BryanDavis) [11:31:10] (03CR) 10Muehlenhoff: "Sure, this is just a cleanup, so I'll wait until he's back." [puppet] - 10https://gerrit.wikimedia.org/r/1331494 (owner: 10Muehlenhoff) [11:32:44] RESOLVED: RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:33:43] (03PS2) 10Blake: memcached: continue rollout of PKI use. [puppet] - 10https://gerrit.wikimedia.org/r/1330436 (https://phabricator.wikimedia.org/T353511) [11:35:57] (03PS1) 10JMeybohm: k8s: Disable user namespaces support [puppet] - 10https://gerrit.wikimedia.org/r/1332695 (https://phabricator.wikimedia.org/T427069) [11:37:04] (03CR) 10Clément Goubert: [C:03+1] memcached: continue rollout of PKI use. [puppet] - 10https://gerrit.wikimedia.org/r/1330436 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [11:37:58] (03CR) 10Blake: [C:03+2] memcached: continue rollout of PKI use. [puppet] - 10https://gerrit.wikimedia.org/r/1330436 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [11:39:36] (03CR) 10JMeybohm: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332695 (https://phabricator.wikimedia.org/T427069) (owner: 10JMeybohm) [11:44:29] (03PS2) 10Slyngshede: wmnet: update CNAME records for DB masters to codfw [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) [11:45:13] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12270495 (10Marostegui) I think it is fine to proceed @VRiley-WMF - thanks! [11:45:42] (03CR) 10Slyngshede: "Let me know if I'm missing some hosts. There are the dbproxy entries which I'm unsure of." [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [11:45:53] PROBLEM - Host wikikube-worker1261 is DOWN: PING CRITICAL - Packet loss = 0%, RTA = 4121.61 ms [11:45:57] (03CR) 10Slyngshede: "Should we add X4?" [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [11:45:59] RECOVERY - Host wikikube-worker1261 is UP: PING OK - Packet loss = 0%, RTA = 0.31 ms [11:49:01] (03PS1) 10Hnowlan: Remove various hardcoded statsd.eqiad references [puppet] - 10https://gerrit.wikimedia.org/r/1332698 (https://phabricator.wikimedia.org/T435340) [11:49:03] (03PS1) 10Hnowlan: scap: remove statsd config [puppet] - 10https://gerrit.wikimedia.org/r/1332699 (https://phabricator.wikimedia.org/T435340) [11:49:37] (03CR) 10Hnowlan: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332698 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [11:49:40] (03CR) 10CI reject: [V:04-1] Remove various hardcoded statsd.eqiad references [puppet] - 10https://gerrit.wikimedia.org/r/1332698 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [11:51:20] (03CR) 10AOkoth: [C:03+1] etherpad-next: add discovery ingress records [dns] - 10https://gerrit.wikimedia.org/r/1331613 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [11:52:32] (03CR) 10Gergő Tisza: [C:03+1] CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330446 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [11:53:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv4 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133212 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:56:41] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12270552 (10FCeratto-WMF) [11:58:37] (03CR) 10Hnowlan: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1332698 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [11:58:43] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv4 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133212 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:00:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [12:01:35] PROBLEM - Thanos swift https on thanos-fe1007 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/Thanos [12:04:25] RECOVERY - Thanos swift https on thanos-fe1007 is OK: HTTP OK: HTTP/1.1 200 OK - 279 bytes in 0.055 second response time https://wikitech.wikimedia.org/wiki/Thanos [12:08:24] (03PS1) 10Mszwarc: Update stream config for user_info_card_interaction [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332684 (https://phabricator.wikimedia.org/T435585) [12:09:38] (03PS2) 10Joal: Update X-is-browser webrequest_sample turnilo conf [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331524 (https://phabricator.wikimedia.org/T435152) [12:20:05] (03CR) 10Marostegui: [C:04-1] "There are some errors. I've pointed out the right host." [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [12:26:08] (03PS2) 10Jelto: scaffold: Fix helm release NOTES.txt text [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332556 (https://phabricator.wikimedia.org/T433589) [12:29:47] (03CR) 10JMeybohm: scaffold: Fix helm release NOTES.txt text (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332556 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [12:29:56] !log installing openjdk-17 security updates [12:29:57] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:33:45] FIRING: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [12:38:56] (03CR) 10Arnaudb: "thanks for putting this together! Inline comments to highlight the findings returned by Claude that I think are worth noting" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327571 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [12:39:39] FIRING: CoreBGPDown: Core BGP session down between cr1-codfw and cr1-eqiad (208.80.153.220) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:39:58] FIRING: CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:41:31] (03PS1) 10Aude: Enable ReadingLists for logged-in users on phase 1 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332717 (https://phabricator.wikimedia.org/T434922) [12:42:28] (03CR) 10Aude: [C:04-2] "this is scheduled for tomorrow, September 1" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332717 (https://phabricator.wikimedia.org/T434922) (owner: 10Aude) [12:43:55] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:44:39] (03CR) 10Brouberol: [C:03+1] Update X-is-browser webrequest_sample turnilo conf [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331524 (https://phabricator.wikimedia.org/T435152) (owner: 10Joal) [12:45:45] (03CR) 10Brouberol: [C:03+2] Update X-is-browser webrequest_sample turnilo conf [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331524 (https://phabricator.wikimedia.org/T435152) (owner: 10Joal) [12:50:25] (03PS1) 10JMeybohm: kubernetes: Add a comment about kubelet cgroup driver selection [puppet] - 10https://gerrit.wikimedia.org/r/1332719 (https://phabricator.wikimedia.org/T427069) [12:52:18] 06SRE, 06Infrastructure-Foundations: Re-IP hosts running Cassandra to per-rack subnets in codfw row A and B. - https://phabricator.wikimedia.org/T354871#12270684 (10ayounsi) > Too many IPs configure on the primary interface, manual migration required. Ah right, of course... That's because those hosts have ex... [12:55:32] (03PS2) 10JMeybohm: k8s: Disable user namespaces support [puppet] - 10https://gerrit.wikimedia.org/r/1332695 (https://phabricator.wikimedia.org/T427069) [12:55:32] (03PS2) 10JMeybohm: kubernetes: Add a comment about kubelet cgroup driver selection [puppet] - 10https://gerrit.wikimedia.org/r/1332719 (https://phabricator.wikimedia.org/T427069) [12:55:34] (03CR) 10Jelto: k8s: Disable user namespaces support (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1332695 (https://phabricator.wikimedia.org/T427069) (owner: 10JMeybohm) [12:56:10] (03CR) 10Jelto: [C:03+1] "lgtm now" [puppet] - 10https://gerrit.wikimedia.org/r/1332695 (https://phabricator.wikimedia.org/T427069) (owner: 10JMeybohm) [12:56:23] (03CR) 10JMeybohm: k8s: Disable user namespaces support (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1332695 (https://phabricator.wikimedia.org/T427069) (owner: 10JMeybohm) [12:57:16] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12270708 (10JArguello-WMF) Hi @FCeratto-WMF , I approve this request, thank you so much for your help. [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: I seem to be stuck in Groundhog week. Sigh. Time for (yet another) UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1300). [13:00:05] No Gerrit patches in the queue for this window AFAICS. [13:00:16] o/ [13:00:26] I don’t see anything to deploy either :) [13:01:30] (03PS2) 10Arnaudb: kubernetes: allow pod egress to text-lb for gitlab [puppet] - 10https://gerrit.wikimedia.org/r/1332720 (https://phabricator.wikimedia.org/T425441) [13:01:44] (03PS3) 10Slyngshede: wmnet: update CNAME records for DB masters to codfw [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) [13:02:21] (03CR) 10Scott French: "Thanks for the reviews, Reuven!" [puppet] - 10https://gerrit.wikimedia.org/r/1330689 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [13:02:26] (03CR) 10Scott French: [C:03+2] P:services_proxy::envoy: Fix sni_rewrites_host_header handling [puppet] - 10https://gerrit.wikimedia.org/r/1330689 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [13:02:32] (03CR) 10Slyngshede: "Not sure how I got that many wrong." [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [13:03:45] RESOLVED: WidespreadPuppetFailure: Puppet has failed in eqsin - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [13:05:10] FIRING: BFDdown: BFD session down between cr1-codfw and fe80::7ee2:ca07:dde:4ba1 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:05:32] (03CR) 10Scott French: [C:03+2] P:services_proxy::envoy: Drop support for split and introduce splits [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [13:06:17] (03PS1) 10Ayounsi: Remove POPs GRE tunnels [homer/public] - 10https://gerrit.wikimedia.org/r/1332723 [13:06:37] (03PS3) 10Arnaudb: kubernetes: allow pod egress to text-lb for gitlab [puppet] - 10https://gerrit.wikimedia.org/r/1332720 (https://phabricator.wikimedia.org/T425441) [13:06:39] (03PS8) 10Scott French: P:services_proxy::envoy: Drop support for split and introduce splits [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) [13:06:39] (03PS12) 10Scott French: P:kubernetes::deployment_server::global_config: Update services_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1328256 (https://phabricator.wikimedia.org/T427666) [13:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [13:07:07] (03PS2) 10Ayounsi: Remove POPs GRE tunnels OSPF [homer/public] - 10https://gerrit.wikimedia.org/r/1332723 [13:10:10] RESOLVED: BFDdown: BFD session down between cr1-codfw and fe80::7ee2:ca07:dde:4ba1 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:10:42] (03CR) 10Scott French: [C:03+2] P:services_proxy::envoy: Drop support for split and introduce splits [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [13:11:22] (03PS3) 10Blake: memcached: Rebuild configuration before reloading certificates. [puppet] - 10https://gerrit.wikimedia.org/r/1332722 (https://phabricator.wikimedia.org/T353511) [13:14:42] (03CR) 10Scott French: [C:03+2] P:kubernetes::deployment_server::global_config: Update services_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1328256 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [13:15:34] 06SRE: Add alerting for gnmic total series counters - https://phabricator.wikimedia.org/T435184#12270772 (10ayounsi) I opened {T436517} to maybe fix the issue by throwing more compute at it. [13:18:10] !log joal@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply [13:18:36] !log joal@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply [13:19:16] (03CR) 10Muehlenhoff: [C:03+2] service::node: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328178 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [13:20:34] (03PS1) 10Slyngshede: mw-web: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332725 (https://phabricator.wikimedia.org/T433363) [13:20:44] (03PS7) 10Tiziano Fogli: prometheus: add cardinality exporter [puppet] - 10https://gerrit.wikimedia.org/r/1331600 (https://phabricator.wikimedia.org/T435334) [13:20:55] (03PS4) 10Tiziano Fogli: prometheus: add cardinality exporter (ops instance) [puppet] - 10https://gerrit.wikimedia.org/r/1332549 (https://phabricator.wikimedia.org/T435334) [13:21:23] !log installing apr-util security updates [13:21:24] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:21:28] (03CR) 10Slyngshede: "We might be able to go lower than 600, but that's the number we used in March." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332725 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [13:21:31] (03PS4) 10Tiziano Fogli: prometheus: add cardinality exporter (ext instance) [puppet] - 10https://gerrit.wikimedia.org/r/1332550 (https://phabricator.wikimedia.org/T435334) [13:21:34] (03PS4) 10Tiziano Fogli: prometheus: add cardinality exporter (analytics instance) [puppet] - 10https://gerrit.wikimedia.org/r/1332551 (https://phabricator.wikimedia.org/T435334) [13:21:36] (03PS4) 10Tiziano Fogli: prometheus: add cardinality exporter (cloud instance) [puppet] - 10https://gerrit.wikimedia.org/r/1332552 (https://phabricator.wikimedia.org/T435334) [13:21:38] (03PS4) 10Tiziano Fogli: prometheus: add cardinality exporter (services instance) [puppet] - 10https://gerrit.wikimedia.org/r/1332553 (https://phabricator.wikimedia.org/T435334) [13:21:40] (03PS4) 10Tiziano Fogli: prometheus: add cardinality exporter (k8s instances) [puppet] - 10https://gerrit.wikimedia.org/r/1332554 (https://phabricator.wikimedia.org/T435334) [13:25:04] (03CR) 10Tiziano Fogli: [C:03+1] graphite: disable uwsgi service for graphite web [puppet] - 10https://gerrit.wikimedia.org/r/1328153 (https://phabricator.wikimedia.org/T435341) (owner: 10Hnowlan) [13:25:12] (03CR) 10Scott French: "Diff vs. current: https://phabricator.wikimedia.org/P96280" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331857 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [13:25:48] (03PS3) 10Scott French: Rakefile: Simplify and update 'upstream' mock data population [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331872 (https://phabricator.wikimedia.org/T427666) [13:25:49] (03CR) 10Scott French: "Diff vs. Ifc82fee7433284eecd7713c539796b756a6a6964: https://phabricator.wikimedia.org/P96281" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331872 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [13:26:16] (03PS3) 10Scott French: Rakefile: Update mock services_proxy data for splits [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) [13:26:16] (03CR) 10Scott French: "Diff vs. Ic0e1aff9984fc3d91d431b4034115d506a6a6964 is empty (expected)." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [13:37:27] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12270868 (10FCeratto-WMF) [13:41:19] (03PS2) 10Ladsgroup: upload: Drop profile::cache::upload::upload_webp_hits_threshold [puppet] - 10https://gerrit.wikimedia.org/r/1327656 (https://phabricator.wikimedia.org/T431150) [13:41:25] (03CR) 10Ladsgroup: [V:03+2 C:03+2] upload: Drop profile::cache::upload::upload_webp_hits_threshold [puppet] - 10https://gerrit.wikimedia.org/r/1327656 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [13:42:45] (03PS1) 10AOkoth: phabricator: decom phab1004 and add phab1006 [puppet] - 10https://gerrit.wikimedia.org/r/1332730 (https://phabricator.wikimedia.org/T377889) [13:44:53] (03CR) 10Muehlenhoff: [C:03+2] Apply cluster::management role to cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1331611 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [13:48:28] (03CR) 10Filippo Giunchedi: sre/cdn: use traffic ratios in ATSBackendErrorsHigh (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [13:49:25] (03CR) 10Andrew Bogott: "> /home/andrew isn't going to work for the destination file" [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) (owner: 10Andrew Bogott) [13:49:47] (03CR) 10AOkoth: "https://puppet-compiler.wmflabs.org/output/1332730/9353/" [puppet] - 10https://gerrit.wikimedia.org/r/1332730 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [13:50:12] (03CR) 10Muehlenhoff: [C:03+2] Add cumin1004 as mysql root client / grant [puppet] - 10https://gerrit.wikimedia.org/r/1331666 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [13:51:09] (03CR) 10Filippo Giunchedi: [C:04-2] "Some extra inspiration, the nfsd textfile collector I am working on: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1331487" [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) (owner: 10Andrew Bogott) [13:52:15] (03PS1) 10Arnaudb: Revert^2 "gitlab: discard firewall throttling on the primary" [puppet] - 10https://gerrit.wikimedia.org/r/1332731 (https://phabricator.wikimedia.org/T425441) [13:54:09] (03PS1) 10Arnaudb: Revert^2 "gitlab: point gitlab.wikimedia.org at the CDN" [dns] - 10https://gerrit.wikimedia.org/r/1332733 (https://phabricator.wikimedia.org/T425441) [13:55:48] (03PS2) 10Arnaudb: Revert^2 "gitlab: discard firewall throttling on the primary" [puppet] - 10https://gerrit.wikimedia.org/r/1332731 (https://phabricator.wikimedia.org/T425441) [13:56:25] 10ops-eqiad, 06SRE, 06DC-Ops, 10Kafka-Infrastructure, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12270920 (10brouberol) @VRiley-WMF Would you be able to give us a ballpark date for when the first host could be... [13:59:18] 10ops-eqiad, 10Cloud-VPS, 06DC-Ops, 06tools-infrastructure-team: cloudcephosd1035 offline - https://phabricator.wikimedia.org/T436520 (10Andrew) 03NEW [14:03:24] (03PS1) 10Marostegui: redact_sanitarium.sh: Add db1270 [puppet] - 10https://gerrit.wikimedia.org/r/1332738 (https://phabricator.wikimedia.org/T407942) [14:04:29] 06SRE: Reactivate Sheila Wangari's Phabricator account - https://phabricator.wikimedia.org/T436521 (10JLam-WMF) 03NEW [14:05:22] (03CR) 10Marostegui: [C:03+1] "Looks good - we will have to review a few days before the DC switch as it is likely we'll be switching some masters in between." [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [14:05:51] (03CR) 10Marostegui: [C:03+2] redact_sanitarium.sh: Add db1270 [puppet] - 10https://gerrit.wikimedia.org/r/1332738 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [14:10:54] (03CR) 10Giuseppe Lavagetto: [C:03+2] drivers: Remove label attribute in favor of arguments [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329691 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [14:13:12] !log attempt xfs_repair /dev/sdd1 on ms-be1090 [14:13:12] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:14:35] 06SRE, 10Phabricator: Reactivate Sheila Wangari's Phabricator account - https://phabricator.wikimedia.org/T436521#12271149 (10JLam-WMF) [14:14:36] (03Merged) 10jenkins-bot: drivers: Remove label attribute in favor of arguments [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329691 (https://phabricator.wikimedia.org/T434957) (owner: 10Dduvall) [14:17:07] 10ops-eqiad, 10Cloud-VPS, 06DC-Ops, 06tools-infrastructure-team: cloudcephosd1035 offline - https://phabricator.wikimedia.org/T436520#12271173 (10VRiley-WMF) a:03VRiley-WMF [14:21:34] (03PS2) 10Eevans: sessionstore: set storage compatability to `UPGRADING` [puppet] - 10https://gerrit.wikimedia.org/r/1326905 (https://phabricator.wikimedia.org/T435154) [14:21:34] (03PS2) 10Eevans: sessionstore: set storage compatability to `NONE` [puppet] - 10https://gerrit.wikimedia.org/r/1326906 (https://phabricator.wikimedia.org/T435154) [14:21:42] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12271191 (10bking) Per my earlier update, I'm not seeing any evidence of disk issues (no alerts or log messages on either host) since I removed `g... [14:21:50] (03CR) 10Eevans: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1326905 (https://phabricator.wikimedia.org/T435154) (owner: 10Eevans) [14:21:58] (03CR) 10Eevans: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1326906 (https://phabricator.wikimedia.org/T435154) (owner: 10Eevans) [14:22:26] (03PS1) 10Jelto: service::catalog: Set ipip_encapsulation for citoid, cxserver in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1332741 (https://phabricator.wikimedia.org/T420436) [14:23:06] (03CR) 10CI reject: [V:04-1] service::catalog: Set ipip_encapsulation for citoid, cxserver in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1332741 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [14:24:17] (03PS2) 10Jelto: service::catalog: Set ipip_encapsulation for citoid, cxserver in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1332741 (https://phabricator.wikimedia.org/T420436) [14:25:50] (03PS1) 10Bking: Kerberos: Move replica to production [puppet] - 10https://gerrit.wikimedia.org/r/1332743 (https://phabricator.wikimedia.org/T435873) [14:27:56] 06SRE, 10Phabricator: Reactivate Sheila Wangari's Phabricator account - https://phabricator.wikimedia.org/T436521#12271232 (10taavi) 05Open→03Resolved a:03taavi Done. [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1430) [14:31:29] (03CR) 10Arnaudb: "thanks! the change looks good, 2 questions inline, and another one here:" [puppet] - 10https://gerrit.wikimedia.org/r/1332730 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [14:32:09] (03CR) 10Jelto: [C:03+1] "lgtm" [dns] - 10https://gerrit.wikimedia.org/r/1332733 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [14:32:18] (03CR) 10LSobanski: "Approved in the Infrastructure Foundations team meeting." [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [14:33:45] (03CR) 10Eevans: [C:03+2] sessionstore: set storage compatability to `UPGRADING` [puppet] - 10https://gerrit.wikimedia.org/r/1326905 (https://phabricator.wikimedia.org/T435154) (owner: 10Eevans) [14:36:53] !log eevans@cumin1003 START - Cookbook sre.cassandra.roll-restart for nodes matching A:sessionstore: Set storage compatability to UPGRADING — T435154 - eevans@cumin1003 [14:36:56] T435154: Upgrade sessionstore cluster to Cassandra 5.0.8 & JDK 17 - https://phabricator.wikimedia.org/T435154 [14:39:26] (03CR) 10Arnaudb: [C:03+1] Revert^2 "gitlab: discard firewall throttling on the primary" [puppet] - 10https://gerrit.wikimedia.org/r/1332731 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [14:41:41] (03PS1) 10Marostegui: mariadb.yaml: Declare pcX hosts as bidir [puppet] - 10https://gerrit.wikimedia.org/r/1332745 [14:42:12] (03CR) 10Marostegui: "@Ladsgroup@gmail.com do you see anything wrong with this patch?" [puppet] - 10https://gerrit.wikimedia.org/r/1332745 (owner: 10Marostegui) [14:42:33] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations: attempts to reimage cp5022 are failing - https://phabricator.wikimedia.org/T436386#12271303 (10LSobanski) @BCornwall this could have been caused by network maintenance in eqsin, could you try again and let us know if it's still a problem? [14:43:11] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations: attempts to reimage cp5022 are failing - https://phabricator.wikimedia.org/T436386#12271306 (10LSobanski) p:05Triage→03Medium [14:45:03] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=cp5022.* [reason: update IP addrs] [14:45:35] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie [14:45:48] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12271335 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [14:47:43] (03CR) 10AOkoth: "Yes." [puppet] - 10https://gerrit.wikimedia.org/r/1332730 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [14:48:13] !log eevans@cumin1003 END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:sessionstore: Set storage compatability to UPGRADING — T435154 - eevans@cumin1003 [14:48:16] T435154: Upgrade sessionstore cluster to Cassandra 5.0.8 & JDK 17 - https://phabricator.wikimedia.org/T435154 [14:48:27] (03PS2) 10AOkoth: phabricator: decom phab1004 and add phab1006 [puppet] - 10https://gerrit.wikimedia.org/r/1332730 (https://phabricator.wikimedia.org/T377889) [14:49:09] (03CR) 10Marostegui: "This looks as expected: https://puppet-compiler.wmflabs.org/output/1332745/9354/" [puppet] - 10https://gerrit.wikimedia.org/r/1332745 (owner: 10Marostegui) [14:50:08] (03PS1) 10Muehlenhoff: Cumin dbbackups config for cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1332746 (https://phabricator.wikimedia.org/T427897) [14:50:12] (03CR) 10Hnowlan: [C:03+1] sre/cdn: use traffic ratios in ATSBackendErrorsHigh (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [14:50:24] (03CR) 10Hnowlan: [C:03+1] sre/cdn: use traffic ratios in ATSBackendErrorsHigh (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [14:51:21] (03CR) 10Marostegui: [C:03+1] "These hosts will need the grant for cumin1004 which will arrive via https://phabricator.wikimedia.org/T436502" [puppet] - 10https://gerrit.wikimedia.org/r/1332746 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [14:52:17] (03CR) 10Arnaudb: [C:03+1] "thanks for the fix!" [puppet] - 10https://gerrit.wikimedia.org/r/1332730 (https://phabricator.wikimedia.org/T377889) (owner: 10AOkoth) [14:53:24] (03CR) 10Clément Goubert: [C:03+1] memcached: Rebuild configuration before reloading certificates. [puppet] - 10https://gerrit.wikimedia.org/r/1332722 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:54:15] (03PS1) 10Federico Ceratto: site.pp,core_test.yaml: shorten log retention test-s4 [puppet] - 10https://gerrit.wikimedia.org/r/1332747 (https://phabricator.wikimedia.org/T435059) [15:01:45] (03CR) 10Blake: [C:03+2] memcached: Rebuild configuration before reloading certificates. [puppet] - 10https://gerrit.wikimedia.org/r/1332722 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [15:02:11] jouncebot: now [15:02:11] No deployments scheduled for the next 0 hour(s) and 27 minute(s) [15:02:26] I’ll deploy a harmless little query builder update [15:02:38] (03CR) 10Lucas Werkmeister (WMDE): [C:03+2] wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331503 (https://phabricator.wikimedia.org/T435632) (owner: 10Lucas Werkmeister (WMDE)) [15:03:10] (03PS2) 10Lucas Werkmeister (WMDE): wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331503 (https://phabricator.wikimedia.org/T435632) [15:03:15] (03CR) 10Lucas Werkmeister (WMDE): wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331503 (https://phabricator.wikimedia.org/T435632) (owner: 10Lucas Werkmeister (WMDE)) [15:05:58] (03Merged) 10jenkins-bot: wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331503 (https://phabricator.wikimedia.org/T435632) (owner: 10Lucas Werkmeister (WMDE)) [15:06:30] (03PS1) 10Ladsgroup: Enable thumb.wikimedia.org everywhere except enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332748 (https://phabricator.wikimedia.org/T427465) [15:07:44] (03PS1) 10Filippo Giunchedi: openstack: stop using icinga for pdns auth checks [puppet] - 10https://gerrit.wikimedia.org/r/1332750 (https://phabricator.wikimedia.org/T328502) [15:07:47] (03PS1) 10Filippo Giunchedi: openldap: stop using icinga in clouddev [puppet] - 10https://gerrit.wikimedia.org/r/1332751 (https://phabricator.wikimedia.org/T328502) [15:08:54] !log lucaswerkmeister-wmde@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [15:09:13] !log lucaswerkmeister-wmde@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [15:09:21] !log lucaswerkmeister-wmde@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [15:09:36] !log lucaswerkmeister-wmde@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [15:09:41] !log lucaswerkmeister-wmde@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [15:09:59] (03PS1) 10Muehlenhoff: Remove dbbackups for cumin1004 for the initial setup [puppet] - 10https://gerrit.wikimedia.org/r/1332752 (https://phabricator.wikimedia.org/T427897) [15:10:02] !log lucaswerkmeister-wmde@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [15:10:38] (03CR) 10Ladsgroup: [C:03+1] "LGTM, I think we originally wanted to cut the replication but decided against it." [puppet] - 10https://gerrit.wikimedia.org/r/1332745 (owner: 10Marostegui) [15:10:43] (03CR) 10Marostegui: [C:03+1] site.pp,core_test.yaml: shorten log retention test-s4 [puppet] - 10https://gerrit.wikimedia.org/r/1332747 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [15:10:56] (03CR) 10Marostegui: [C:03+2] mariadb.yaml: Declare pcX hosts as bidir [puppet] - 10https://gerrit.wikimedia.org/r/1332745 (owner: 10Marostegui) [15:11:51] (03CR) 10Hnowlan: [C:03+2] graphite: disable uwsgi service for graphite web [puppet] - 10https://gerrit.wikimedia.org/r/1328153 (https://phabricator.wikimedia.org/T435341) (owner: 10Hnowlan) [15:12:49] (03CR) 10Eevans: [C:03+2] sessionstore: set storage compatability to `NONE` [puppet] - 10https://gerrit.wikimedia.org/r/1326906 (https://phabricator.wikimedia.org/T435154) (owner: 10Eevans) [15:13:19] (03CR) 10Marostegui: "Oh sweet. I am CC'ing Jaime here so he's aware when he's back." [puppet] - 10https://gerrit.wikimedia.org/r/1332752 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [15:13:41] (03CR) 10Marostegui: [C:03+1] Remove dbbackups for cumin1004 for the initial setup [puppet] - 10https://gerrit.wikimedia.org/r/1332752 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [15:14:36] (03PS1) 10Filippo Giunchedi: site: temp decom for cloudvirt10(6[567]|7[234]) ahead of rack move [puppet] - 10https://gerrit.wikimedia.org/r/1332754 (https://phabricator.wikimedia.org/T435921) [15:15:03] (03CR) 10Ladsgroup: "Yup, I want to reduce it slowly to avoid major sudden shift of traffic" [puppet] - 10https://gerrit.wikimedia.org/r/1327657 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [15:16:47] (03PS2) 10Ladsgroup: cache: Reduce the webp threshold to 80 [puppet] - 10https://gerrit.wikimedia.org/r/1327657 (https://phabricator.wikimedia.org/T431150) [15:16:54] (03CR) 10Ladsgroup: [V:03+2 C:03+2] cache: Reduce the webp threshold to 80 [puppet] - 10https://gerrit.wikimedia.org/r/1327657 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [15:17:04] jouncebot: nowandnext [15:17:04] No deployments scheduled for the next 0 hour(s) and 13 minute(s) [15:17:04] In 0 hour(s) and 13 minute(s): Wikimedia Portals Update (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1530) [15:17:08] PROBLEM - graphite.wikimedia.org api on graphite1005 is CRITICAL: HTTP CRITICAL: HTTP/1.1 502 Bad Gateway - 438 bytes in 0.001 second response time https://wikitech.wikimedia.org/wiki/Graphite%23Operations_troubleshooting [15:17:08] PROBLEM - graphite.wikimedia.org render on graphite1005 is CRITICAL: HTTP CRITICAL: HTTP/1.1 502 Bad Gateway - 438 bytes in 0.001 second response time https://wikitech.wikimedia.org/wiki/Graphite%23Operations_troubleshooting [15:17:08] (03PS1) 10Filippo Giunchedi: site: temp decom for cloudvirt1075 ahead of rack move [puppet] - 10https://gerrit.wikimedia.org/r/1332755 (https://phabricator.wikimedia.org/T435923) [15:17:10] PROBLEM - graphite.wikimedia.org render on graphite2004 is CRITICAL: HTTP CRITICAL: HTTP/1.1 502 Bad Gateway - 438 bytes in 0.082 second response time https://wikitech.wikimedia.org/wiki/Graphite%23Operations_troubleshooting [15:17:10] PROBLEM - graphite.wikimedia.org api on graphite2004 is CRITICAL: HTTP CRITICAL: HTTP/1.1 502 Bad Gateway - 438 bytes in 0.080 second response time https://wikitech.wikimedia.org/wiki/Graphite%23Operations_troubleshooting [15:17:11] !log eevans@cumin1003 START - Cookbook sre.cassandra.roll-restart for nodes matching A:sessionstore: Set storage compatability to NONE — T435154 - eevans@cumin1003 [15:17:14] T435154: Upgrade sessionstore cluster to Cassandra 5.0.8 & JDK 17 - https://phabricator.wikimedia.org/T435154 [15:18:31] ^ the graphite noise is me, apologies. nothing to worry about [15:20:14] (03CR) 10Raymond Ndibe: "No I don't have the permission @fnegri@wikimedia.org , It'd be nice if you can help!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332528 (https://phabricator.wikimedia.org/T434556) (owner: 10Raymond Ndibe) [15:20:15] (03CR) 10Muehlenhoff: [C:03+2] Remove dbbackups for cumin1004 for the initial setup [puppet] - 10https://gerrit.wikimedia.org/r/1332752 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [15:21:28] (03PS1) 10Hnowlan: graphite: disable monitoring [puppet] - 10https://gerrit.wikimedia.org/r/1332756 (https://phabricator.wikimedia.org/T435341) [15:26:08] (03CR) 10SBassett: [C:03+1] "Normally we'd want CSP changes for Wikimedia domains to be Report-Only for a while, just to be safe, but this is probably fine to just do." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330446 (https://phabricator.wikimedia.org/T419684) (owner: 10Arendpieter) [15:27:53] (03CR) 10Federico Ceratto: [C:03+2] site.pp,core_test.yaml: shorten log retention test-s4 [puppet] - 10https://gerrit.wikimedia.org/r/1332747 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [15:28:34] !log eevans@cumin1003 END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:sessionstore: Set storage compatability to NONE — T435154 - eevans@cumin1003 [15:28:37] T435154: Upgrade sessionstore cluster to Cassandra 5.0.8 & JDK 17 - https://phabricator.wikimedia.org/T435154 [15:30:05] jan_drewniak: #bothumor My software never has bugs. It just develops random features. Rise for Wikimedia Portals Update. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1530). [15:32:19] !log Deployed patch for T435086 [15:32:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:34:17] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332743 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [15:34:28] FIRING: KeyholderUnarmed: 2 unarmed Keyholder key(s) on cumin1004:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [15:37:43] (03CR) 10Ladsgroup: "I ran it on a text and upload cache host just in case: https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=p" [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) (owner: 10Ladsgroup) [15:37:54] (03PS3) 10Ladsgroup: turnilo: Expose X-analytics thumb_generated in webrequest_sampled_live [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) [15:37:54] (03CR) 10Elukey: [C:03+2] docker_registry: expand the ML regex [puppet] - 10https://gerrit.wikimedia.org/r/1330254 (https://phabricator.wikimedia.org/T428022) (owner: 10Elukey) [15:38:03] (03CR) 10Ladsgroup: [V:03+2 C:03+2] turnilo: Expose X-analytics thumb_generated in webrequest_sampled_live [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) (owner: 10Ladsgroup) [15:39:30] (03CR) 10Muehlenhoff: [C:03+2] Replace grant for cumin2002 with cumin2003 [puppet] - 10https://gerrit.wikimedia.org/r/1331411 (https://phabricator.wikimedia.org/T427884) (owner: 10Muehlenhoff) [15:39:43] Amir1: okay to merge your patch along? [15:39:47] sure [15:40:54] I was about to ask you [15:41:19] (03PS1) 10Hnowlan: klaxon: add support for managers escalation [puppet] - 10https://gerrit.wikimedia.org/r/1332762 [15:41:36] jouncebot: nowandnext [15:41:36] For the next 0 hour(s) and 18 minute(s): Wikimedia Portals Update (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1530) [15:41:36] In 1 hour(s) and 18 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1700) [15:41:36] In 1 hour(s) and 18 minute(s): Wikidata Query Service weekly deploy (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1700) [15:44:10] (03CR) 10Tiziano Fogli: [C:03+1] graphite: disable monitoring [puppet] - 10https://gerrit.wikimedia.org/r/1332756 (https://phabricator.wikimedia.org/T435341) (owner: 10Hnowlan) [15:44:28] RESOLVED: KeyholderUnarmed: 1 unarmed Keyholder key(s) on cumin1004:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [15:44:43] (03CR) 10Hnowlan: [C:03+2] graphite: disable monitoring [puppet] - 10https://gerrit.wikimedia.org/r/1332756 (https://phabricator.wikimedia.org/T435341) (owner: 10Hnowlan) [15:44:52] moritzm: merged [15:48:23] thx [15:48:59] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, 06tools-infrastructure-team: cloudcephosd1035 offline - https://phabricator.wikimedia.org/T436520#12271704 (10VRiley-WMF) 05Open→03Resolved This unit is back up. I updated the BIOS and that seemed to correct the issue. [15:50:10] (03CR) 10Hnowlan: [C:03+2] idp: remove graphite configuration [puppet] - 10https://gerrit.wikimedia.org/r/1328165 (https://phabricator.wikimedia.org/T435341) (owner: 10Hnowlan) [15:50:13] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12271707 (10VRiley-WMF) [15:50:55] (03CR) 10Bking: [C:03+2] "self-merging, as the previous change was already approved and only rolled back due to hardware issues." [puppet] - 10https://gerrit.wikimedia.org/r/1332743 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [15:51:12] (03CR) 10Hnowlan: [C:03+2] "Done!" [puppet] - 10https://gerrit.wikimedia.org/r/1328165 (https://phabricator.wikimedia.org/T435341) (owner: 10Hnowlan) [15:55:26] (03CR) 10BCornwall: [C:03+1] "Yay cleanup!" [dns] - 10https://gerrit.wikimedia.org/r/1332687 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [16:00:38] FIRING: [2x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [16:00:52] (03CR) 10BCornwall: [C:03+1] varnish: Push webp on Safari except 14 and 15 [puppet] - 10https://gerrit.wikimedia.org/r/1332678 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [16:02:49] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{aux-k8s-worker[2006-2009].codfw.wmnet} and (A:aux-master-codfw or A:aux-worker-codfw) [16:02:49] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet [16:02:49] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet [16:02:49] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, 06tools-infrastructure-team: cloudcephosd1035 offline - https://phabricator.wikimedia.org/T436520#12271784 (10Andrew) I do not see anything obvious in the logs explaining the crash. I will put this back in service and we'll see how things go. [16:04:24] (03CR) 10BCornwall: [C:03+2] varnish: Push webp on Safari except 14 and 15 [puppet] - 10https://gerrit.wikimedia.org/r/1332678 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [16:06:06] (03CR) 10BCornwall: [V:03+1 C:03+2] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9356/co" [puppet] - 10https://gerrit.wikimedia.org/r/1332678 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [16:06:48] (03PS12) 10BryanDavis: MediaWiki: Only proxy existing .php files, otherwise return nice 404 [puppet] - 10https://gerrit.wikimedia.org/r/1100534 (https://phabricator.wikimedia.org/T382357) (owner: 10Bartosz Dziewoński) [16:06:52] !log cdobbins@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie [16:07:09] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12271795 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (... [16:07:13] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet [16:07:15] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet [16:07:22] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet [16:07:58] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet [16:08:48] (03CR) 10BryanDavis: "This appears to be the only local patch on deployment-puppetserver-1 that is not titled as a `[LOCAL HACK]`. Rebased to bump attention set" [puppet] - 10https://gerrit.wikimedia.org/r/1100534 (https://phabricator.wikimedia.org/T382357) (owner: 10Bartosz Dziewoński) [16:09:53] PROBLEM - SSH on cloudnet2005-dev is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [16:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:12:45] RECOVERY - SSH on cloudnet2005-dev is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [16:12:57] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet [16:12:58] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet [16:13:04] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet [16:14:44] 06SRE, 10Pywikibot, 06Traffic, 10Wikidata, and 3 others: Pywikibot reports maxlag retry error on Wikidata - https://phabricator.wikimedia.org/T421642#12271821 (10JeanFred) As I wrote back in T244030 (which was {T242081} back then): >>! In T244030#5842377, @JeanFred wrote: > I was under the impression that... [16:14:51] (03CR) 10Elukey: "To keep archives happy - me and Jesse will review this asap!" [puppet] - 10https://gerrit.wikimedia.org/r/1331768 (https://phabricator.wikimedia.org/T436393) (owner: 10Cwhite) [16:15:10] FIRING: BFDdown: BFD session down between cr1-esams and 185.15.59.138 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-esams:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:15:55] (03PS1) 10DLynch: editcheck-headless: Add deployment for the technical pilot [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) [16:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:18:08] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet [16:20:10] RESOLVED: BFDdown: BFD session down between cr1-esams and 185.15.59.138 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-esams:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:20:17] PROBLEM - Host cloudnet2006-dev is DOWN: PING CRITICAL - Packet loss = 100% [16:22:39] RECOVERY - Host cloudnet2006-dev is UP: PING OK - Packet loss = 0%, RTA = 39.52 ms [16:23:20] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet [16:23:22] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet [16:23:27] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet [16:24:01] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet [16:29:22] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet [16:29:23] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet [16:29:23] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{aux-k8s-worker[2006-2009].codfw.wmnet} and (A:aux-master-codfw or A:aux-worker-codfw) [16:29:49] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12271912 (10elukey) Hey folks, I didn't read the whole task since I got aware of it only from the SRE's te... [16:31:47] (03CR) 10DLynch: "I guess I don't feel too bad about not considering that one when looking at something that's not a python webapp. 😅" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326332 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [16:32:49] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12271932 (10elukey) @Andrew could you please list here what parameters you modified? It will help me into... [16:33:01] (03CR) 10Majavah: [C:03+2] aptrepo: Stop mirroring Helm upstream repos [puppet] - 10https://gerrit.wikimedia.org/r/1330356 (owner: 10Majavah) [16:34:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr1-codfw and cr1-esams (185.15.59.139) - group Confed_esams - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [16:35:23] FIRING: [3x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [16:43:55] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [16:50:23] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [16:53:20] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12271995 (10elukey) Dumping some stuff found on cloudvirt1077 - I see the following BIOS option via Redfis... [17:00:05] swfrench-wmf: Time to snap out of that daydream and deploy MediaWiki infrastructure (UTC late). (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1700). [17:00:05] ryankemper: That opportune time for a Wikidata Query Service weekly deploy deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1700). [17:00:16] o/ [17:01:14] it'll take me a couple of minutes to prepare, but I plan to deploy a couple of cleanups to rest-gateway during this infra window [17:01:50] (03PS7) 10Scott French: api-gateway: Drop support for debug_hosts [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305236 (https://phabricator.wikimedia.org/T433752) [17:01:50] (03PS6) 10Scott French: api-gateway: Remove stale test assertion and noop Lua code [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311874 (https://phabricator.wikimedia.org/T433752) [17:01:50] (03PS6) 10Scott French: api-gateway: Drop support for php_engine_routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311875 (https://phabricator.wikimedia.org/T433752) [17:01:51] (03PS5) 10Scott French: api-gateway: Basic cluster specifier support and Lua plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) [17:01:52] (03PS7) 10Scott French: api-gateway: Support x-wikimedia-debug routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311964 (https://phabricator.wikimedia.org/T433752) [17:01:53] (03PS4) 10Scott French: api-gateway: Support Host-based diversion in the mw-api plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322140 (https://phabricator.wikimedia.org/T433752) [17:03:28] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations: attempts to reimage cp5022 are failing - https://phabricator.wikimedia.org/T436386#12272026 (10CDobbins) @LSobanski another reimage attempt today failed for the same reason [17:04:36] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1020.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:05:08] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:05:59] (03PS2) 10CDanis: klaxon: add support for managers escalation [puppet] - 10https://gerrit.wikimedia.org/r/1332762 (owner: 10Hnowlan) [17:06:00] (03CR) 10CDanis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332762 (owner: 10Hnowlan) [17:06:36] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [17:07:08] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:09:27] (03PS3) 10CDanis: klaxon: add support for managers escalation [puppet] - 10https://gerrit.wikimedia.org/r/1332762 (owner: 10Hnowlan) [17:09:29] (03CR) 10CDanis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1332762 (owner: 10Hnowlan) [17:09:48] FIRING: PuppetFailure: Puppet has failed on cumin1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [17:11:45] (03CR) 10CDanis: [C:03+1] klaxon: add support for managers escalation [puppet] - 10https://gerrit.wikimedia.org/r/1332762 (owner: 10Hnowlan) [17:14:45] (03PS1) 10Hnowlan: graphite: remove module, references, config [puppet] - 10https://gerrit.wikimedia.org/r/1332767 (https://phabricator.wikimedia.org/T435340) [17:16:42] (03CR) 10Scott French: [C:03+2] api-gateway: Drop support for debug_hosts [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305236 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:19:40] (03Merged) 10jenkins-bot: api-gateway: Drop support for debug_hosts [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305236 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:22:39] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [17:23:22] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [17:25:52] (03CR) 10Scott French: [C:03+2] api-gateway: Remove stale test assertion and noop Lua code [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311874 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:28:02] jouncebot: nowandnext [17:28:02] For the next 0 hour(s) and 31 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1700) [17:28:02] For the next 0 hour(s) and 1 minute(s): Wikidata Query Service weekly deploy (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T1700) [17:28:02] In 2 hour(s) and 31 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T2000) [17:28:51] (03Merged) 10jenkins-bot: api-gateway: Remove stale test assertion and noop Lua code [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311874 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:32:01] (03CR) 10Ssingh: [C:03+1] Add cookbook to roll-restart/reboot URL downloaders [cookbooks] - 10https://gerrit.wikimedia.org/r/1329523 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [17:32:04] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [17:32:21] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [17:32:28] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [17:32:33] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [17:33:25] (03PS1) 10Majavah: P:toolforge::bastion: Remove temporary index migration script [puppet] - 10https://gerrit.wikimedia.org/r/1332769 (https://phabricator.wikimedia.org/T401818) [17:33:28] (03PS1) 10Majavah: O:wmcs::toolforge: Remove Elasticsearch related classes [puppet] - 10https://gerrit.wikimedia.org/r/1332770 (https://phabricator.wikimedia.org/T401818) [17:33:30] (03PS1) 10Majavah: P:toolforge: Remove unused apt_pinning class [puppet] - 10https://gerrit.wikimedia.org/r/1332771 [17:33:30] (03PS1) 10Majavah: P:elasticsearch: Remove unused profiles [puppet] - 10https://gerrit.wikimedia.org/r/1332772 [17:33:30] (03PS1) 10Majavah: elasticsearch: Remove unused classes [puppet] - 10https://gerrit.wikimedia.org/r/1332773 [17:34:35] (03CR) 10Scott French: [C:03+2] api-gateway: Drop support for php_engine_routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311875 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:38:30] (03Merged) 10jenkins-bot: api-gateway: Drop support for php_engine_routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311875 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [17:39:15] (03CR) 10Ahmon Dancy: [C:03+1] scap: remove statsd config [puppet] - 10https://gerrit.wikimedia.org/r/1332699 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [17:41:02] !log swfrench@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [17:41:36] !log swfrench@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [17:42:49] !log swfrench@deploy1003 helmfile [codfw] START helmfile.d/services/rest-gateway: apply [17:43:24] !log swfrench@deploy1003 helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply [17:47:42] (03CR) 10Ssingh: "very interesting approach. naive question: is there a reason to prefer this approach over having a class parameter?" [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [17:52:52] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12272191 (10bking) I've rolled out the patch, and load remains low on the `krb1004` replica. If nothing has changed by Thursday, I'll try removing... [17:53:09] !log swfrench@deploy1003 helmfile [eqiad] START helmfile.d/services/rest-gateway: apply [17:53:35] !log swfrench@deploy1003 helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply [17:55:06] alright, I'll continue to monitor things on the rest-gateway side, but in the absence of surprises, I'm done with the infra window. [18:00:16] (03CR) 10RLazarus: [C:03+1] "Good catch." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331857 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [18:02:38] jouncebot: nowandnext [18:02:39] No deployments scheduled for the next 1 hour(s) and 57 minute(s) [18:02:39] In 1 hour(s) and 57 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T2000) [18:02:44] cool cool [18:03:43] (03PS1) 10Ladsgroup: Revert^4 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332776 [18:04:38] (03PS1) 10Majavah: P:toolforge::prometheus: Remove Elasticsearch targets [puppet] - 10https://gerrit.wikimedia.org/r/1332777 (https://phabricator.wikimedia.org/T401818) [18:04:41] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 31 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [18:05:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332748 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [18:06:30] (03Merged) 10jenkins-bot: Enable thumb.wikimedia.org everywhere except enwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332748 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [18:06:54] (03PS2) 10Ladsgroup: Revert^4 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332776 [18:08:19] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1332748|Enable thumb.wikimedia.org everywhere except enwiki (T427465)]] [18:08:22] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [18:10:10] 10ops-ulsfo, 06SRE, 06DC-Ops: ulsfo: OOB migration from copper to fiber - https://phabricator.wikimedia.org/T436499#12272241 (10wiki_willy) a:03RobH Assigning to Rob, and replied back to the Digital Realty email that he'll be the point person. [18:10:29] (03CR) 10Scott French: "Thanks for the review!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331857 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [18:10:50] (03CR) 10Scott French: [C:03+2] Rakefile: Fix 'upstream' mock in services_proxy data [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331857 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [18:12:52] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1332748|Enable thumb.wikimedia.org everywhere except enwiki (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:14:41] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [18:15:35] (03CR) 10Scott French: "@rlazarus@wikimedia.org FYI (so the UI doesn't hide it from you)" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331872 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [18:16:22] (03CR) 10BCornwall: [V:03+1] "There already is a class parameter. The issue is that the role will attempt to parse /etc/network/interfaces anytime the role is included " [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [18:19:58] (03PS1) 10Fabfur: external_cloud_vendors: set datetime from yesterday on ripe query [puppet] - 10https://gerrit.wikimedia.org/r/1332778 [18:20:13] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1332748|Enable thumb.wikimedia.org everywhere except enwiki (T427465)]] (duration: 11m 53s) [18:20:16] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [18:20:41] (03CR) 10CI reject: [V:04-1] external_cloud_vendors: set datetime from yesterday on ripe query [puppet] - 10https://gerrit.wikimedia.org/r/1332778 (owner: 10Fabfur) [18:23:20] (03CR) 10CDanis: ipip: Don't touch interfaces file if nonexistent (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [18:26:49] (03CR) 10Southparkfan: deployment-prep: add svc for Wikifeeds [puppet] - 10https://gerrit.wikimedia.org/r/1332499 (https://phabricator.wikimedia.org/T436462) (owner: 10Southparkfan) [18:29:36] (03PS3) 10BCornwall: ipip: Don't touch interfaces file if nonexistent [puppet] - 10https://gerrit.wikimedia.org/r/1330654 [18:35:59] (03CR) 10CDanis: [C:04-1] external_cloud_vendors: set datetime from yesterday on ripe query (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1332778 (owner: 10Fabfur) [18:37:27] (03CR) 10Ssingh: "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [18:45:32] (03Merged) 10jenkins-bot: Rakefile: Fix 'upstream' mock in services_proxy data [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331857 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [19:04:03] (03CR) 10RLazarus: [C:03+1] Rakefile: Simplify and update 'upstream' mock data population (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331872 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [19:16:37] (03PS1) 10Tsevener: Add main page exclusions to Apple App Site Association File [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332794 (https://phabricator.wikimedia.org/T432412) [19:19:24] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, August 31 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-ite" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332794 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [19:37:07] (03CR) 10RLazarus: [C:03+1] "Thanks! One last suggestion but feel free to merge without another roundtrip." [puppet] - 10https://gerrit.wikimedia.org/r/1329607 (https://phabricator.wikimedia.org/T433547) (owner: 10Clément Goubert) [19:41:09] (03CR) 10DLynch: "Context: this is more or less just the helmfile.d parts of If5e485667922b354c2446dbf0563e95536a431dc, after @cgoubert@wikimedia.org pointe" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [19:49:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-codfw and cr1-eqiad (208.80.153.220) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [19:49:58] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/0/2 (Transport: cr1-eqiad:et-1/1/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [19:52:24] (03CR) 10Majavah: [C:03+2] deployment-prep: add svc for Wikifeeds [puppet] - 10https://gerrit.wikimedia.org/r/1332499 (https://phabricator.wikimedia.org/T436462) (owner: 10Southparkfan) [19:52:36] (03CR) 10Majavah: [C:03+2] "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1332499 (https://phabricator.wikimedia.org/T436462) (owner: 10Southparkfan) [19:58:08] (03CR) 10Ssingh: [C:03+1] "Looks good, thanks for working on this!" [puppet] - 10https://gerrit.wikimedia.org/r/1329663 (owner: 10CDanis) [19:59:03] (03CR) 10CDanis: [C:03+2] geodns: also export pooledness within each PoP [puppet] - 10https://gerrit.wikimedia.org/r/1329663 (owner: 10CDanis) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: It is that lovely time of the day again! You are hereby commanded to deploy UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T2000). [20:00:05] arlolra and toni_: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:30] o/ [20:01:10] FIRING: BFDdown: BFD session down between cr1-codfw and fe80::5e5e:abff:fe3d:8198 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [20:01:49] o/ [20:02:34] I can get started [20:04:39] (03PS4) 10Scott French: Rakefile: Simplify and update 'upstream' mock data population [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331872 (https://phabricator.wikimedia.org/T427666) [20:04:39] (03PS4) 10Scott French: Rakefile: Update mock services_proxy data for splits [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) [20:05:10] (03PS1) 10Majavah: openstack: puppet-enc: Allow MEDIUMTEXT worth of Hiera [puppet] - 10https://gerrit.wikimedia.org/r/1332802 (https://phabricator.wikimedia.org/T436561) [20:05:44] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [20:06:10] RESOLVED: BFDdown: BFD session down between cr1-codfw and fe80::5e5e:abff:fe3d:8198 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [20:06:44] (03Merged) 10jenkins-bot: switch testwiki to use parsoid [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1323180 (https://phabricator.wikimedia.org/T431636) (owner: 10Peterxy12) [20:06:58] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1323180|switch testwiki to use parsoid (T431636)]] [20:07:01] T431636: switch testwiki to use parsoid - https://phabricator.wikimedia.org/T431636 [20:08:47] (03PS2) 10Majavah: openstack: puppet-enc: Allow MEDIUMTEXT worth of Hiera [puppet] - 10https://gerrit.wikimedia.org/r/1332802 (https://phabricator.wikimedia.org/T436561) [20:09:46] (03CR) 10Scott French: "Thank you, Reuven!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331872 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [20:10:55] !log arlolra@deploy1003 peterxy12, arlolra: Backport for [[gerrit:1323180|switch testwiki to use parsoid (T431636)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:11:21] (03CR) 10Scott French: "This is the last one of these for now. As noted in the other comment, and as you would expect given the state of puppet, it produces no di" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [20:11:38] !log arlolra@deploy1003 peterxy12, arlolra: Continuing with deployment [20:11:52] 06SRE, 06Traffic: Prometheus Ferm MSS export does not work for IPv6 - https://phabricator.wikimedia.org/T433672#12272699 (10CDobbins) I wasn't able to reproduce this on clouddumps1002: ` cdobbins@clouddumps1002:~$ sudo /usr/local/bin/prometheus-ferm-mss -o /tmp/test_run.prom -e 208.80.154.242:2049 -e 208.80.1... [20:14:58] (03PS1) 10Majavah: hieradata: Migrate deployment-prep data to ENC [puppet] - 10https://gerrit.wikimedia.org/r/1332804 (https://phabricator.wikimedia.org/T277680) [20:16:57] FIRING: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:18:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 20.7% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:18:27] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1323180|switch testwiki to use parsoid (T431636)]] (duration: 11m 29s) [20:18:31] T431636: switch testwiki to use parsoid - https://phabricator.wikimedia.org/T431636 [20:19:03] toni_: on to you, unless you need me to deploy [20:19:32] actually I do need someone to deploy, thanks! 🙏 [20:19:48] np [20:20:30] (03CR) 10TrainBranchBot: [C:03+2] "Approved by arlolra@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332794 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [20:21:57] RESOLVED: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:22:02] (03Merged) 10jenkins-bot: Add main page exclusions to Apple App Site Association File [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332794 (https://phabricator.wikimedia.org/T432412) (owner: 10Tsevener) [20:22:13] !log arlolra@deploy1003 Started scap sync-world: Backport for [[gerrit:1332794|Add main page exclusions to Apple App Site Association File (T432412)]] [20:22:16] T432412: [Eng] iOS - Update app site association file - https://phabricator.wikimedia.org/T432412 [20:22:27] FIRING: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:25:01] (03CR) 10RLazarus: [C:03+1] "🚀" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [20:26:19] !log arlolra@deploy1003 tsev, arlolra: Backport for [[gerrit:1332794|Add main page exclusions to Apple App Site Association File (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:27:04] toni_: if you have something to test [20:27:49] (03CR) 10Southparkfan: [C:03+1] "LGTM, one non-blocking nit." [puppet] - 10https://gerrit.wikimedia.org/r/1332804 (https://phabricator.wikimedia.org/T277680) (owner: 10Majavah) [20:27:57] FIRING: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:28:18] (03PS2) 10Majavah: hieradata: Migrate deployment-prep data to ENC [puppet] - 10https://gerrit.wikimedia.org/r/1332804 (https://phabricator.wikimedia.org/T277680) [20:28:27] looks good! [20:28:44] !log arlolra@deploy1003 tsev, arlolra: Continuing with deployment [20:29:35] (03CR) 10Majavah: [C:03+2] hieradata: Migrate deployment-prep data to ENC (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1332804 (https://phabricator.wikimedia.org/T277680) (owner: 10Majavah) [20:29:42] (03CR) 10Southparkfan: [C:03+1] hieradata: Migrate deployment-prep data to ENC [puppet] - 10https://gerrit.wikimedia.org/r/1332804 (https://phabricator.wikimedia.org/T277680) (owner: 10Majavah) [20:32:12] RESOLVED: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:32:27] FIRING: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:32:57] RESOLVED: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:33:17] !log arlolra@deploy1003 Finished scap sync-world: Backport for [[gerrit:1332794|Add main page exclusions to Apple App Site Association File (T432412)]] (duration: 11m 04s) [20:33:20] T432412: [Eng] iOS - Update app site association file - https://phabricator.wikimedia.org/T432412 [20:33:34] !incidents [20:33:35] 8314 (RESOLVED) ProbeDown sre (10.2.2.70 ip4 chart-renderer:30443 probes/service http_chart-renderer_ip4 eqiad) [20:33:49] all done [20:34:16] !log kubectl -n chart-renderer scale deployment chart-renderer-production --replicas 4 [20:34:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:35:24] (03CR) 10Jdlrobson: [C:03+1] Enable ReadingLists for logged-in users on phase 1 wikis (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332717 (https://phabricator.wikimedia.org/T434922) (owner: 10Aude) [20:37:12] RESOLVED: ProbeDown: Service chart-renderer:30443 has failed probes (http_chart-renderer_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#chart-renderer:30443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:38:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-int releases routed via main at eqiad: 22.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-int&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [20:39:08] thank you! [20:50:38] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [20:51:38] (03PS1) 10CDanis: chart-renderer: scale 2 --> 4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332810 (https://phabricator.wikimedia.org/T372081) [20:53:39] (03CR) 10BCornwall: ipip: Don't touch interfaces file if nonexistent (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [20:53:42] (03CR) 10Jasmine: [C:03+1] chart-renderer: scale 2 --> 4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332810 (https://phabricator.wikimedia.org/T372081) (owner: 10CDanis) [20:55:12] (03CR) 10CDanis: [C:03+1] ipip: Don't touch interfaces file if nonexistent (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [20:55:23] (03CR) 10CDanis: [C:03+2] chart-renderer: scale 2 --> 4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332810 (https://phabricator.wikimedia.org/T372081) (owner: 10CDanis) [20:55:44] !log tsev@deploy1003 mwscript-k8s job started: purgeList.php # T432412 [20:55:47] T432412: [Eng] iOS - Update app site association file - https://phabricator.wikimedia.org/T432412 [20:57:19] (03CR) 10Scott French: [C:03+2] Rakefile: Simplify and update 'upstream' mock data population [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331872 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [20:57:33] (03CR) 10BCornwall: [C:03+2] ipip: Don't touch interfaces file if nonexistent [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [20:57:41] (03Merged) 10jenkins-bot: chart-renderer: scale 2 --> 4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332810 (https://phabricator.wikimedia.org/T372081) (owner: 10CDanis) [20:58:32] !log cdanis@deploy1003 helmfile [eqiad] START helmfile.d/services/chart-renderer: apply [20:58:35] !log cdanis@deploy1003 helmfile [eqiad] DONE helmfile.d/services/chart-renderer: apply [20:58:50] !log cdanis@deploy1003 helmfile [codfw] START helmfile.d/services/chart-renderer: apply [20:59:07] !log cdanis@deploy1003 helmfile [codfw] DONE helmfile.d/services/chart-renderer: apply [21:00:05] alexsanford, Reedy, sbassett, Maryum, and manfredi: I, the Bot under the Fountain, call upon thee, The Deployer, to do Weekly Security deployment window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260831T2100). [21:04:38] (03CR) 10Fabfur: external_cloud_vendors: set datetime from yesterday on ripe query (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1332778 (owner: 10Fabfur) [21:05:37] (03PS2) 10Fabfur: external_cloud_vendors: set datetime from yesterday on ripe query [puppet] - 10https://gerrit.wikimedia.org/r/1332778 [21:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [21:10:03] FIRING: PuppetFailure: Puppet has failed on cumin1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [21:10:06] (03PS1) 10BCornwall: interface: Don't touch interfaces file if absent [puppet] - 10https://gerrit.wikimedia.org/r/1332813 [21:10:19] (03CR) 10CDanis: [C:03+1] external_cloud_vendors: set datetime from yesterday on ripe query [puppet] - 10https://gerrit.wikimedia.org/r/1332778 (owner: 10Fabfur) [21:11:02] (03PS2) 10Aude: Enable ReadingLists for logged-in users on phase 1 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1332717 (https://phabricator.wikimedia.org/T434922) [21:17:23] Hey all - I have one sec patch I’d like to deploy. Are the backports done? [21:23:03] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): eqiad row A&B host migration details request for Search Platform - https://phabricator.wikimedia.org/T432651#12272958 (10bking) 05Open→03Resolved Looks like Search Platform's work here is finished, so I'm closing out this... [21:25:23] FIRING: [4x] CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [21:25:42] 06SRE, 10corto, 10Incident Tooling: Increase trusted volunteer's visibility into production incidents - https://phabricator.wikimedia.org/T426137#12272967 (10SomeRandomDeveloper) Since the view policy for incident tasks was changed to #wmf-nda, users like me who are only in #acl_security but not in #wmf-nda... [21:30:45] !log Deployed security fix for T435863 [21:30:46] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:32:45] (03Merged) 10jenkins-bot: Rakefile: Simplify and update 'upstream' mock data population [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331872 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [21:46:07] (03CR) 10Scott French: [C:03+2] Rakefile: Update mock services_proxy data for splits [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [21:46:18] (03CR) 10CI reject: [V:04-1] Rakefile: Update mock services_proxy data for splits [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [21:46:45] (03PS5) 10Scott French: Rakefile: Update mock services_proxy data for splits [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328274 (https://phabricator.wikimedia.org/T427666)