[00:05:30] (03CR) 10RLazarus: "Thanks for this!" [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [00:15:30] Amir1: Sorry for the belated ping but I'm done [00:15:49] no worries. I just started the deploy! [00:15:58] it wasn't anything important [00:16:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329317 (https://phabricator.wikimedia.org/T435979) (owner: 10Ladsgroup) [00:17:10] (03Merged) 10jenkins-bot: Allow linking to thumb.wikimedia.org in TemplateStyles [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329317 (https://phabricator.wikimedia.org/T435979) (owner: 10Ladsgroup) [00:18:22] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1329317|Allow linking to thumb.wikimedia.org in TemplateStyles (T435979)]] [00:18:26] T435979: TemplateStyles should allow linking to thumb.wikimedia.org - https://phabricator.wikimedia.org/T435979 [00:19:07] (03CR) 10RLazarus: [C:03+1] "I haven't looked at the template changes in the dependent I236e1661 yet, but in a vacuum, LGTM. (Happy to though, let me know if that's re" [puppet] - 10https://gerrit.wikimedia.org/r/1328256 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [00:22:49] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1329317|Allow linking to thumb.wikimedia.org in TemplateStyles (T435979)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [00:32:46] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [00:37:04] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329317|Allow linking to thumb.wikimedia.org in TemplateStyles (T435979)]] (duration: 18m 42s) [00:37:10] T435979: TemplateStyles should allow linking to thumb.wikimedia.org - https://phabricator.wikimedia.org/T435979 [01:03:41] (03PS1) 10BCornwall: wmfuniq: Use narrower JSONDecodeError exception [puppet] - 10https://gerrit.wikimedia.org/r/1329702 [01:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [01:11:17] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1329704 [01:11:17] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1329704 (owner: 10TrainBranchBot) [01:19:50] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1329704 (owner: 10TrainBranchBot) [01:44:38] (03PS1) 10David Martin: wikifunctions: Switch ORCHESTRATOR_CALLBACK_API_URI to localhost:6520 in staging evaluators [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329706 (https://phabricator.wikimedia.org/T415616) [01:47:50] (03CR) 10David Martin: [C:03+2] wikifunctions: Switch ORCHESTRATOR_CALLBACK_API_URI to localhost:6520 in staging evaluators [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329706 (https://phabricator.wikimedia.org/T415616) (owner: 10David Martin) [01:50:16] (03Merged) 10jenkins-bot: wikifunctions: Switch ORCHESTRATOR_CALLBACK_API_URI to localhost:6520 in staging evaluators [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329706 (https://phabricator.wikimedia.org/T415616) (owner: 10David Martin) [01:52:52] PROBLEM - Host wikikube-worker1260 is DOWN: PING CRITICAL - Packet loss = 100% [01:55:16] !log dmartin@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [01:57:52] !log dmartin@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [02:00:41] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:08:41] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 08m 00s) [02:23:10] FIRING: BFDdown: BFD session down between cr3-eqsin and fe80::669:8f02:1ece:f4ef - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr3-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:28:10] RESOLVED: BFDdown: BFD session down between cr3-eqsin and fe80::669:8f02:1ece:f4ef - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr3-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:37:04] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [02:46:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:29:21] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324965 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [03:30:40] (03Merged) 10jenkins-bot: Enable Produnto on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324965 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [03:31:07] !log tstarling@deploy1003 Started scap sync-world: Backport for [[gerrit:1324965|Enable Produnto on testwiki (T421436)]] [03:31:11] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [03:35:51] !log tstarling@deploy1003 tstarling: Backport for [[gerrit:1324965|Enable Produnto on testwiki (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [04:06:31] !log tstarling@deploy1003 tstarling: Continuing with deployment [04:11:45] !log tstarling@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324965|Enable Produnto on testwiki (T421436)]] (duration: 40m 38s) [04:11:50] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [04:25:53] (03PS1) 10Tim Starling: Produnto: Add IPv6 range for GitLab [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329728 (https://phabricator.wikimedia.org/T421436) [04:34:06] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [puppet] - 10https://gerrit.wikimedia.org/r/1329393 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [05:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [05:41:11] (03PS2) 10Giuseppe Lavagetto: aptrepo: add thirdparty/gvisor to newer distros as well [puppet] - 10https://gerrit.wikimedia.org/r/1328540 (https://phabricator.wikimedia.org/T435758) [05:41:11] (03PS1) 10Giuseppe Lavagetto: profile::containerd: format according to our style guide [puppet] - 10https://gerrit.wikimedia.org/r/1330128 [05:41:11] (03PS1) 10Giuseppe Lavagetto: profile::containerd: add gVisor support [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) [05:41:15] (03PS1) 10Giuseppe Lavagetto: containerd: use hiera parameters to detect if dragonfly is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1330130 [05:41:15] (03PS1) 10Giuseppe Lavagetto: kubernetes::staging: make containerd support gVisor runtime [puppet] - 10https://gerrit.wikimedia.org/r/1330131 (https://phabricator.wikimedia.org/T435796) [05:42:41] (03CR) 10CI reject: [V:04-1] profile::containerd: add gVisor support [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T0600) [06:00:05] marostegui, Amir1, and federico3: I, the Bot under the Fountain, call upon thee, The Deployer, to do Primary database switchover deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T0600). [06:32:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:34:41] (03PS1) 10Giuseppe Lavagetto: profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) [06:36:08] (03CR) 10CI reject: [V:04-1] profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [06:37:04] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [06:37:10] RESOLVED: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:46:39] jouncebot: next [06:46:40] In 0 hour(s) and 13 minute(s): UTC morning backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T0700) [06:46:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:56:33] (03CR) 10Svantje Lilienthal: [C:03+1] Add feature flag to beta cluster and test wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328590 (https://phabricator.wikimedia.org/T431544) (owner: 10Mareike Heuer) [07:00:05] Amir1, urbanecm, and awight: Your horoscope predicts another UTC morning backport window deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T0700). [07:00:05] No Gerrit patches in the queue for this window AFAICS. [07:14:02] (03CR) 10Slyngshede: [C:03+2] P:idp allow selective mfa enablement [puppet] - 10https://gerrit.wikimedia.org/r/1304784 (https://phabricator.wikimedia.org/T277841) (owner: 10Slyngshede) [07:16:58] FYI: We'll deploy 1328590 now, just sneaked that into the calendar [07:20:09] (03CR) 10TrainBranchBot: [C:03+2] "Approved by wmde-fisch@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328590 (https://phabricator.wikimedia.org/T431544) (owner: 10Mareike Heuer) [07:21:10] (03Merged) 10jenkins-bot: Add feature flag to beta cluster and test wiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328590 (https://phabricator.wikimedia.org/T431544) (owner: 10Mareike Heuer) [07:21:55] !log wmde-fisch@deploy1003 Started scap sync-world: Backport for [[gerrit:1328590|Add feature flag to beta cluster and test wiki (T431544)]] [07:22:00] T431544: Use translatable citation type name to generated autonames instead of :n - https://phabricator.wikimedia.org/T431544 [07:24:05] (03CR) 10Majavah: [V:03+1 C:03+2] Remove various references to the puppetmaster class [puppet] - 10https://gerrit.wikimedia.org/r/1329540 (owner: 10Majavah) [07:26:52] !log wmde-fisch@deploy1003 wmde-fisch, mareikeheuer: Backport for [[gerrit:1328590|Add feature flag to beta cluster and test wiki (T431544)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:29:39] !log wmde-fisch@deploy1003 wmde-fisch, mareikeheuer: Continuing with deployment [07:34:02] !log wmde-fisch@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328590|Add feature flag to beta cluster and test wiki (T431544)]] (duration: 12m 07s) [07:34:07] T431544: Use translatable citation type name to generated autonames instead of :n - https://phabricator.wikimedia.org/T431544 [07:35:16] FIRING: [2x] ProbeDown: Service idp2005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:35:19] Done, nothing else in the Window. :-) [07:40:15] RESOLVED: [2x] ProbeDown: Service idp2005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:54:06] (03PS1) 10Elukey: docker_registry: expand the ML regex [puppet] - 10https://gerrit.wikimedia.org/r/1330254 (https://phabricator.wikimedia.org/T428022) [07:58:42] (03CR) 10Filippo Giunchedi: [V:03+1 C:03+2] labstore: clean up traffic_shaping [puppet] - 10https://gerrit.wikimedia.org/r/1327853 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [07:58:44] (03CR) 10Filippo Giunchedi: [V:03+1 C:03+2] wmcs: remove support for clouddumps client symlinks [puppet] - 10https://gerrit.wikimedia.org/r/1327854 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [07:58:46] (03CR) 10Filippo Giunchedi: [V:03+1 C:03+2] dumps: remove production support for non-lb mounts [puppet] - 10https://gerrit.wikimedia.org/r/1327855 (https://phabricator.wikimedia.org/T435581) (owner: 10Filippo Giunchedi) [08:00:05] hashar and andre: Deploy window MediaWiki train - Utc-0 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T0800) [08:01:46] (03PS1) 10TrainBranchBot: group2 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330272 (https://phabricator.wikimedia.org/T430836) [08:01:50] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by hashar@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330272 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [08:01:58] I am running the MediaWiki train [08:02:14] WMDE-Fisch: +1 on the deploy well done thanks :) [08:03:02] (03Merged) 10jenkins-bot: group2 to 1.47.0-wmf.17 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330272 (https://phabricator.wikimedia.org/T430836) (owner: 10TrainBranchBot) [08:11:17] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [08:12:15] !log hashar@deploy1003 rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.17 refs T430836 [08:12:20] T430836: 1.47.0-wmf.17 deployment blockers - https://phabricator.wikimedia.org/T430836 [08:26:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [08:27:51] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [08:28:53] (03PS1) 10Joal: Enable presto join spill config [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) [08:35:19] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [08:35:56] (03PS4) 10Elukey: cfssl: limit key sizes to values defined in error string [puppet] - 10https://gerrit.wikimedia.org/r/1329572 (owner: 10Cwhite) [08:36:05] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329572 (owner: 10Cwhite) [08:40:04] (03CR) 10Elukey: [C:03+1] "I sampled some hosts using c:profile::pki::client (that include the cfssl::cert define) and afaics it is a no-op. Since the scope is very " [puppet] - 10https://gerrit.wikimedia.org/r/1329572 (owner: 10Cwhite) [08:42:19] (03CR) 10Elukey: [C:03+2] "Nevermind, it won't cause any issues, possibly puppet failures but they will be easy to spot. Thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1329572 (owner: 10Cwhite) [08:46:23] (03PS1) 10Slyngshede: P:IDP typed mfa options [puppet] - 10https://gerrit.wikimedia.org/r/1330298 [08:47:54] (03PS2) 10Raymond Ndibe: webservice-runner: default to 8000 only if PORT and TOOL_WEB_PORT has no value [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1310212 (https://phabricator.wikimedia.org/T432078) [08:47:54] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1330298 (owner: 10Slyngshede) [08:48:58] (03CR) 10Raymond Ndibe: "resuming this because we've gotten to the place were we need to merge this" [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1310212 (https://phabricator.wikimedia.org/T432078) (owner: 10Raymond Ndibe) [08:49:16] (03CR) 10Elukey: kafka: converge topic config from hieradata (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [09:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [09:11:05] (03PS2) 10Slyngshede: P:IDP typed mfa options [puppet] - 10https://gerrit.wikimedia.org/r/1330298 [09:18:41] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1330298 (owner: 10Slyngshede) [09:33:46] (03PS1) 10Jelto: profile::lists::monitoring: remove hardcoded number of uwsgi processes [puppet] - 10https://gerrit.wikimedia.org/r/1330312 (https://phabricator.wikimedia.org/T436047) [09:34:21] (03CR) 10CI reject: [V:04-1] profile::lists::monitoring: remove hardcoded number of uwsgi processes [puppet] - 10https://gerrit.wikimedia.org/r/1330312 (https://phabricator.wikimedia.org/T436047) (owner: 10Jelto) [09:34:49] (03PS1) 10Fabfur: cache::haproxy: fix logging of normalized Host header [puppet] - 10https://gerrit.wikimedia.org/r/1330314 (https://phabricator.wikimedia.org/T434766) [09:35:35] (03PS2) 10Jelto: profile::lists::monitoring: remove hardcoded number of uwsgi processes [puppet] - 10https://gerrit.wikimedia.org/r/1330312 (https://phabricator.wikimedia.org/T436047) [09:37:32] (03CR) 10Jelto: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9335/co" [puppet] - 10https://gerrit.wikimedia.org/r/1330312 (https://phabricator.wikimedia.org/T436047) (owner: 10Jelto) [09:39:47] (03CR) 10Jelto: [V:03+1] "The bump of uwsgi processes triggered icinga alerts: https://alerts.wikimedia.org/?q=%40state%3Dactive&q=instance%3Dlists1004" [puppet] - 10https://gerrit.wikimedia.org/r/1330312 (https://phabricator.wikimedia.org/T436047) (owner: 10Jelto) [09:41:05] (03CR) 10Fabfur: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1330314 (https://phabricator.wikimedia.org/T434766) (owner: 10Fabfur) [09:55:26] (03PS3) 10Slyngshede: P:IDP typed mfa options [puppet] - 10https://gerrit.wikimedia.org/r/1330298 [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T1000) [10:05:23] (03CR) 10Blake: [C:03+2] mw-videoscaler: bump envoy to 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329565 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:07:54] (03Merged) 10jenkins-bot: mw-videoscaler: bump envoy to 1.39.0-1. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329565 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:09:20] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-videoscaler: apply [10:12:37] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-videoscaler: apply [10:12:43] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-videoscaler: apply [10:12:53] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-videoscaler: apply [10:22:35] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 27 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [skins/MinervaNeue] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329672 (https://phabricator.wikimedia.org/T436141) (owner: 10LWatson) [10:28:31] (03PS1) 10Giuseppe Lavagetto: admin: add support for gVisor RuntimeClass handlers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) [10:30:23] (03PS1) 10Majavah: kubeadm::helm: Use helm3.17 from bookworm-wikimedia [puppet] - 10https://gerrit.wikimedia.org/r/1330345 [10:30:58] (03CR) 10FNegri: [C:03+1] kubeadm::helm: Use helm3.17 from bookworm-wikimedia [puppet] - 10https://gerrit.wikimedia.org/r/1330345 (owner: 10Majavah) [10:32:03] (03CR) 10FNegri: [C:03+2] kubeadm::helm: Use helm3.17 from bookworm-wikimedia [puppet] - 10https://gerrit.wikimedia.org/r/1330345 (owner: 10Majavah) [10:33:22] (03CR) 10Hnowlan: [C:03+2] team-sre: Add data-engineering tag [alerts] - 10https://gerrit.wikimedia.org/r/1319827 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [10:36:23] (03Merged) 10jenkins-bot: team-sre: Add data-engineering tag [alerts] - 10https://gerrit.wikimedia.org/r/1319827 (https://phabricator.wikimedia.org/T432376) (owner: 10Hnowlan) [10:41:29] (03CR) 10Hnowlan: opensearch: pass curator username/password through server profile (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1329393 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [10:46:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:54:50] (03CR) 10Ladsgroup: [C:03+1] "I'm not sure why an alert checks for that but this fixes the problem." [puppet] - 10https://gerrit.wikimedia.org/r/1330312 (https://phabricator.wikimedia.org/T436047) (owner: 10Jelto) [11:02:11] (03PS1) 10Majavah: aptrepo: Stop mirroring Helm upstream repos [puppet] - 10https://gerrit.wikimedia.org/r/1330356 [11:07:00] (03CR) 10Hnowlan: [C:03+1] grafana: Disable loading of not installed core plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [11:15:46] (03CR) 10Jelto: [V:03+1 C:03+2] profile::lists::monitoring: remove hardcoded number of uwsgi processes [puppet] - 10https://gerrit.wikimedia.org/r/1330312 (https://phabricator.wikimedia.org/T436047) (owner: 10Jelto) [11:19:06] RECOVERY - mailman3-web on lists1004 is OK: PROCS OK: 25 processes with UID = 33 (www-data), regex args /usr/bin/uwsgi https://wikitech.wikimedia.org/wiki/Mailman/Monitoring [11:32:15] (03Abandoned) 10Clément Goubert: mw-on-k8s: Use statsd exporter for fatal-error.php [puppet] - 10https://gerrit.wikimedia.org/r/1327527 (https://phabricator.wikimedia.org/T435364) (owner: 10Clément Goubert) [11:32:20] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - thanos-query_443: Servers titan1002.eqiad.wmnet are marked down but pooled: thanos-web_443: Servers titan1001.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [11:32:57] FIRING: [3x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:34:20] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [11:34:24] PROBLEM - Host titan1002 is DOWN: PING CRITICAL - Packet loss = 60%, RTA = 7837.83 ms [11:35:14] RECOVERY - Host titan1002 is UP: PING OK - Packet loss = 0%, RTA = 0.23 ms [11:37:18] (03CR) 10Clément Goubert: [C:03+1] api-gateway: Drop support for debug_hosts [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305236 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [11:37:57] RESOLVED: [3x] ProbeDown: Service titan1002:443 has failed probes (http_thanos_wikimedia_org_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:38:13] (03CR) 10Clément Goubert: [C:03+1] api-gateway: Remove stale test assertion and noop Lua code [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311874 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [11:40:07] (03CR) 10Clément Goubert: [C:03+1] api-gateway: Drop support for php_engine_routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311875 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [11:41:44] (03CR) 10Clément Goubert: api-gateway: Basic cluster specifier support and Lua plugin (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [11:45:10] FIRING: BFDdown: BFD session down between cr1-eqiad and 195.200.68.137 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:46:32] (03PS1) 10JMeybohm: Update to v1.34.11 [debs/kubernetes] (v1.34) - 10https://gerrit.wikimedia.org/r/1330382 (https://phabricator.wikimedia.org/T427069) [11:50:10] RESOLVED: [2x] BFDdown: BFD session down between cr1-eqiad and 195.200.68.137 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:55:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [11:58:39] FIRING: CoreBGPDown: Core BGP session down between cr1-magru and cr1-eqiad (195.200.68.136) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=magru&var-device=cr1-magru:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T1200) [12:03:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-magru and cr1-eqiad (195.200.68.136) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=magru&var-device=cr1-magru:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:04:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:05:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95140317 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [12:09:10] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:10:10] FIRING: BFDdown: BFD session down between cr2-eqiad and fe80::aad0:e5ff:fee3:80c3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:13:21] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:13:25] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [12:14:25] RESOLVED: [2x] BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:16:01] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:16:05] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [12:20:19] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [12:22:21] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [12:26:53] (03PS1) 10PipelineBot: mobileapps: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330403 [12:27:08] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:27:12] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [12:28:15] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'. [12:30:17] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'. [12:31:33] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:31:37] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [12:36:13] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [12:36:18] !log dpogorzelski@deploy1003 helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'llm' . [12:41:10] (03PS1) 10STran: Enable SuggestedInvestigations front-end for eswiki and jawiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330420 (https://phabricator.wikimedia.org/T435757) [12:44:31] (03CR) 10Mszwarc: [C:03+1] Enable SuggestedInvestigations front-end for eswiki and jawiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330420 (https://phabricator.wikimedia.org/T435757) (owner: 10STran) [12:45:07] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 27 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330420 (https://phabricator.wikimedia.org/T435757) (owner: 10STran) [12:46:53] (03PS1) 10Brouberol: mediawiki-dumps-legacy: update the clouddumps1002 public ssh key [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330422 [12:53:52] (03CR) 10Dpogorzelski: [C:03+1] mediawiki-dumps-legacy: update the clouddumps1002 public ssh key [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330422 (owner: 10Brouberol) [12:54:01] (03CR) 10Brouberol: [C:03+2] mediawiki-dumps-legacy: update the clouddumps1002 public ssh key [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330422 (owner: 10Brouberol) [12:55:38] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply [12:55:47] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: How many deployers does it take to do UTC afternoon backport window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T1300). [13:00:05] mfossati and Tran: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:09] o/ [13:00:12] o/ [13:00:19] I can self-deploy [13:00:39] o/ [13:01:01] mfossati: go ahead imho [13:01:06] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 27 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311058 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [13:01:10] cool [13:01:17] I can self-deploy as well but I've also got a private code change to go with mine [13:01:18] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 27 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311059 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [13:01:18] Tran: will you also self-deploy afterwards or do you need someone? [13:01:20] ok [13:01:46] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mfossati@deploy1003 using scap backport" [skins/MinervaNeue] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329672 (https://phabricator.wikimedia.org/T436141) (owner: 10LWatson) [13:02:25] * Lucas_WMDE doesn’t know how private config(?) changes work anyway [13:03:23] (03CR) 10CDanis: [C:03+1] cache::haproxy: fix logging of normalized Host header [puppet] - 10https://gerrit.wikimedia.org/r/1330314 (https://phabricator.wikimedia.org/T434766) (owner: 10Fabfur) [13:06:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [13:08:53] (03Merged) 10jenkins-bot: Minimal Minerva: change label from "comments" to "discussion" [skins/MinervaNeue] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1329672 (https://phabricator.wikimedia.org/T436141) (owner: 10LWatson) [13:09:11] !log mfossati@deploy1003 Started scap sync-world: Backport for [[gerrit:1329672|Minimal Minerva: change label from "comments" to "discussion" (T436141)]] [13:09:17] T436141: [MinMin] Label change from "Comments" to "Discussions" - https://phabricator.wikimedia.org/T436141 [13:17:11] (03PS2) 10Giuseppe Lavagetto: admin: add support for gVisor RuntimeClass handlers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) [13:19:51] (03CR) 10CI reject: [V:04-1] admin: add support for gVisor RuntimeClass handlers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [13:23:09] (03CR) 10Xcollazo: "Can we add `Hosts:` and run PPC to test the config generates what we want?" [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [13:25:39] FIRING: TransitBGPDown: Transit BGP session down between cr2-esams and Init7 (2001:1620:1000::85) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=esams&var-device=cr2-esams:9804&var-bgp_group=Transit6&var-bgp_neighbor=Init7 - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [13:25:44] (03CR) 10CDobbins: [V:03+1 C:03+2] hieradata: apply pdns v5 flag to all dns hosts [puppet] - 10https://gerrit.wikimedia.org/r/1329618 (https://phabricator.wikimedia.org/T401832) (owner: 10CDobbins) [13:27:51] (03PS2) 10CDanis: geodns: also export pooledness within each PoP [puppet] - 10https://gerrit.wikimedia.org/r/1329663 [13:27:55] (03CR) 10CDanis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329663 (owner: 10CDanis) [13:28:58] !log mfossati@deploy1003 lwatson, mfossati: Backport for [[gerrit:1329672|Minimal Minerva: change label from "comments" to "discussion" (T436141)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:29:03] T436141: [MinMin] Label change from "Comments" to "Discussions" - https://phabricator.wikimedia.org/T436141 [13:29:07] testing [13:29:52] !log mfossati@deploy1003 lwatson, mfossati: Continuing with deployment [13:30:39] RESOLVED: [2x] TransitBGPDown: Transit BGP session down between cr2-esams and Init7 (2001:1620:1000::85) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [13:32:21] (03PS3) 10Giuseppe Lavagetto: admin: add support for gVisor RuntimeClass handlers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) [13:34:46] (03CR) 10CI reject: [V:04-1] admin: add support for gVisor RuntimeClass handlers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [13:43:17] !log mfossati@deploy1003 Finished scap sync-world: Backport for [[gerrit:1329672|Minimal Minerva: change label from "comments" to "discussion" (T436141)]] (duration: 34m 06s) [13:43:22] T436141: [MinMin] Label change from "Comments" to "Discussions" - https://phabricator.wikimedia.org/T436141 [13:43:37] done, thanks for bearing with me :-) [13:44:34] going to start my private code deploy, this might take a bit as well so please ping me if something comes up. [13:46:40] Tran: ack, let me know once done :) [13:53:46] !log daniel@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-experimental: apply [13:54:59] (03PS1) 10Arendpieter: CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330446 (https://phabricator.wikimedia.org/T419684) [13:56:27] testing now [13:57:59] lgtm, continuing [14:01:00] !log daniel@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply [14:03:32] (03CR) 10Clément Goubert: [C:03+1] mediawiki: optionally use initContainers for sidecars. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325883 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [14:07:29] (03CR) 10Fabfur: "Thanks, I'll merge next Monday to be sure no unintended consequences goes overlook" [puppet] - 10https://gerrit.wikimedia.org/r/1330314 (https://phabricator.wikimedia.org/T434766) (owner: 10Fabfur) [14:08:21] private deploy component done, now deploying its associated config patch [14:08:27] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330420 (https://phabricator.wikimedia.org/T435757) (owner: 10STran) [14:09:24] (03Merged) 10jenkins-bot: Enable SuggestedInvestigations front-end for eswiki and jawiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330420 (https://phabricator.wikimedia.org/T435757) (owner: 10STran) [14:09:40] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1330420|Enable SuggestedInvestigations front-end for eswiki and jawiki (T435757)]] [14:09:45] T435757: Enable SI on ja and eswiki - https://phabricator.wikimedia.org/T435757 [14:11:50] (03CR) 10Hashar: "I'll need to adjust the commit message to mention the Puppet upgrade brings in newer version of Facter that has the facts for Trixie. That" [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [14:14:01] !log stran@deploy1003 stran: Backport for [[gerrit:1330420|Enable SuggestedInvestigations front-end for eswiki and jawiki (T435757)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:14:44] testing now [14:15:30] !log stran@deploy1003 stran: Continuing with deployment [14:15:34] lgtm, continuing [14:19:46] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1330420|Enable SuggestedInvestigations front-end for eswiki and jawiki (T435757)]] (duration: 10m 06s) [14:19:52] T435757: Enable SI on ja and eswiki - https://phabricator.wikimedia.org/T435757 [14:21:13] Krinkle: I'm done [14:21:38] (03CR) 10Hnowlan: [C:03+1] monitoring groups: add asw1-60[34]-eqsin to monitoring groups [puppet] - 10https://gerrit.wikimedia.org/r/1329653 (https://phabricator.wikimedia.org/T418439) (owner: 10Cwhite) [14:28:24] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [14:29:05] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [14:29:37] (03CR) 10JHathaway: "ready for review" [puppet] - 10https://gerrit.wikimedia.org/r/1329665 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:29:57] Tran: ack [14:30:04] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T1430) [14:31:07] (03CR) 10JHathaway: "ready for review" [puppet] - 10https://gerrit.wikimedia.org/r/1329667 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:31:42] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet [14:31:51] (03CR) 10JHathaway: "ready for review" [puppet] - 10https://gerrit.wikimedia.org/r/1329669 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:31:57] (03PS4) 10Krinkle: MathML+MathJax rollout to phase 3 (Wikibooks) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311058 (https://phabricator.wikimedia.org/T271001) [14:32:11] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace legacy facts in profile::tlsproxy::envoy spec [puppet] - 10https://gerrit.wikimedia.org/r/1329667 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:32:36] (03PS4) 10Krinkle: MathML+MathJax rollout to phase 4 (Wikisource) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311059 (https://phabricator.wikimedia.org/T271001) [14:32:49] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in profile::tlsproxy::envoy spec [puppet] - 10https://gerrit.wikimedia.org/r/1329667 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:34:45] (03PS1) 10JMeybohm: Update pause image to 3.10.1 (k8s 1.34) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1330459 (https://phabricator.wikimedia.org/T427069) [14:36:09] (03PS3) 10Hashar: rake_modules: support Debian 13 (Trixie) facts [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) [14:36:11] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet [14:37:13] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet [14:37:15] FIRING: HttpdUnreachable: httpd unavailable for deployment mw-experimental/pinkllama at eqiad - https://wikitech.wikimedia.org/wiki/Application_servers - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=257&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-experimental&var-release=pinkllama - https://alerts.wikimedia.org/?q=alertname%3DHttpdUnreachable [14:37:26] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [14:37:56] (03CR) 10Hashar: [C:04-1] "Do not merge until the `Depends-On` commit got merged:" [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [14:40:57] (03PS1) 10CDobbins: hieradata: update cp5022 IP addrs [puppet] - 10https://gerrit.wikimedia.org/r/1330461 (https://phabricator.wikimedia.org/T414411) [14:41:15] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:aux-worker-codfw [14:41:19] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet [14:41:44] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet [14:41:49] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:42:20] (03CR) 10CDobbins: [V:03+1] "PCC SUCCESS (NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9336/console" [puppet] - 10https://gerrit.wikimedia.org/r/1330461 (https://phabricator.wikimedia.org/T414411) (owner: 10CDobbins) [14:42:21] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet [14:42:25] (03PS3) 10BryanDavis: P:mediawiki::php: add 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) [14:43:01] (03CR) 10Ssingh: [C:03+1] hieradata: update cp5022 IP addrs [puppet] - 10https://gerrit.wikimedia.org/r/1330461 (https://phabricator.wikimedia.org/T414411) (owner: 10CDobbins) [14:43:34] (03PS1) 10Cathal Mooney: wmf-netbox: add function _get_rpm_probes() to build probe data [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1330462 (https://phabricator.wikimedia.org/T435855) [14:45:26] 10ops-eqiad, 06SRE, 06DC-Ops: Alert for device ps1-f4-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T436064#12261875 (10VRiley-WMF) a:03VRiley-WMF [14:46:05] 10ops-eqiad, 06SRE, 06DC-Ops: Find console port for leaf - https://phabricator.wikimedia.org/T432044#12261877 (10VRiley-WMF) a:03VRiley-WMF [14:46:21] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet [14:46:27] 10ops-eqiad, 06SRE, 06DC-Ops: Find console port for leaf - https://phabricator.wikimedia.org/T432044#12261878 (10VRiley-WMF) 05Open→03Resolved Netbox has been updated. [14:46:56] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet [14:47:15] RESOLVED: HttpdUnreachable: httpd unavailable for deployment mw-experimental/pinkllama at eqiad - https://wikitech.wikimedia.org/wiki/Application_servers - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=257&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-experimental&var-release=pinkllama - https://alerts.wikimedia.org/?q=alertname%3DHttpdUnreachable [14:48:33] (03CR) 10CDobbins: [V:03+1 C:03+2] hieradata: update cp5022 IP addrs [puppet] - 10https://gerrit.wikimedia.org/r/1330461 (https://phabricator.wikimedia.org/T414411) (owner: 10CDobbins) [14:50:16] !log elukey@cumin1003 START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet [14:50:27] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet [14:50:29] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet [14:50:36] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet [14:51:05] !log cdobbins@cumin1003 conftool action : set/pooled=no; selector: name=cp5022.* [reason: update IP addrs] [14:51:08] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet [14:52:04] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12261910 (10VRiley-WMF) p:05High→03Medium Updating prioirty. [14:52:10] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie [14:52:24] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12261920 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [14:54:51] !log elukey@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet [14:55:13] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet [14:55:15] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet [14:55:20] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet [14:55:53] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet [14:57:46] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12261952 (10ecarg) Thanks for the review @RLazarus 1. On granularity: agreed; this SLI is intentionally binary ("did the reader... [14:59:57] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet [14:59:59] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet [15:00:05] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet [15:00:26] (03CR) 10Cwhite: [C:03+2] monitoring groups: add asw1-60[34]-eqsin to monitoring groups [puppet] - 10https://gerrit.wikimedia.org/r/1329653 (https://phabricator.wikimedia.org/T418439) (owner: 10Cwhite) [15:00:38] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet [15:04:42] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet [15:04:44] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet [15:04:50] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet [15:05:27] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet [15:08:50] (03PS2) 10Cathal Mooney: wmf-netbox: add function _get_rpm_probes() to build probe data [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1330462 (https://phabricator.wikimedia.org/T435855) [15:10:59] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet [15:11:01] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet [15:11:06] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet [15:11:43] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet [15:12:18] (03CR) 10Cathal Mooney: [C:03+2] wmf-netbox: add function _get_rpm_probes() to build probe data [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1330462 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [15:12:57] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311058 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [15:12:58] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311059 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [15:13:05] (03PS1) 10Cathal Mooney: Remove YAML definitions for rpm-probes [homer/public] - 10https://gerrit.wikimedia.org/r/1330466 (https://phabricator.wikimedia.org/T435855) [15:13:52] (03Merged) 10jenkins-bot: MathML+MathJax rollout to phase 3 (Wikibooks) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311058 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [15:13:53] (03PS1) 10Cwhite: monitoring: add asw1-60[34]-eqsin to infra_devices [puppet] - 10https://gerrit.wikimedia.org/r/1330468 (https://phabricator.wikimedia.org/T418439) [15:13:55] (03Merged) 10jenkins-bot: MathML+MathJax rollout to phase 4 (Wikisource) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1311059 (https://phabricator.wikimedia.org/T271001) (owner: 10Krinkle) [15:14:09] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1311058|MathML+MathJax rollout to phase 3 (Wikibooks) (T271001)]], [[gerrit:1311059|MathML+MathJax rollout to phase 4 (Wikisource) (T271001)]] [15:14:13] T271001: Transition to client-side MathJax SVG rendering as default - https://phabricator.wikimedia.org/T271001 [15:15:53] !log cmooney@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet,cumin1003.eqiad.wmnet with reason: Homer - add function to wmf plugin to build rpm probe data - cmooney@cumin1003 [15:17:09] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet [15:17:11] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet [15:17:17] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet [15:17:30] !log cmooney@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet,cumin1003.eqiad.wmnet with reason: Homer - add function to wmf plugin to build rpm probe data - cmooney@cumin1003 [15:17:30] (03CR) 10Cathal Mooney: [C:03+1] "Thanks! Apologies for the ommission this should have been done yesterday when we migrated to the new switches in eqsin" [puppet] - 10https://gerrit.wikimedia.org/r/1330468 (https://phabricator.wikimedia.org/T418439) (owner: 10Cwhite) [15:17:35] (03CR) 10Ahmon Dancy: mediawiki::deployment::server: Automate scope=pretrain deployments (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329690 (https://phabricator.wikimedia.org/T436178) (owner: 10BryanDavis) [15:17:53] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet [15:18:08] (03CR) 10Cwhite: [C:03+2] monitoring: add asw1-60[34]-eqsin to infra_devices [puppet] - 10https://gerrit.wikimedia.org/r/1330468 (https://phabricator.wikimedia.org/T418439) (owner: 10Cwhite) [15:18:19] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1311058|MathML+MathJax rollout to phase 3 (Wikibooks) (T271001)]], [[gerrit:1311059|MathML+MathJax rollout to phase 4 (Wikisource) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:18:58] !log cmooney@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet,cumin1003.eqiad.wmnet with reason: Homer - add function to wmf plugin to build rpm probe data - cmooney@cumin1003 [15:19:20] (03PS1) 10Rscout: T434609: Deploy update to Security site [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330473 [15:20:35] !log cmooney@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet,cumin1003.eqiad.wmnet with reason: Homer - add function to wmf plugin to build rpm probe data - cmooney@cumin1003 [15:21:58] (03PS1) 10Dwisehaupt: fr-tech: shift read traffic off of frdb1008 for db work [dns] - 10https://gerrit.wikimedia.org/r/1330476 [15:23:13] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet [15:23:15] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet [15:23:20] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet [15:23:25] !log krinkle@deploy1003 krinkle: Continuing with deployment [15:24:17] !log cdobbins@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage [15:24:45] RECOVERY - Host wikikube-worker1260 is UP: PING OK - Packet loss = 0%, RTA = 0.41 ms [15:24:57] (03CR) 10Cathal Mooney: [C:03+2] Remove YAML definitions for rpm-probes [homer/public] - 10https://gerrit.wikimedia.org/r/1330466 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [15:25:36] (03CR) 10Jgreen: [C:03+2] fr-tech: shift read traffic off of frdb1008 for db work [dns] - 10https://gerrit.wikimedia.org/r/1330476 (owner: 10Dwisehaupt) [15:25:45] (03CR) 10Jgreen: [C:03+1] fr-tech: shift read traffic off of frdb1008 for db work [dns] - 10https://gerrit.wikimedia.org/r/1330476 (owner: 10Dwisehaupt) [15:26:24] (03Merged) 10jenkins-bot: Remove YAML definitions for rpm-probes [homer/public] - 10https://gerrit.wikimedia.org/r/1330466 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [15:27:29] (03CR) 10Dwisehaupt: [C:03+2] fr-tech: shift read traffic off of frdb1008 for db work [dns] - 10https://gerrit.wikimedia.org/r/1330476 (owner: 10Dwisehaupt) [15:27:45] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1311058|MathML+MathJax rollout to phase 3 (Wikibooks) (T271001)]], [[gerrit:1311059|MathML+MathJax rollout to phase 4 (Wikisource) (T271001)]] (duration: 13m 36s) [15:27:50] T271001: Transition to client-side MathJax SVG rendering as default - https://phabricator.wikimedia.org/T271001 [15:28:12] !log dwisehaupt@dns1005 START - running authdns-update [15:28:23] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet [15:28:27] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage [15:28:54] RECOVERY - Check correctness of the icinga configuration on alert1002 is OK: Icinga configuration is correct https://wikitech.wikimedia.org/wiki/Icinga [15:30:25] (03CR) 10SBassett: [C:03+1] "LGTM. Feel free to self +2 or I can whenever you're ready to deploy on miscweb." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330473 (owner: 10Rscout) [15:30:52] !log dwisehaupt@dns1005 END - running authdns-update [15:30:53] 10ops-eqiad, 06SRE, 06DC-Ops: Check list of PXE miss-configs for eqiad - https://phabricator.wikimedia.org/T401441#12262088 (10VRiley-WMF) [15:31:14] PROBLEM - Juniper alarms on asw1-603-eqsin is CRITICAL: JNX_ALARMS CRITICAL - No response from remote host 103.102.166.129 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Juniper_alarm [15:32:31] (03CR) 10Andrea Denisse: [C:03+2] grafana: Disable automatic updates for plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329416 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [15:32:41] (03CR) 10Andrea Denisse: [V:03+1 C:03+2] grafana: Disable loading of not installed core plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [15:32:59] (03PS8) 10Andrea Denisse: grafana: Disable loading of not installed core plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) [15:33:15] (03CR) 10Andrea Denisse: [V:03+2 C:03+2] grafana: Disable loading of not installed core plugins [puppet] - 10https://gerrit.wikimedia.org/r/1329517 (https://phabricator.wikimedia.org/T436045) (owner: 10Andrea Denisse) [15:33:42] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet [15:33:43] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet [15:33:44] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:aux-worker-codfw [15:35:29] (03CR) 10Aghirelli: [C:03+1] Remove mode from RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328215 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [15:36:03] (03PS1) 10Clare Ming: Remove constructive edits maintenance job [puppet] - 10https://gerrit.wikimedia.org/r/1330485 (https://phabricator.wikimedia.org/T436175) [15:36:53] (03PS1) 10Krinkle: Fix incorrect number of children in m(under|over) [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1330486 (https://phabricator.wikimedia.org/T435705) [15:37:36] (03PS1) 10Hnowlan: Add button for mgmt escalation [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330487 [15:38:09] !log eevans@cumin1003 START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [15:39:01] (03CR) 10Hnowlan: "Just an initial pass at this - not a high priority change. Probably an overall quality of life improvement compared to manually filling te" [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330487 (owner: 10Hnowlan) [15:41:13] (03PS4) 10Hnowlan: team-sre: add infrastructure-foundations tag [alerts] - 10https://gerrit.wikimedia.org/r/1319832 (https://phabricator.wikimedia.org/T432376) [15:41:14] PROBLEM - Juniper alarms on asw1-604-eqsin is CRITICAL: JNX_ALARMS CRITICAL - No response from remote host 103.102.166.136 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Juniper_alarm [15:43:27] (03CR) 10Dwisehaupt: [C:03+1] "Looks good. Thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1329669 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:43:38] (03CR) 10Rscout: [C:03+2] T434609: Deploy update to Security site [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330473 (owner: 10Rscout) [15:46:03] (03PS1) 10Ladsgroup: Switch jawiki, zhwiki and ptwiki to thumb.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330491 (https://phabricator.wikimedia.org/T427465) [15:46:03] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [15:46:17] (03Merged) 10jenkins-bot: T434609: Deploy update to Security site [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330473 (owner: 10Rscout) [15:50:04] jouncebot: nowandnext [15:50:04] No deployments scheduled for the next 0 hour(s) and 9 minute(s) [15:50:05] In 0 hour(s) and 9 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T1600) [15:50:29] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330491 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [15:50:55] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm [15:51:29] (03Merged) 10jenkins-bot: Switch jawiki, zhwiki and ptwiki to thumb.wikimedia.org [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1330491 (https://phabricator.wikimedia.org/T427465) (owner: 10Ladsgroup) [15:51:38] (03PS1) 10Cathal Mooney: eqsin tidyup: remove offline old switch asw1-eqsin following move [puppet] - 10https://gerrit.wikimedia.org/r/1330495 (https://phabricator.wikimedia.org/T418439) [15:51:42] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1330491|Switch jawiki, zhwiki and ptwiki to thumb.wikimedia.org (T427465)]] [15:51:51] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [15:54:05] (03PS3) 10Hnowlan: sre/cdn: create recording rules in advance of moving to ratio for CDN [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) [15:54:21] (03CR) 10Hnowlan: sre/cdn: create recording rules in advance of moving to ratio for CDN (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [15:54:53] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in community-civi.my.cnf.erb [puppet] - 10https://gerrit.wikimedia.org/r/1329669 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:55:57] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1330491|Switch jawiki, zhwiki and ptwiki to thumb.wikimedia.org (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:57:09] Amir1: we'll have one for the puppet window but I'm about to post a comment on it that we'll need to address first, so we won't be going right at the top of the hour -- no rush, just lmk when you're clear [15:57:22] sure thing [15:57:24] thanks! [15:58:13] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [15:58:40] rzl: thanks! what needs to be updated? [15:58:58] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cassandra-dev2002.codfw.wmnet with OS bookworm [15:59:32] (03PS2) 10Hnowlan: sre/cdn: use traffic ratios in ATSBackendErrorsHigh [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) [15:59:54] (03CR) 10Hnowlan: "Both changes updated to use this approach" [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [15:59:58] (03CR) 10RLazarus: Remove constructive edits maintenance job (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1330485 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [16:00:04] jhathaway and rzl: I, the Bot under the Fountain, call upon thee, The Deployer, to do Puppet request window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T1600). [16:00:04] cjming: A patch you scheduled for Puppet request window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [16:00:07] cjming: ah hi :) I was just typing at you in gerrit, see my comment there [16:00:14] sorry I didn't get a chance to look until just before the window [16:00:32] thanks for grabbing it rzl [16:00:32] no worries - thanks for your review [16:00:54] !log elukey@cumin1003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:aux-worker-eqiad [16:00:57] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet [16:01:00] !log eevans@cumin1003 START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [16:01:32] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet [16:01:43] (03CR) 10Clare Ming: Remove constructive edits maintenance job (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1330485 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [16:02:29] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1330491|Switch jawiki, zhwiki and ptwiki to thumb.wikimedia.org (T427465)]] (duration: 10m 47s) [16:02:34] T427465: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465 [16:03:06] cjming: yep exactly [16:03:22] rzl: pushing imminently - thanks [16:03:48] (and you'll have a few minutes to shuffle the second patch around while we're merging and deploying the first one) [16:04:08] sounds good - will do [16:04:54] (03PS1) 10Clare Ming: Tell puppet to delete constructive edits maintenance script [puppet] - 10https://gerrit.wikimedia.org/r/1330504 (https://phabricator.wikimedia.org/T436175) [16:04:57] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Q1:rack/setup/install mc20[56-73] - https://phabricator.wikimedia.org/T436272 (10RobH) 03NEW [16:05:00] (03PS1) 10Clément Goubert: mediawiki: Redirect /api/ to /w/rest.php [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330505 (https://phabricator.wikimedia.org/T433547) [16:05:30] (03CR) 10CI reject: [V:04-1] Tell puppet to delete constructive edits maintenance script [puppet] - 10https://gerrit.wikimedia.org/r/1330504 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [16:05:36] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet [16:05:37] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet [16:05:43] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet [16:06:02] cjming: ah sorry, just needs a trailing comma, and I think CI will also yell if you don't add a bunch of spaces to align the => with all the ones below [16:06:12] (03PS2) 10Aleksandar Mastilovic: Enable presto join spill config [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [16:06:18] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet [16:06:18] (03CR) 10Aleksandar Mastilovic: "I've just added them." [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [16:06:22] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Q1:rack/setup/install mc20[56-73] - https://phabricator.wikimedia.org/T436272#12262275 (10RobH) a:03Clement_Goubert @Clement_Goubert, The racking details weren't filled out on the order task T434583 so I've taken the liberty o... [16:06:23] (03CR) 10Aleksandar Mastilovic: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [16:06:48] (03PS3) 10Clément Goubert: mediawiki-vhost: Redirect /api/ to /w/rest.php [puppet] - 10https://gerrit.wikimedia.org/r/1329607 (https://phabricator.wikimedia.org/T433547) [16:07:03] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: mc20[56-73] implementation tracking - https://phabricator.wikimedia.org/T436273 (10RobH) 03NEW [16:07:24] (03PS2) 10Clare Ming: Tell puppet to delete constructive edits maintenance script [puppet] - 10https://gerrit.wikimedia.org/r/1330504 (https://phabricator.wikimedia.org/T436175) [16:07:33] 10ops-codfw, 06SRE, 06DC-Ops, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: Q1:rack/setup/install mc20[56-73] - https://phabricator.wikimedia.org/T436272#12262298 (10RobH) [16:07:40] cjming: perfect, that should pass [16:07:59] rzl: I'm done! [16:08:08] Amir1: thanks! [16:08:10] PROBLEM - MariaDB Replica IO: matomo on db1208 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl@matomo1003.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on matomo1003.eqiad.wmnet (111 Connection refused) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:08:14] ideal timing [16:08:57] i think it's a random database, not real prod [16:09:01] 10ops-eqiad, 06SRE, 06DC-Ops: Alert for device ps1-f4-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T436064#12262312 (10VRiley-WMF) 05Open→03Resolved Rebalanced power, will monitor. [16:09:03] let me triple check [16:09:19] (03PS2) 10Clare Ming: Remove constructive edits maintenance job [puppet] - 10https://gerrit.wikimedia.org/r/1330485 (https://phabricator.wikimedia.org/T436175) [16:09:48] (03CR) 10CI reject: [V:04-1] Remove constructive edits maintenance job [puppet] - 10https://gerrit.wikimedia.org/r/1330485 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [16:10:07] yeah, data engineering database (T334055) [16:10:08] T334055: Replace db1108 with db1208 - https://phabricator.wikimedia.org/T334055 [16:10:21] (03PS3) 10Clare Ming: Remove constructive edits maintenance job [puppet] - 10https://gerrit.wikimedia.org/r/1330485 (https://phabricator.wikimedia.org/T436175) [16:10:21] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet [16:10:23] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet [16:10:28] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet [16:10:49] (03CR) 10CI reject: [V:04-1] Remove constructive edits maintenance job [puppet] - 10https://gerrit.wikimedia.org/r/1330485 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [16:10:59] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet [16:11:34] (03CR) 10RLazarus: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9337/co" [puppet] - 10https://gerrit.wikimedia.org/r/1330504 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [16:12:02] PCC looks good, deploying #1 [16:12:06] (03CR) 10RLazarus: [V:03+1 C:03+2] Tell puppet to delete constructive edits maintenance script [puppet] - 10https://gerrit.wikimedia.org/r/1330504 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [16:13:16] running puppet on the deployment server, see you in seven to ten business days [16:13:30] (or, like, twelve minutes) [16:14:53] alerting host vs deployment server, fight! [16:15:02] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet [16:15:03] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet [16:15:09] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet [16:15:44] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet [16:16:12] lol [16:17:47] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [16:18:31] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [16:18:44] (03PS1) 10Clare Ming: Remove constructive edits maintenance job [puppet] - 10https://gerrit.wikimedia.org/r/1330514 (https://phabricator.wikimedia.org/T436175) [16:18:53] (03CR) 10Andrea Denisse: [C:03+1] "LGTM, thank you!!" [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330487 (owner: 10Hnowlan) [16:18:57] 10ops-eqiad, 06SRE, 10Ceph, 06cloud-services-team, and 3 others: cloudcephosd1044 boot issues - https://phabricator.wikimedia.org/T429267#12262353 (10VRiley-WMF) Will be looking at this today and hopefully we can resolve that issue. [16:19:22] and done, twelve minutes just exactly like I said [16:19:36] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-cron: apply [16:19:47] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet [16:19:48] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet [16:19:50] (03Abandoned) 10Clare Ming: Remove constructive edits maintenance job [puppet] - 10https://gerrit.wikimedia.org/r/1330485 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [16:19:55] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet [16:19:58] (03PS1) 10Blake: kubernetes: switch the default envoy version to 1.39.0 [puppet] - 10https://gerrit.wikimedia.org/r/1330515 (https://phabricator.wikimedia.org/T421418) [16:20:08] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply [16:20:13] rzl: sorry - here's the follow up patch https://gerrit.wikimedia.org/r/c/operations/puppet/+/1330514 [16:20:27] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet [16:20:42] rzl: do you provide consulting on which numbers to pick for lotto tickets as well? asking for a friend. [16:20:49] * sukhe stops trolling [16:21:24] cjming: looks good! in the meantime, I ran the helmfile apply that actually deletes the job, so it's officially gone [16:21:34] woohoo! [16:21:39] thanks for advising [16:22:16] thanks for that cleanup patch -- I'll let it sit for at least 30 minutes (which allows puppet to run on every host, making sure the absented job is actually deleted everywhere) and then merge that, but it'll be a no-op, so I'll just do it sometime today and no need for you to hang around and wait for it [16:22:42] rzl: awesome - thanks for much for your help! [16:22:47] thanks for the cleanup! [16:23:10] RECOVERY - MariaDB Replica IO: matomo on db1208 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [16:23:15] and, puppet window's complete but I think bjensen and I might start something early for the infra window, if no one else is in line [16:23:49] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm [16:24:00] (03CR) 10RLazarus: [C:03+1] kubernetes: switch the default envoy version to 1.39.0 [puppet] - 10https://gerrit.wikimedia.org/r/1330515 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [16:24:26] (03PS2) 10Cathal Mooney: eqsin tidyup: remove offline old switch asw1-eqsin following move [puppet] - 10https://gerrit.wikimedia.org/r/1330495 (https://phabricator.wikimedia.org/T418439) [16:25:45] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet [16:25:46] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet [16:25:47] (03PS1) 10Dwisehaupt: fr-tech: shift read traffic back to frdb1008 [dns] - 10https://gerrit.wikimedia.org/r/1330516 [16:25:52] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet [16:27:16] (03CR) 10Jgreen: [C:03+1] fr-tech: shift read traffic back to frdb1008 [dns] - 10https://gerrit.wikimedia.org/r/1330516 (owner: 10Dwisehaupt) [16:29:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:29:12] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [16:29:15] (03PS1) 10RLazarus: wikifunctions: Update restricted callback listener to the correct URL [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330518 (https://phabricator.wikimedia.org/T427863) [16:29:57] (03CR) 10Blake: [C:03+2] kubernetes: switch the default envoy version to 1.39.0 [puppet] - 10https://gerrit.wikimedia.org/r/1330515 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [16:30:55] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet [16:31:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs-next/scholarly-external on k8s-dse@codfw in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [16:34:10] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:35:12] (03CR) 10Jforrester: [C:03+2] wikifunctions: Update restricted callback listener to the correct URL [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330518 (https://phabricator.wikimedia.org/T427863) (owner: 10RLazarus) [16:36:14] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet [16:36:15] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet [16:36:20] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet [16:36:56] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet [16:37:29] (03Merged) 10jenkins-bot: wikifunctions: Update restricted callback listener to the correct URL [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330518 (https://phabricator.wikimedia.org/T427863) (owner: 10RLazarus) [16:38:11] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [16:38:50] (03PS4) 10BryanDavis: P:mediawiki::php: add 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) [16:38:52] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [16:39:26] (03CR) 10CI reject: [V:04-1] P:mediawiki::php: add 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) (owner: 10BryanDavis) [16:39:38] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/apertium: apply [16:39:46] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/apertium: apply [16:39:52] 10ops-eqiad, 06SRE, 10Ceph, 06cloud-services-team, and 3 others: cloudcephosd1044 boot issues - https://phabricator.wikimedia.org/T429267#12262463 (10VRiley-WMF) 05In progress→03Resolved reseated backplan and cables. After that I was able to boot it up and from iDRAC, everything looks clean. This s... [16:40:28] !log eevans@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage [16:40:40] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/chart-renderer: apply [16:40:54] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/chart-renderer: apply [16:41:08] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply [16:41:12] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply [16:41:22] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [16:41:44] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [16:41:51] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/commons-impact-analytics: apply [16:41:59] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/commons-impact-analytics: apply [16:42:14] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet [16:42:15] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet [16:42:21] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet [16:42:23] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/cxserver: apply [16:42:41] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/cxserver: apply [16:42:56] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet [16:42:59] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/data-gateway: apply [16:43:24] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage [16:44:45] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/data-gateway: apply [16:44:55] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/developer-portal: apply [16:45:05] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/developer-portal: apply [16:45:10] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/device-analytics: apply [16:45:18] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/device-analytics: apply [16:45:27] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/echostore: apply [16:46:03] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/echostore: apply [16:46:10] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/edit-analytics: apply [16:46:19] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/edit-analytics: apply [16:46:25] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/editor-analytics: apply [16:46:33] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/editor-analytics: apply [16:46:39] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-analytics: apply [16:47:03] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply [16:47:11] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-analytics-external: apply [16:47:19] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-analytics-external: apply [16:47:25] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply [16:47:29] (03CR) 10Bking: [C:03+2] Kerberos: Apply kerberos role to newly-provisioned host [puppet] - 10https://gerrit.wikimedia.org/r/1329636 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [16:47:32] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply [16:47:38] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-main: apply [16:47:46] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-main: apply [16:47:51] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/eventstreams: apply [16:48:16] !log elukey@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet [16:48:16] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet [16:48:17] !log elukey@cumin1003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:aux-worker-eqiad [16:48:20] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/eventstreams: apply [16:49:32] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/geo-analytics: apply [16:49:44] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/geo-analytics: apply [16:49:49] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/kartotherian: apply [16:49:57] (03CR) 10Dwisehaupt: [C:03+2] fr-tech: shift read traffic back to frdb1008 [dns] - 10https://gerrit.wikimedia.org/r/1330516 (owner: 10Dwisehaupt) [16:50:06] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/kartotherian: apply [16:50:11] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/linked-artifacts: apply [16:50:25] !log dwisehaupt@dns1005 START - running authdns-update [16:51:08] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply [16:51:10] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12262529 (10VRiley-WMF) @Marostegui essentially, it's making sure that everything is set the way it's supposed to for booting up properly and reliably. We normally run the scripts for fresh i... [16:51:14] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/linkrecommendation: apply [16:51:24] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply [16:51:30] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/machinetranslation: apply [16:52:48] !log dwisehaupt@dns1005 END - running authdns-update [16:53:20] (03CR) 10Ssingh: [C:03+1] "I tried a bunch of different greps and I think you have covered all asw1-eqsin references." [puppet] - 10https://gerrit.wikimedia.org/r/1330495 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [16:54:26] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/machinetranslation: apply [16:54:32] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/media-analytics: apply [16:55:28] !log cdobbins@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie [16:55:43] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie [16:55:43] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12262551 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (... [16:55:58] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12262553 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [16:56:25] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:56:45] (03CR) 10Xcollazo: [C:03+1] Enable presto join spill config [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [17:00:05] bd808: Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T1700). Please do the needful. [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T1700) [17:00:05] (03PS1) 10Cathal Mooney: wmf-netbox: fix logic error in finding the far-side IP for rpm probes [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1330540 (https://phabricator.wikimedia.org/T435855) [17:00:36] (03CR) 10Cathal Mooney: [C:03+2] eqsin tidyup: remove offline old switch asw1-eqsin following move [puppet] - 10https://gerrit.wikimedia.org/r/1330495 (https://phabricator.wikimedia.org/T418439) (owner: 10Cathal Mooney) [17:01:05] (03PS5) 10BryanDavis: P:mediawiki::php: add 8.5 [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) [17:01:29] (03CR) 10CI reject: [V:04-1] wmf-netbox: fix logic error in finding the far-side IP for rpm probes [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1330540 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [17:02:04] Nothing for my window this week [17:02:21] rzl: ready with the MW infra deploy? might do a MW backport now otherwise [17:03:15] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm [17:04:37] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/media-analytics: apply [17:05:00] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1014.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:05:02] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/media-analytics: apply [17:05:35] (03PS2) 10Cathal Mooney: wmf-netbox: fix logic error in finding the far-side IP for rpm probes [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1330540 (https://phabricator.wikimedia.org/T435855) [17:06:00] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [17:07:02] Krinkle: coordinate with bjensen, they're already rolling [17:11:16] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12262627 (10bking) I have deployed `krb1004` as a Kerberos replica as promised, the next step is to add it to `kerberos_kdc_... [17:11:25] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:15:07] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/media-analytics: apply [17:15:40] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/mobileapps: apply [17:15:48] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/mobileapps: apply [17:16:46] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/page-analytics: apply [17:16:54] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/page-analytics: apply [17:16:59] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/proton: apply [17:17:37] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/proton: apply [17:17:41] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/push-notifications: apply [17:17:49] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/push-notifications: apply [17:18:20] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/recommendation-api: apply [17:18:27] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/recommendation-api: apply [17:18:33] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/sessionstore: apply [17:18:40] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/sessionstore: apply [17:18:48] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/shellbox: apply [17:19:08] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox: apply [17:19:16] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-constraints: apply [17:19:24] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply [17:19:30] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-media: apply [17:19:39] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-media: apply [17:19:45] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply [17:19:53] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [17:19:58] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-timeline: apply [17:20:15] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply [17:20:21] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/shellbox-video: apply [17:20:42] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/shellbox-video: apply [17:21:08] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/termbox: apply [17:21:16] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/termbox: apply [17:21:25] RESOLVED: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:22:07] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [17:22:21] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [17:22:26] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/wikifeeds: apply [17:22:44] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifeeds: apply [17:22:50] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/zotero: apply [17:22:52] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12262666 (10cmooney) @VRiley-WMF I've been having it out with Lumen on this one. I still think the issue is there side, but if possible can you try to see if we have a spare 100GBase-LR... [17:22:55] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:22:57] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/zotero: apply [17:23:10] RESOLVED: SystemdUnitFailed: krb5-kdc.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:23:40] 10ops-eqiad, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install cirrussearch11[26-30] - https://phabricator.wikimedia.org/T436285 (10RobH) 03NEW [17:24:19] 10ops-eqiad, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install cirrussearch11[26-30] - https://phabricator.wikimedia.org/T436285#12262685 (10RobH) a:03bking Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/SRE/Dc-operations) and add the... [17:24:44] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12262689 (10VRiley-WMF) Okay, no problem. I'll take a look at it now. [17:25:00] PROBLEM - very high load average likely xfs on ms-be1097 is CRITICAL: LOAD CRITICAL - total load average: 200.70, 94.69, 48.45 https://wikitech.wikimedia.org/wiki/Swift [17:25:16] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/apertium: apply [17:25:50] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/apertium: apply [17:25:59] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/chart-renderer: apply [17:26:07] !log rscout@deploy1003 helmfile [codfw] START helmfile.d/services/miscweb: apply [17:26:23] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/chart-renderer: apply [17:26:30] !log rscout@deploy1003 helmfile [codfw] DONE helmfile.d/services/miscweb: apply [17:26:32] 10ops-eqiad, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install cirrussearch11[26-30] - https://phabricator.wikimedia.org/T436285#12262708 (10bking) [17:26:39] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply [17:26:41] !log rscout@deploy1003 helmfile [eqiad] START helmfile.d/services/miscweb: apply [17:26:45] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply [17:26:51] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/citoid: apply [17:26:57] !log rscout@deploy1003 helmfile [eqiad] DONE helmfile.d/services/miscweb: apply [17:27:00] RECOVERY - very high load average likely xfs on ms-be1097 is OK: LOAD OK - total load average: 52.25, 73.20, 46.02 https://wikitech.wikimedia.org/wiki/Swift [17:27:15] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/citoid: apply [17:27:20] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/commons-impact-analytics: apply [17:27:36] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/commons-impact-analytics: apply [17:27:42] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/cxserver: apply [17:27:43] 10ops-codfw, 06SRE, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install cirrussearch21[16-20] - https://phabricator.wikimedia.org/T436287 (10RobH) 03NEW [17:27:55] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:28:07] 10ops-codfw, 06SRE, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install cirrussearch21[16-20] - https://phabricator.wikimedia.org/T436287#12262733 (10RobH) a:03bking Please update the site.pp file with the insetup role for your team (detailed on https://wikitech.wikimedia.org/wiki/SRE/Dc-operations) and... [17:28:12] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/cxserver: apply [17:28:17] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/data-gateway: apply [17:28:24] 10ops-codfw, 06SRE, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install cirrussearch21[16-20] - https://phabricator.wikimedia.org/T436287#12262740 (10RobH) [17:28:31] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/data-gateway: apply [17:28:36] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/developer-portal: apply [17:28:50] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/developer-portal: apply [17:28:56] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/device-analytics: apply [17:29:09] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/device-analytics: apply [17:29:10] 10ops-eqiad, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install cirrussearch11[26-30] - https://phabricator.wikimedia.org/T436285#12262741 (10bking) [17:29:10] (03CR) 10Cathal Mooney: [C:03+2] wmf-netbox: fix logic error in finding the far-side IP for rpm probes [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1330540 (https://phabricator.wikimedia.org/T435855) (owner: 10Cathal Mooney) [17:29:15] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/echostore: apply [17:29:29] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/echostore: apply [17:29:35] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/edit-analytics: apply [17:29:48] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/edit-analytics: apply [17:29:55] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/editor-analytics: apply [17:30:03] 10ops-codfw, 06SRE, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install cirrussearch21[16-20] - https://phabricator.wikimedia.org/T436287#12262746 (10bking) [17:30:08] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/editor-analytics: apply [17:30:12] !log cmooney@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet,cumin1003.eqiad.wmnet with reason: Homer - add function to wmf plugin to build rpm probe data - cmooney@cumin1003 [17:30:15] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply [17:30:55] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply [17:30:56] PROBLEM - Check correctness of the icinga configuration on alert1002 is CRITICAL: Icinga configuration contains errors https://wikitech.wikimedia.org/wiki/Icinga [17:31:03] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: apply [17:31:30] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: apply [17:31:37] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply [17:31:48] !log cmooney@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet,cumin1003.eqiad.wmnet with reason: Homer - add function to wmf plugin to build rpm probe data - cmooney@cumin1003 [17:32:05] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply [17:32:15] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-main: apply [17:32:27] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply [17:32:34] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/eventstreams: apply [17:33:09] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventstreams: apply [17:33:14] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/geo-analytics: apply [17:33:28] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/geo-analytics: apply [17:33:34] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/kartotherian: apply [17:34:40] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/kartotherian: apply [17:34:45] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/linked-artifacts: apply [17:35:03] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply [17:35:12] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/linkrecommendation: apply [17:35:16] (03CR) 10Cathal Mooney: [C:03+1] monitoring groups: add asw1-60[34]-eqsin to monitoring groups [puppet] - 10https://gerrit.wikimedia.org/r/1329653 (https://phabricator.wikimedia.org/T418439) (owner: 10Cwhite) [17:35:33] (03CR) 10Cathal Mooney: [C:03+1] "sorry folks thanks for updating this!" [puppet] - 10https://gerrit.wikimedia.org/r/1329653 (https://phabricator.wikimedia.org/T418439) (owner: 10Cwhite) [17:36:46] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply [17:36:52] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/machinetranslation: apply [17:37:02] (03CR) 10Ssingh: "Post merge fly-by-comment:" [puppet] - 10https://gerrit.wikimedia.org/r/1329653 (https://phabricator.wikimedia.org/T418439) (owner: 10Cwhite) [17:40:33] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/machinetranslation: apply [17:40:40] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/media-analytics: apply [17:40:58] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/media-analytics: apply [17:41:22] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mobileapps: apply [17:41:34] (03PS1) 10Bking: Kerberos: Move replica to production [puppet] - 10https://gerrit.wikimedia.org/r/1330562 (https://phabricator.wikimedia.org/T435873) [17:42:03] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mobileapps: apply [17:42:26] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1330562 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [17:42:33] (03PS1) 10BCornwall: varnish: rm wmfuniq-experiment-fetcher when off [puppet] - 10https://gerrit.wikimedia.org/r/1330563 [17:43:28] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [17:43:46] (03CR) 10BCornwall: [V:03+1] "PCC SUCCESS (NOOP 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9338/console" [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [17:44:07] (03CR) 10BCornwall: varnish: rm wmfuniq-experiment-fetcher when off [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [17:44:16] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [17:44:26] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [17:44:46] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [17:44:46] (03CR) 10Bking: "The PCC command fails because `krb1004` is too new. As such, it can be safely ignored." [puppet] - 10https://gerrit.wikimedia.org/r/1330562 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [17:45:00] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/page-analytics: apply [17:45:16] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/page-analytics: apply [17:45:21] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/proton: apply [17:46:10] (03CR) 10Ssingh: [C:03+1] varnish: rm wmfuniq-experiment-fetcher when off [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [17:46:27] (03CR) 10Ssingh: varnish: rm wmfuniq-experiment-fetcher when off [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [17:46:27] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/proton: apply [17:46:32] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/push-notifications: apply [17:47:02] (03CR) 10Ssingh: varnish: rm wmfuniq-experiment-fetcher when off (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [17:47:08] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/push-notifications: apply [17:47:31] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/recommendation-api: apply [17:47:54] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/recommendation-api: apply [17:48:04] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/sessionstore: apply [17:48:18] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/sessionstore: apply [17:48:30] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox: apply [17:49:10] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox: apply [17:49:16] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply [17:49:44] (03PS2) 10BCornwall: varnish: rm wmfuniq-experiment-fetcher when off [puppet] - 10https://gerrit.wikimedia.org/r/1330563 [17:49:46] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply [17:49:57] (03CR) 10BCornwall: varnish: rm wmfuniq-experiment-fetcher when off (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [17:49:58] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-media: apply [17:50:13] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply [17:50:21] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply [17:50:42] (03CR) 10BCornwall: [V:03+1] "PCC SUCCESS (NOOP 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9340/console" [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [17:50:43] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [17:50:50] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply [17:51:12] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply [17:51:23] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/shellbox-video: apply [17:52:01] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply [17:52:22] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/termbox: apply [17:53:05] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/termbox: apply [17:53:23] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [17:53:42] apologies, this is going to run over the infra window a bit, still need to deploy to eqiad [17:53:43] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [17:53:48] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/wikifeeds: apply [17:54:12] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply [17:55:08] !log blake@deploy1003 helmfile [codfw] START helmfile.d/services/zotero: apply [17:55:32] !log blake@deploy1003 helmfile [codfw] DONE helmfile.d/services/zotero: apply [17:55:50] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/apertium: apply [17:56:23] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/apertium: apply [17:56:30] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/chart-renderer: apply [17:56:57] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/chart-renderer: apply [17:57:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [17:57:16] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply [17:57:20] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply [17:57:25] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/citoid: apply [17:57:48] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/citoid: apply [17:57:56] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/commons-impact-analytics: apply [17:58:11] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/commons-impact-analytics: apply [17:58:17] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/cxserver: apply [17:58:48] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/cxserver: apply [17:58:56] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/data-gateway: apply [17:59:10] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/data-gateway: apply [17:59:16] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/developer-portal: apply [17:59:30] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply [17:59:34] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/device-analytics: apply [17:59:48] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/device-analytics: apply [17:59:55] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/echostore: apply [18:00:08] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/echostore: apply [18:00:19] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/edit-analytics: apply [18:00:21] (03CR) 10Ssingh: [C:03+1] varnish: rm wmfuniq-experiment-fetcher when off [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [18:00:33] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/edit-analytics: apply [18:00:43] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/editor-analytics: apply [18:00:57] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/editor-analytics: apply [18:01:04] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply [18:01:41] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply [18:01:54] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: apply [18:02:10] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:02:20] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: apply [18:02:35] (03CR) 10BCornwall: [V:03+1 C:03+2] varnish: rm wmfuniq-experiment-fetcher when off [puppet] - 10https://gerrit.wikimedia.org/r/1330563 (owner: 10BCornwall) [18:02:36] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply [18:03:00] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply [18:03:07] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-main: apply [18:03:42] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply [18:04:31] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/eventstreams: apply [18:04:55] FIRING: SystemdUnitFailed: replicate-krb-database.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:05:13] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply [18:05:17] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/geo-analytics: apply [18:05:31] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/geo-analytics: apply [18:05:35] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/kartotherian: apply [18:09:22] PROBLEM - Check unit status of replicate-krb-database on krb2002 is CRITICAL: CRITICAL: Status of the systemd unit replicate-krb-database https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [18:10:06] (03CR) 10Ladsgroup: "Joal has given his +1 and is asking this to be merged and deployed first and then he'll merge that patch." [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) (owner: 10Ladsgroup) [18:10:23] !log cdobbins@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie [18:10:39] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12262855 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (... [18:13:10] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:15:45] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply [18:15:56] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply [18:16:09] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply [18:16:14] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply [18:16:33] jouncebot now [18:16:33] No deployments scheduled for the next 1 hour(s) and 43 minute(s) [18:16:45] I shall do some scap testing. [18:17:02] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply [18:17:07] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/machinetranslation: apply [18:17:43] !log dancy@deploy1003 Started scap sync-world: testing [18:21:27] !log dancy@deploy1003 Finished scap sync-world: testing (duration: 03m 44s) [18:22:46] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/machinetranslation: apply [18:23:01] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/media-analytics: apply [18:23:14] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/media-analytics: apply [18:23:31] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mobileapps: apply [18:24:11] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply [18:24:30] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-experimental: apply [18:25:01] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply [18:25:30] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [18:25:53] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [18:25:54] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:26:06] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/page-analytics: apply [18:26:08] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:26:18] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/page-analytics: apply [18:26:19] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:26:22] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/proton: apply [18:26:24] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:27:26] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/proton: apply [18:27:27] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:27:30] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/push-notifications: apply [18:27:32] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:28:03] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply [18:28:04] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:28:31] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/recommendation-api: apply [18:28:32] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:28:54] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/recommendation-api: apply [18:28:55] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:28:59] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/sessionstore: apply [18:29:01] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:29:12] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/sessionstore: apply [18:29:14] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:29:20] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox: apply [18:29:22] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:30:02] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox: apply [18:30:05] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:30:20] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply [18:30:21] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:30:28] bjensen: hm, curious what's up with stashbot there [18:30:47] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply [18:30:49] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:30:50] me too, but i'm lacking context on how to introspect [18:30:50] Somebody should check the error logs. [18:31:11] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-media: apply [18:31:13] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:31:25] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply [18:31:27] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:31:33] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply [18:31:35] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:31:51] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply [18:31:52] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:31:56] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply [18:31:58] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:32:04] we've occasionally seen that when wikitech isn't accepting edits, so that was my first worry, but just double-checked and that's not the case [18:32:12] > mwclient.errors.APIError: ('contenttoobig', 'The text you have submitted is 2,048.046 kilobytes long, which is more than the maximum of 2,048 kilobytes.', 'See https://wikitech.wikimedia.org/w/api.php for API usage. Subscribe to the mediawiki-api-announce mailing list at [18:32:12] <https://lists.wikimedia.org/postorius/lists/mediawiki-api-announce.lists.wikimedia.org/> for notice of API deprecations and breaking changes.') [18:32:18] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply [18:32:19] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:32:32] I think it's saying that https://wikitech.wikimedia.org/wiki/Server_Admin_Log as a page is too big [18:32:41] oh, that's fun [18:32:47] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/shellbox-video: apply [18:32:49] blake@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [18:33:11] I guess we've done too much server adminning this month, and we should stop until September [18:33:19] is there a process around that? archiving older logs or so? [18:33:51] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply [18:33:52] I was just starting to look into that, dunno! it's always Just Worked, I've never had to think about it [18:34:06] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/termbox: apply [18:34:23] oh, thanks taavi [18:34:39] how has no-one automated this yet [18:34:43] (the number of things around here that have always Just Worked and then it turns out to be taavi)++ [18:34:46] i am disappointed in all of you (collective) [18:34:50] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/termbox: apply [18:35:26] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [18:36:06] aha, this is T378369 [18:36:06] T378369: Find or invent a method for archiving SAL page messages - https://phabricator.wikimedia.org/T378369 [18:36:07] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [18:36:11] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifeeds: apply [18:36:34] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply [18:36:45] man yeah, I assumed that was already happening automatically and that's why the page started at 2026-07-01. I guess now [18:36:47] *guess not [18:36:52] !log blake@deploy1003 helmfile [eqiad] START helmfile.d/services/zotero: apply [18:37:15] !log blake@deploy1003 helmfile [eqiad] DONE helmfile.d/services/zotero: apply [18:38:11] (03CR) 10Atsuko: [C:03+1] Kerberos: Move replica to production [puppet] - 10https://gerrit.wikimedia.org/r/1330562 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [18:40:01] (03CR) 10Bking: [C:03+2] Kerberos: Move replica to production [puppet] - 10https://gerrit.wikimedia.org/r/1330562 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [18:40:57] 06SRE, 10Stashbot: Find or invent a method for archiving SAL page messages - https://phabricator.wikimedia.org/T378369#12262908 (10taavi) I just had to archive the page manually since it had grown to the maximum MW page size. We are supposed to be good at automating boring repetitive work, so tagging #SRE in c... [18:42:10] FIRING: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:47:10] RESOLVED: BFDdown: BFD session down between cr2-codfw and fe80::8618:88ff:fe0d:d9a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [18:49:56] 10ops-codfw, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Noisy alerts: investigate fstrim.service failure on dse-k8s-wdqs2001 (new Supermicro host) - https://phabricator.wikimedia.org/T434299#12262922 (10bking) 05Open→03In progress p:05Triage→03Low a:03bking [18:53:12] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12262943 (10VRiley-WMF) I just swapped out the optic. Looking at grafana now... [19:06:04] 10ops-codfw, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Noisy alerts: investigate fstrim.service failure on dse-k8s-wdqs2001 (new Supermicro host) - https://phabricator.wikimedia.org/T434299#12262973 (10bking) 05In progress→03Resolved Per IRC conversation with @Jhancock.wm , I... [19:09:22] RECOVERY - Check unit status of replicate-krb-database on krb2002 is OK: OK: Status of the systemd unit replicate-krb-database https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:10:55] RESOLVED: SystemdUnitFailed: replicate-krb-database.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:15:09] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12262990 (10bking) Since deploying the replica, I've noticed several flapping systemd unit failures for both primary and rep... [19:16:06] (03PS1) 10Bking: Revert "Kerberos: Move replica to production" [puppet] - 10https://gerrit.wikimedia.org/r/1330603 [19:16:21] (03CR) 10Bking: [V:03+2 C:03+2] Revert "Kerberos: Move replica to production" [puppet] - 10https://gerrit.wikimedia.org/r/1330603 (owner: 10Bking) [19:20:20] !log dancy@deploy1003 Installing scap version "4.286.0" for 3 host(s) [19:21:46] (03PS1) 10Cwhite: add curator example passwords for pcc [labs/private] - 10https://gerrit.wikimedia.org/r/1330608 (https://phabricator.wikimedia.org/T350516) [19:22:18] !log dancy@deploy1003 Installation of scap version "4.286.0" completed for 3 hosts [19:22:49] !log dancy@deploy1003 Started scap sync-world: testing [19:23:17] !log cdobbins@cumin1003 START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie [19:23:21] !log cdobbins@cumin1003 START - Cookbook sre.hosts.move-vlan for host cp5022 [19:23:21] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cp5022 [19:23:36] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12262995 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [19:28:16] !log dancy@deploy1003 Finished scap sync-world: testing (duration: 05m 45s) [19:29:16] (03CR) 10Cwhite: [V:03+2 C:03+2] add curator example passwords for pcc [labs/private] - 10https://gerrit.wikimedia.org/r/1330608 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [19:32:44] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12263011 (10Cyberpower678) I’m told it’s seeing enwiki as dead again [19:34:01] !log [bking@ganeti1046] ~$ sudo gnt-instance migrate krb1004 ganeti1035 -> ganeti1029 T435873 [19:34:04] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:34:05] T435873: Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873 [19:34:22] !log dancy@deploy1003 Started scap sync-world: testing [19:35:34] !log dancy@deploy1003 sync-world failed: dictionary changed size during iteration (scap version: 4.286.0) (duration: 01m 12s) [19:36:10] !log dancy@deploy1003 Started scap sync-world: testing [19:36:21] ooh interesting [19:36:44] (03CR) 10RLazarus: [C:03+2] Remove constructive edits maintenance job [puppet] - 10https://gerrit.wikimedia.org/r/1330514 (https://phabricator.wikimedia.org/T436175) (owner: 10Clare Ming) [19:39:10] !log dancy@deploy1003 Finished scap sync-world: testing (duration: 03m 00s) [19:41:07] !log [bking@ganeti1046] ~$ sudo gnt-instance replace-disks -n ganeti1057.eqiad.wmnet krb1004.eqiad.wmnet T435873 [19:41:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:41:10] T435873: Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873 [20:00:04] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298 (10JAATPH) 03NEW [20:00:04] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: Time to do the UTC late backport window deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T2000). [20:00:05] No Gerrit patches in the queue for this window AFAICS. [20:06:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:10:56] RECOVERY - Check correctness of the icinga configuration on alert1002 is OK: Icinga configuration is correct https://wikitech.wikimedia.org/wiki/Icinga [20:11:30] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12263206 (10VRiley-WMF) Pinged Moritz on this on IRC [20:11:31] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, August 27 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-i" [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1330486 (https://phabricator.wikimedia.org/T435705) (owner: 10Krinkle) [20:12:01] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1330486 (https://phabricator.wikimedia.org/T435705) (owner: 10Krinkle) [20:13:44] (03Merged) 10jenkins-bot: Fix incorrect number of children in m(under|over) [extensions/Math] (wmf/1.47.0-wmf.17) - 10https://gerrit.wikimedia.org/r/1330486 (https://phabricator.wikimedia.org/T435705) (owner: 10Krinkle) [20:13:59] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1330486|Fix incorrect number of children in m(under|over) (T435705)]] [20:14:02] T435705: \overset{} yield syntax errors in clientside SVG mode - https://phabricator.wikimedia.org/T435705 [20:15:46] (03PS3) 10Krinkle: mediawiki: Ignore autocreate errors in backfillLocalAccounts.php [puppet] - 10https://gerrit.wikimedia.org/r/1328372 (https://phabricator.wikimedia.org/T432613) (owner: 10Gergő Tisza) [20:16:45] (03CR) 10Krinkle: [C:03+1] "LGTM. The MW changs ridden the train meanwhile to all wikis so this should be good to go. You'll need to ask an SRE or try a puppet backpo" [puppet] - 10https://gerrit.wikimedia.org/r/1328372 (https://phabricator.wikimedia.org/T432613) (owner: 10Gergő Tisza) [20:18:10] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1330486|Fix incorrect number of children in m(under|over) (T435705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:22:24] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [20:25:58] !log krinkle@deploy1003 krinkle: Continuing with deployment [20:27:46] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12263299 (10cmooney) @VRiley-WMF thanks, light levels are ok, doesn't seem to be any change in the link flapping but I wasn't expecting it really. Please put the optic that came out in... [20:29:52] !log lerickson@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [20:30:27] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for EChukwukere-WMF - https://phabricator.wikimedia.org/T428827#12263316 (10Sfaci) Your experiment [[ https://growthbook-next.wikimedia.org/experiment/exp_1d70nmtbq0cki | valencia ]] failed validation and its status has been rev... [20:31:00] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1330486|Fix incorrect number of children in m(under|over) (T435705)]] (duration: 17m 00s) [20:31:02] T435705: \overset{} yield syntax errors in clientside SVG mode - https://phabricator.wikimedia.org/T435705 [20:35:19] 10ops-eqiad, 06SRE, 10Ceph, 06cloud-services-team, and 3 others: cloudcephosd1044 boot issues - https://phabricator.wikimedia.org/T429267#12263348 (10Andrew) Looks good -- I've put it back in service. Thank you! [20:36:44] (03PS5) 10Andrea Denisse: WIP [DNM] Add SQLite-backed incident store and analytics CLI [software/klaxon] - 10https://gerrit.wikimedia.org/r/1274026 (https://phabricator.wikimedia.org/T431101) (owner: 10CDanis) [20:37:05] !log cdobbins@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie [20:37:20] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12263359 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (... [20:41:00] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for EChukwukere-WMF - https://phabricator.wikimedia.org/T428827#12263377 (10EChukwukere-WMF) so if I set it to a date btw Mon - thurs I should be good then ? [20:41:16] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12263382 (10bking) Since `krb1004` was originally by hosted on`ganeti1035` with `ganeti1029` as secondary, I failed over `krb1004` to `ganeti1029`... [20:41:26] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12263383 (10bking) [20:41:40] FIRING: KubernetesRsyslogDown: rsyslog on wikikube-worker1049:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1049 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [20:41:58] (03PS1) 10Jforrester: wikifunctions: Update restricted callback listener to have a trailing slash [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330635 (https://phabricator.wikimedia.org/T427863) [20:43:15] (03PS2) 10Jforrester: wikifunctions: Update restricted callback listener to have a trailing slash [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330635 (https://phabricator.wikimedia.org/T427863) [20:43:25] (03CR) 10RLazarus: [C:03+1] wikifunctions: Update restricted callback listener to have a trailing slash [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330635 (https://phabricator.wikimedia.org/T427863) (owner: 10Jforrester) [20:43:39] (03CR) 10Jforrester: [C:03+2] wikifunctions: Update restricted callback listener to have a trailing slash [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330635 (https://phabricator.wikimedia.org/T427863) (owner: 10Jforrester) [20:46:24] (03Merged) 10jenkins-bot: wikifunctions: Update restricted callback listener to have a trailing slash [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330635 (https://phabricator.wikimedia.org/T427863) (owner: 10Jforrester) [20:47:07] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [20:47:22] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [20:54:21] (03PS6) 10Cwhite: opensearch: remove sensitive type from parameter [puppet] - 10https://gerrit.wikimedia.org/r/1329393 (https://phabricator.wikimedia.org/T350516) [20:56:27] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12263461 (10cmooney) [20:56:40] RESOLVED: KubernetesRsyslogDown: rsyslog on wikikube-worker1049:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=wikikube-worker1049 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [20:56:40] Hey all - I’d like to deploy a couple of security patches now. Let me know if I should hold off, thanks. [20:57:42] 10ops-eqsin, 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12263463 (10cmooney) 05Open→03Resolved I'm gonna close this one, everything went smoothly on the day. We did not get time to factory... [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260827T2100) [21:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [21:09:17] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for EChukwukere-WMF - https://phabricator.wikimedia.org/T428827#12263536 (10Sfaci) >>! In T428827#12263316, @Sfaci wrote: > Your experiment [[ https://growthbook-next.wikimedia.org/experiment/exp_1d70nmtbq0cki | valencia ]] fail... [21:10:34] (03CR) 10BryanDavis: "Raine: this is cherry-picked to the deployment-prep puppetserver and applied on servers with and without `8.5` in `profile::mediawiki::php" [puppet] - 10https://gerrit.wikimedia.org/r/1328728 (https://phabricator.wikimedia.org/T435393) (owner: 10BryanDavis) [21:11:36] (03PS3) 10Andrea Denisse: Add support to store SplunkOnCall users by email [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330630 (https://phabricator.wikimedia.org/T434729) [21:15:30] (03PS1) 10Bking: dse-k8s-eqiad: Add Matomo namespaces [puppet] - 10https://gerrit.wikimedia.org/r/1330644 (https://phabricator.wikimedia.org/T436258) [21:15:44] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1330644 (https://phabricator.wikimedia.org/T436258) (owner: 10Bking) [21:30:07] (03CR) 10RLazarus: "I like it in principle! Just questions on the config. Sorry to have so many nits on a small patch." [puppet] - 10https://gerrit.wikimedia.org/r/1329607 (https://phabricator.wikimedia.org/T433547) (owner: 10Clément Goubert) [21:31:30] (03PS1) 10Bking: dse-k8s-eqiad: Add Matomo namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330653 (https://phabricator.wikimedia.org/T436258) [21:32:47] (03CR) 10RLazarus: "Same comments on the config as I7bdd93cc. I know this is the effective copy, I just started there because it's a chance to review the conf" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330505 (https://phabricator.wikimedia.org/T433547) (owner: 10Clément Goubert) [21:32:54] (03PS1) 10BCornwall: ipip: Don't touch interfaces file if nonexistent [puppet] - 10https://gerrit.wikimedia.org/r/1330654 [21:33:31] (03CR) 10CI reject: [V:04-1] ipip: Don't touch interfaces file if nonexistent [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [21:39:48] (03PS2) 10BCornwall: ipip: Don't touch interfaces file if nonexistent [puppet] - 10https://gerrit.wikimedia.org/r/1330654 [21:40:54] !log Deployed security mitigations for T432713, T435026 [21:40:55] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:43:04] (03CR) 10BCornwall: [V:03+1] "PCC SUCCESS (NOOP 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9347/console" [puppet] - 10https://gerrit.wikimedia.org/r/1330654 (owner: 10BCornwall) [22:11:08] !log jasmine@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1261.eqiad.wmnet [22:11:11] !log jasmine@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1261.eqiad.wmnet [22:11:44] !log jasmine@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1261.eqiad.wmnet [22:12:08] !log jasmine@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1261.eqiad.wmnet with OS trixie [22:12:35] !log jasmine@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1261 [22:13:52] !log jasmine@cumin1003 START - Cookbook sre.dns.netbox [22:19:10] !log jasmine@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1261 - jasmine@cumin1003" [22:19:14] !log jasmine@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1261 - jasmine@cumin1003" [22:19:14] !log jasmine@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [22:19:14] !log jasmine@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1261.eqiad.wmnet 71.32.64.10.in-addr.arpa 1.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:19:17] !log jasmine@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1261.eqiad.wmnet 71.32.64.10.in-addr.arpa 1.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:19:18] !log jasmine@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1261 [22:20:39] !log jasmine@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1261 [22:20:39] !log jasmine@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1261 [22:25:29] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, 06MediaWiki-Core-Platform-Team (Kanban): Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12263782 (10Krinkle) Okay, so from `WikiFilePage::doPurge > LocalFile::purge... [22:35:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:40:51] !log jasmine@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1261.eqiad.wmnet with reason: host reimage [22:44:20] !log jasmine@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1261.eqiad.wmnet with reason: host reimage [22:45:42] RESOLVED: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:48:04] (03PS2) 10Cwhite: opensearch_dashboards: remove certificate authority details [puppet] - 10https://gerrit.wikimedia.org/r/1329379 (https://phabricator.wikimedia.org/T350516) [22:48:04] (03PS1) 10Cwhite: lookup_options: convert dashboards opensearch_api_password to sensitive [puppet] - 10https://gerrit.wikimedia.org/r/1330686 (https://phabricator.wikimedia.org/T350516) [22:48:06] (03PS1) 10Cwhite: beta-logs: install CA root [puppet] - 10https://gerrit.wikimedia.org/r/1330687 (https://phabricator.wikimedia.org/T350516) [22:48:09] (03PS1) 10Cwhite: beta-logs: add logstash_real role [puppet] - 10https://gerrit.wikimedia.org/r/1330688 (https://phabricator.wikimedia.org/T350516) [22:50:15] (03CR) 10Cwhite: [C:03+2] lookup_options: convert dashboards opensearch_api_password to sensitive [puppet] - 10https://gerrit.wikimedia.org/r/1330686 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [22:51:07] (03PS6) 10Scott French: P:services_proxy::envoy: Drop support for split and introduce splits [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) [22:51:08] (03PS10) 10Scott French: P:kubernetes::deployment_server::global_config: Update services_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1328256 (https://phabricator.wikimedia.org/T427666) [22:51:08] (03PS1) 10Scott French: P:services_proxy::envoy: Fix sni_rewrites_host_header handling [puppet] - 10https://gerrit.wikimedia.org/r/1330689 (https://phabricator.wikimedia.org/T427666) [22:53:32] (03PS2) 10Cwhite: beta-logs: add logstash_real role [puppet] - 10https://gerrit.wikimedia.org/r/1330688 (https://phabricator.wikimedia.org/T350516) [22:57:28] (03PS2) 10Scott French: P:services_proxy::envoy: Fix sni_rewrites_host_header handling [puppet] - 10https://gerrit.wikimedia.org/r/1330689 (https://phabricator.wikimedia.org/T427666) [22:57:28] (03PS7) 10Scott French: P:services_proxy::envoy: Drop support for split and introduce splits [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) [22:57:28] (03PS11) 10Scott French: P:kubernetes::deployment_server::global_config: Update services_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1328256 (https://phabricator.wikimedia.org/T427666) [22:57:39] (03CR) 10Scott French: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1330689 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [23:05:46] !log jasmine@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1261.eqiad.wmnet with OS trixie [23:06:02] 10SRE-SLO, 06Abstract Wikipedia team (27Q1 (Jul–Sep)), 07OKR-Work: new SLI (1 of 2): server-side metrics on Abstract Wikipedia preview - https://phabricator.wikimedia.org/T434231#12263826 (10RLazarus) Okay sounds good! For next steps, two things: - Documenting the SLO, for humans to understand what we've ag... [23:07:13] !log homer lsw1-c6-eqiad* commit 'T421711' [23:07:15] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [23:07:16] T421711: ServiceOps: Re-IP eqiad private baremetal hosts to new per-rack vlans/subnets - https://phabricator.wikimedia.org/T421711 [23:14:52] !log jasmine@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1261.eqiad.wmnet [23:14:53] !log jasmine@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1261.eqiad.wmnet [23:14:54] !log jasmine@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1261.eqiad.wmnet [23:35:41] (03CR) 10Scott French: "Thanks for the review, Reuven!" [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [23:39:05] (03CR) 10Scott French: "Thanks for the review! And good call not looking at I236e1661 yet - that's the "low pain" version that tries to sneak this into patch vers" [puppet] - 10https://gerrit.wikimedia.org/r/1328256 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [23:41:19] (03PS3) 10Cwhite: opensearch: expose a chained certificate [puppet] - 10https://gerrit.wikimedia.org/r/1329379 (https://phabricator.wikimedia.org/T350516) [23:41:19] (03PS3) 10Cwhite: beta-logs: add logstash_real role [puppet] - 10https://gerrit.wikimedia.org/r/1330688 (https://phabricator.wikimedia.org/T350516) [23:41:37] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1330705 [23:41:37] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1330705 (owner: 10TrainBranchBot) [23:42:14] (03CR) 10CI reject: [V:04-1] opensearch: expose a chained certificate [puppet] - 10https://gerrit.wikimedia.org/r/1329379 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:43:14] (03PS4) 10Cwhite: opensearch: expose a chained certificate [puppet] - 10https://gerrit.wikimedia.org/r/1329379 (https://phabricator.wikimedia.org/T350516) [23:49:11] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1330705 (owner: 10TrainBranchBot) [23:57:52] (03PS1) 10Ladsgroup: Revert^2 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330710 [23:58:02] (03CR) 10CI reject: [V:04-1] Revert^2 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330710 (owner: 10Ladsgroup)