[01:02:14] (03CR) 10BCornwall: [C:03+1] P:tofurkey enable Tofurkey for MAGRU [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [01:03:23] (03CR) 10BCornwall: [V:04-1] "Yeah, it was manual." [puppet] - 10https://gerrit.wikimedia.org/r/1333244 (https://phabricator.wikimedia.org/T433672) (owner: 10CDobbins) [01:10:50] (03CR) 10BCornwall: "The old stuff probably should stick around forever, but maybe we can move it all to the bottom and surrounded by comments saying not to ev" [dns] - 10https://gerrit.wikimedia.org/r/1338267 (https://phabricator.wikimedia.org/T437403) (owner: 10Dzahn) [01:11:06] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1339290 [01:11:06] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1339290 (owner: 10TrainBranchBot) [01:13:11] (03CR) 10Samwilson: [C:03+2] "Looks good, works in my local tests." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1338951 (https://phabricator.wikimedia.org/T290345) (owner: 10TheDJ) [01:16:46] (03Merged) 10jenkins-bot: WebP: apply coalesce before resizing [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1338951 (https://phabricator.wikimedia.org/T290345) (owner: 10TheDJ) [01:21:40] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1339290 (owner: 10TrainBranchBot) [01:31:16] (03CR) 10Andrew Bogott: "This is a little weird... the capi worker runs in the cloud realm, but it serves as the backend for a regular prod-realm openstack API so " [puppet] - 10https://gerrit.wikimedia.org/r/1339182 (https://phabricator.wikimedia.org/T429557) (owner: 10Andrew Bogott) [01:43:58] (03Abandoned) 10Andrea Denisse: Create views to correlate shift data [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330790 (https://phabricator.wikimedia.org/T434730) (owner: 10Andrea Denisse) [02:00:04] Deploy window Automatic deployment of MediaWiki to pretrain wikis - see mw:Pretrain (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260911T0200) [02:00:57] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:01:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:08:48] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 50s) [02:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [03:13:05] (03CR) 10Samwilson: [C:03+1] "Looks good to me, but I've not tested it. @hnowlan@wikimedia.org you're probably more familiar with this than me!" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337312 (https://phabricator.wikimedia.org/T431767) (owner: 10Ladsgroup-claude) [03:29:16] (03Abandoned) 10Andrea Denisse: Add support to store SplunkOnCall users by email [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330630 (https://phabricator.wikimedia.org/T434729) (owner: 10Andrea Denisse) [04:49:10] FIRING: BFDdown: BFD session down between cr1-codfw and 208.80.153.220 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:54:10] RESOLVED: BFDdown: BFD session down between cr1-codfw and 208.80.153.220 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:01:49] (03PS1) 10Tim Starling: Revert Lua 5.4 support patches [extensions/Scribunto] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1339441 (https://phabricator.wikimedia.org/T437056) [05:08:08] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [extensions/Scribunto] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1339441 (https://phabricator.wikimedia.org/T437056) (owner: 10Tim Starling) [05:15:32] (03Merged) 10jenkins-bot: Revert Lua 5.4 support patches [extensions/Scribunto] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1339441 (https://phabricator.wikimedia.org/T437056) (owner: 10Tim Starling) [05:15:54] !log tstarling@deploy1003 Started scap sync-world: Backport for [[gerrit:1339441|Revert Lua 5.4 support patches (T437056)]] [05:15:57] T437056: Increased "Lua error: not enough memory" on some wikis since 1.47.0-wmf.18 - https://phabricator.wikimedia.org/T437056 [05:17:20] (03CR) 10Andrea Denisse: sre.opensearch.roll-restart-reboot: include checklist items (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [05:20:18] !log tstarling@deploy1003 tstarling: Backport for [[gerrit:1339441|Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [05:21:23] !log tstarling@deploy1003 tstarling: Continuing with deployment [05:25:53] !log tstarling@deploy1003 Finished scap sync-world: Backport for [[gerrit:1339441|Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s) [05:25:57] T437056: Increased "Lua error: not enough memory" on some wikis since 1.47.0-wmf.18 - https://phabricator.wikimedia.org/T437056 [05:27:36] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning [05:29:25] !log Start cloning db1245:s5 T437563 [05:29:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:29:29] T437563: Put db1245 backup source back in service - https://phabricator.wikimedia.org/T437563 [05:29:46] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12310826 (10Marostegui) The host is back now [05:30:38] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning [05:31:20] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one [05:31:34] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one [05:36:24] (03PS1) 10Marostegui: mariadb: Remove note from x4 [puppet] - 10https://gerrit.wikimedia.org/r/1339468 [05:37:04] (03CR) 10Marostegui: "This is a noop." [puppet] - 10https://gerrit.wikimedia.org/r/1339468 (owner: 10Marostegui) [05:37:08] (03CR) 10Marostegui: [C:03+2] mariadb: Remove note from x4 [puppet] - 10https://gerrit.wikimedia.org/r/1339468 (owner: 10Marostegui) [05:57:39] FIRING: CoreBGPDown: Core BGP session down between cr1-codfw and cr1-eqiad (208.80.153.220) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260911T0600) [06:01:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:02:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-codfw and cr1-eqiad (208.80.153.220) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr1-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [06:13:39] FIRING: TransitBGPDown: Transit BGP session down between cr2-codfw and Lumen (64.156.73.169) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Transit4&var-bgp_neighbor=Lumen - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [06:14:52] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-eqiad:et-1/1/2 (Transport: cr1-codfw:et-1/0/2 (Arelion, IC-374549)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [06:18:39] FIRING: [2x] TransitBGPDown: Transit BGP session down between cr2-codfw and Lumen (2001:1900:2100::4b41) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [06:19:52] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr2-codfw:xe-0/0/1:3 (Transit: Lumen (442550279)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [06:23:18] (03PS1) 10Marostegui: common.yaml: Add arbcom_arwiki [puppet] - 10https://gerrit.wikimedia.org/r/1339509 (https://phabricator.wikimedia.org/T437584) [06:24:23] (03CR) 10Ryan Kemper: [C:04-1] "Probably a fat finger, but I don't see the new alert definitions. Right now it would remove 2 alerts and add a new one" [alerts] - 10https://gerrit.wikimedia.org/r/1339209 (https://phabricator.wikimedia.org/T436968) (owner: 10Bking) [06:24:37] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1338236 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [06:24:49] (03CR) 10Ryan Kemper: [C:04-1] "erm, sorry it just removes the 2* my eyes tricked me into thinking I was seeing green lines" [alerts] - 10https://gerrit.wikimedia.org/r/1339209 (https://phabricator.wikimedia.org/T436968) (owner: 10Bking) [06:27:33] !log for T437056: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category [06:27:36] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:27:37] T437056: Increased "Lua error: not enough memory" on some wikis since 1.47.0-wmf.18 - https://phabricator.wikimedia.org/T437056 [06:38:00] !log also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether T437056 [06:38:03] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:38:03] T437056: Increased "Lua error: not enough memory" on some wikis since 1.47.0-wmf.18 - https://phabricator.wikimedia.org/T437056 [07:00:05] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260911T0700) [07:03:39] RESOLVED: [2x] TransitBGPDown: Transit BGP session down between cr2-codfw and Lumen (2001:1900:2100::4b41) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [07:10:20] !log killed jobs for T437056 since they weren't purging [07:10:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:10:23] T437056: Increased "Lua error: not enough memory" on some wikis since 1.47.0-wmf.18 - https://phabricator.wikimedia.org/T437056 [07:15:15] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4 [07:16:11] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1159: Repooling db1159 [07:16:34] (03PS1) 10Milazg: Add URLs for WikimediaCustomizations MobileAppRedirect Special Page [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339561 (https://phabricator.wikimedia.org/T434930) [07:19:21] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one [07:20:08] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one [07:27:10] !log arnaudb@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [07:28:26] !log arnaudb@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab [07:29:04] !log arnaudb@cumin1003 END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [07:29:28] !log arnaudb@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [07:29:53] !log arnaudb@cumin1003 END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [07:44:34] arnaudb@cumin1003 arnaudb: The backup on gitlab2002 is complete, ready to proceed with upgrade. [07:54:08] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab [07:58:27] !log arnaudb@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [07:58:35] !log arnaudb@cumin1003 END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [08:01:17] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159 [08:11:49] (03PS1) 10Marostegui: x4: Add candidate masters [puppet] - 10https://gerrit.wikimedia.org/r/1339594 (https://phabricator.wikimedia.org/T437678) [08:13:22] (03CR) 10Marostegui: [C:03+2] x4: Add candidate masters [puppet] - 10https://gerrit.wikimedia.org/r/1339594 (https://phabricator.wikimedia.org/T437678) (owner: 10Marostegui) [08:14:54] (03CR) 10Elukey: [C:03+2] Add conftool-data for the pki[12]003 VMs [puppet] - 10https://gerrit.wikimedia.org/r/1338236 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [08:15:47] PROBLEM - orchestrator resolve cache non-FQDNs on dborch1002 is CRITICAL: CRITICAL: 3 non-FQDN entries in orchestrator resolve cache: https://wikitech.wikimedia.org/wiki/Orchestrator [08:16:16] (03CR) 10Ladsgroup: [C:03+1] common.yaml: Add arbcom_arwiki [puppet] - 10https://gerrit.wikimedia.org/r/1339509 (https://phabricator.wikimedia.org/T437584) (owner: 10Marostegui) [08:16:34] (03CR) 10Marostegui: [C:03+2] common.yaml: Add arbcom_arwiki [puppet] - 10https://gerrit.wikimedia.org/r/1339509 (https://phabricator.wikimedia.org/T437584) (owner: 10Marostegui) [08:16:47] RECOVERY - orchestrator resolve cache non-FQDNs on dborch1002 is OK: OK: all orchestrator resolve cache entries are FQDNs https://wikitech.wikimedia.org/wiki/Orchestrator [08:18:34] (03CR) 10Elukey: "After some thoughts, I think we may have a bit of problem here. IIUC the loopback interface configs are based on the content of service.ya" [puppet] - 10https://gerrit.wikimedia.org/r/1338936 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [08:19:25] (03CR) 10Muehlenhoff: [C:03+1] "Looks good!" [dns] - 10https://gerrit.wikimedia.org/r/1338234 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [08:19:47] PROBLEM - orchestrator resolve cache non-FQDNs on dborch1002 is CRITICAL: CRITICAL: 3 non-FQDN entries in orchestrator resolve cache: https://wikitech.wikimedia.org/wiki/Orchestrator [08:20:44] !log arnaudb@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [08:20:51] !log arnaudb@cumin1003 END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [08:22:11] (03PS2) 10Elukey: role::pki: add lvs configurations [puppet] - 10https://gerrit.wikimedia.org/r/1338936 (https://phabricator.wikimedia.org/T436809) [08:22:18] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1338936 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [08:23:30] (03PS2) 10Arnaudb: gitlab: ask to skip backup when upgrading replicas [cookbooks] - 10https://gerrit.wikimedia.org/r/1339602 (https://phabricator.wikimedia.org/T437680) [08:24:07] (03PS1) 10Slyngshede: IDP: CAS updates [dns] - 10https://gerrit.wikimedia.org/r/1339605 [08:27:26] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5 [08:28:21] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5 [08:29:05] !log arnaudb@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [08:29:12] !log arnaudb@cumin1003 END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [08:32:22] FIRING: GnmiInterfaceCountersDrop: ... [08:32:22] lsw1-c5-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-c5-codfw:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [08:33:25] (03CR) 10Blake: "Thanks very much!" [puppet] - 10https://gerrit.wikimedia.org/r/1338238 (owner: 10Majavah) [08:33:29] (03CR) 10Blake: [C:03+2] P:memcached::instance: Cleanup deleted options [puppet] - 10https://gerrit.wikimedia.org/r/1338238 (owner: 10Majavah) [08:33:40] (03CR) 10Blake: [C:03+1] P:memcached::instance: Cleanup deleted options [puppet] - 10https://gerrit.wikimedia.org/r/1338238 (owner: 10Majavah) [08:35:23] (03CR) 10Blake: [C:03+1] "Thanks!" [software/httpbb] - 10https://gerrit.wikimedia.org/r/1338398 (https://phabricator.wikimedia.org/T434337) (owner: 10RLazarus) [08:36:08] !log arnaudb@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [08:37:22] RESOLVED: GnmiInterfaceCountersDrop: ... [08:37:22] lsw1-c5-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-c5-codfw:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [08:41:38] (03CR) 10Muehlenhoff: [C:03+1] "Looks good,smoke tests on 2005 and idp-test were fine!" [dns] - 10https://gerrit.wikimedia.org/r/1339605 (owner: 10Slyngshede) [08:41:55] (03CR) 10Slyngshede: [C:03+2] IDP: CAS updates [dns] - 10https://gerrit.wikimedia.org/r/1339605 (owner: 10Slyngshede) [08:43:13] !log slyngshede@dns1004 START - running authdns-update [08:44:06] 10SRE-swift-storage, 10MediaWiki-File-management, 06MediaWiki-Media-Platform-Team: Unknown error while undeleting a file on Wikimedia Commons - https://phabricator.wikimedia.org/T437114#12311172 (10MatthewVernon) If it's useful, there is a copy in backups, I think: ` 0) wiki | commonswiki tit... [08:44:31] (03PS1) 10Muehlenhoff: Remove obsolete buildkitd image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339620 [08:45:28] !log slyngshede@dns1004 END - running authdns-update [08:45:31] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab [08:45:47] (03PS1) 10Elukey: Allow registryctl to use a custom Docker config.json filepath [docker-images/docker-report] - 10https://gerrit.wikimedia.org/r/1339621 (https://phabricator.wikimedia.org/T437297) [08:46:33] !log Update CAS/SSO to CAS 7.3.8.3 [08:46:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:49:23] (03PS1) 10Joal: Grow webrequest retention time to 120 days [puppet] - 10https://gerrit.wikimedia.org/r/1339622 (https://phabricator.wikimedia.org/T437597) [08:51:11] (03CR) 10Brouberol: [C:03+1] Grow webrequest retention time to 120 days [puppet] - 10https://gerrit.wikimedia.org/r/1339622 (https://phabricator.wikimedia.org/T437597) (owner: 10Joal) [08:51:14] (03CR) 10Brouberol: [C:03+2] Grow webrequest retention time to 120 days [puppet] - 10https://gerrit.wikimedia.org/r/1339622 (https://phabricator.wikimedia.org/T437597) (owner: 10Joal) [08:51:17] (03CR) 10Brouberol: [V:03+2 C:03+2] Grow webrequest retention time to 120 days [puppet] - 10https://gerrit.wikimedia.org/r/1339622 (https://phabricator.wikimedia.org/T437597) (owner: 10Joal) [09:04:56] RECOVERY - SSH on an-worker1144 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [09:05:27] (03Abandoned) 10Btullis: an-worker1144: move to insetup role [puppet] - 10https://gerrit.wikimedia.org/r/1339115 (https://phabricator.wikimedia.org/T435873) (owner: 10Bking) [09:05:30] RECOVERY - MegaRAID on an-worker1144 is OK: OK: optimal, 12 logical, 13 physical, WriteBack policy https://wikitech.wikimedia.org/wiki/MegaCli%23Monitoring [09:05:38] (03PS1) 10Arnaudb: modules: Prepare mesh.configuration minor version bump [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338959 (https://phabricator.wikimedia.org/T436657) [09:05:48] (03PS6) 10Arnaudb: mesh: add opt-in websocket support in configuration 1.17.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338751 (https://phabricator.wikimedia.org/T436657) [09:05:51] (03CR) 10Btullis: [C:03+1] airflow3: updating airflow-client [deployment-charts] - 10https://gerrit.wikimedia.org/r/1335780 (https://phabricator.wikimedia.org/T423248) (owner: 10Atsuko) [09:05:56] (03PS1) 10Arnaudb: mesh: document the idle timeout key the templates actually read [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338960 (https://phabricator.wikimedia.org/T436657) [09:06:01] (03CR) 10Atsuko: [C:03+2] airflow3: updating airflow-client [deployment-charts] - 10https://gerrit.wikimedia.org/r/1335780 (https://phabricator.wikimedia.org/T423248) (owner: 10Atsuko) [09:06:12] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12311237 (10Ladsgroup) Which email address one of yours you want to be admin? the wikimedia.be one or the gmail.com one? [09:06:12] (03PS6) 10Arnaudb: scaffold: point new services at mesh 1.17 and istio 1.5 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338752 (https://phabricator.wikimedia.org/T436657) [09:08:29] (03Merged) 10jenkins-bot: airflow3: updating airflow-client [deployment-charts] - 10https://gerrit.wikimedia.org/r/1335780 (https://phabricator.wikimedia.org/T423248) (owner: 10Atsuko) [09:16:39] (03CR) 10Brouberol: "The helm part looks good, except for some small typos/details." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [09:18:48] (03CR) 10Brouberol: airflow: pluggable authmanager for K8s tokens (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [09:19:58] (03CR) 10Btullis: [C:03+2] cloudnative-pg: alert on WAL archiving failures rather than staleness [alerts] - 10https://gerrit.wikimedia.org/r/1338719 (https://phabricator.wikimedia.org/T432104) (owner: 10Btullis) [09:21:34] (03CR) 10Atsuko: "Will address typos, answered the quiestion." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [09:22:05] (03CR) 10Brouberol: airflow: pluggable authmanager for K8s tokens (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [09:22:14] (03Merged) 10jenkins-bot: cloudnative-pg: alert on WAL archiving failures rather than staleness [alerts] - 10https://gerrit.wikimedia.org/r/1338719 (https://phabricator.wikimedia.org/T432104) (owner: 10Btullis) [09:22:43] (03CR) 10Atsuko: "Thanks a lot!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [09:23:30] (03PS1) 10Muehlenhoff: Remove python-build Bullseye image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339635 (https://phabricator.wikimedia.org/T416452) [09:24:25] (03PS2) 10Muehlenhoff: Remove obsolete buildkitd image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339620 (https://phabricator.wikimedia.org/T416452) [09:24:55] (03CR) 10Atsuko: airflow: pluggable authmanager for K8s tokens (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [09:30:51] (03PS1) 10Muehlenhoff: Remove python-devel container image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1339636 (https://phabricator.wikimedia.org/T416452) [09:36:17] (03Abandoned) 10Muehlenhoff: Revert "admin: temp SSH key removal for jynus" [puppet] - 10https://gerrit.wikimedia.org/r/1338091 (owner: 10Jcrespo) [09:37:12] PROBLEM - Host wikikube-worker1138 is DOWN: PING CRITICAL - Packet loss = 66%, RTA = 5463.57 ms [09:37:46] RECOVERY - Host wikikube-worker1138 is UP: PING OK - Packet loss = 0%, RTA = 0.35 ms [09:45:45] (03PS1) 10Ladsgroup: cache: Reduce the webp threshold to 50 [puppet] - 10https://gerrit.wikimedia.org/r/1339641 (https://phabricator.wikimedia.org/T431150) [09:46:06] (03CR) 10Muehlenhoff: [C:03+2] Create a separate role openldap::replica_mdb [puppet] - 10https://gerrit.wikimedia.org/r/1337910 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [09:51:47] (03PS2) 10Ladsgroup: cache: Reduce the webp threshold to 50 [puppet] - 10https://gerrit.wikimedia.org/r/1339641 (https://phabricator.wikimedia.org/T431150) [09:51:52] (03CR) 10Ladsgroup: [V:03+2 C:03+2] cache: Reduce the webp threshold to 50 [puppet] - 10https://gerrit.wikimedia.org/r/1339641 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [09:52:04] (03CR) 10Federico Ceratto: data-persistence: Alert on depooled hosts without silence (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [10:01:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:15:27] !log aokoth@cumin1004 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T437680 [10:17:03] !log restart versitygw@objectstorage00.service on backup1015 [10:17:05] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:20:32] !log restart versitygw@objectstorage0[1-3].service on backup1015 in turn [10:20:33] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:30:31] !log restart versitygw@objectstorage0[0-3].service on backup1016 in turn [10:30:33] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:32:04] !log restart versitygw@objectstorage0[0-3].service on backup1017 in turn [10:32:06] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:33:08] aokoth@cumin1004 aokoth: The backup on gitlab1004 is complete, ready to proceed with upgrade. [10:33:33] !log restart versitygw@objectstorage0[0-3].service on backup1018 in turn [10:33:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:34:53] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1199: db1199 repool [10:35:17] !log restart versitygw@objectstorage0[0-3].service on backup1019 in turn [10:35:19] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:36:18] !log restart versitygw@objectstorage0[0-3].service on backup1020 in turn [10:36:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:36:28] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12311518 (10Marostegui) I've cloned and started replication on db1245 for both s4 and s5. Let's see if it crashes. I will reopen if that happens. If not, putting the back in service is be... [10:36:31] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12311520 (10Marostegui) 05Open→03Resolved [10:37:03] !log restart versitygw@objectstorage0[0-3].service on backup2015 in turn [10:37:05] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:38:57] (03CR) 10Santiago Faci: [C:03+2] Test Kitchen UI: Deploy v1.5.5. release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339043 (https://phabricator.wikimedia.org/T429524) (owner: 10Santiago Faci) [10:39:18] !log restart versitygw@objectstorage0[0-3].service on backup2016 in turn [10:39:19] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:40:08] PROBLEM - Gitlab HTTPS healthcheck on gitlab.wikimedia.org is CRITICAL: HTTP CRITICAL: HTTP/1.1 502 Bad Gateway - 2353 bytes in 0.018 second response time https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [10:40:25] !log restart versitygw@objectstorage0[0-3].service on backup2017 in turn [10:40:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:41:06] !log restart versitygw@objectstorage0[0-3].service on backup2018 in turn [10:41:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:41:08] RECOVERY - Gitlab HTTPS healthcheck on gitlab.wikimedia.org is OK: HTTP OK: HTTP/1.1 200 OK - 28821 bytes in 0.102 second response time https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [10:41:17] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v1.5.5. release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339043 (https://phabricator.wikimedia.org/T429524) (owner: 10Santiago Faci) [10:41:49] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db1261 to x4 master [puppet] - 10https://gerrit.wikimedia.org/r/1339672 (https://phabricator.wikimedia.org/T437695) [10:41:56] (03PS1) 10Gerrit maintenance bot: wmnet: Update x4-master alias [dns] - 10https://gerrit.wikimedia.org/r/1339673 (https://phabricator.wikimedia.org/T437695) [10:42:30] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db2245 to x4 master [puppet] - 10https://gerrit.wikimedia.org/r/1339674 (https://phabricator.wikimedia.org/T437696) [10:42:44] (03Abandoned) 10Marostegui: wmnet: Update x4-master alias [dns] - 10https://gerrit.wikimedia.org/r/1339673 (https://phabricator.wikimedia.org/T437695) (owner: 10Gerrit maintenance bot) [10:42:50] (03Abandoned) 10Marostegui: mariadb: Promote db1261 to x4 master [puppet] - 10https://gerrit.wikimedia.org/r/1339672 (https://phabricator.wikimedia.org/T437695) (owner: 10Gerrit maintenance bot) [10:43:00] !log restart versitygw@objectstorage0[0-3].service on backup2019 in turn [10:43:01] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:43:07] (03Abandoned) 10Marostegui: mariadb: Promote db2245 to x4 master [puppet] - 10https://gerrit.wikimedia.org/r/1339674 (https://phabricator.wikimedia.org/T437696) (owner: 10Gerrit maintenance bot) [10:43:44] !log restart versitygw@objectstorage0[0-3].service on backup2020 in turn [10:43:45] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:44:36] !log aokoth@cumin1004 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T437680 [11:00:04] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260911T0700) [11:00:04] jelto, arnoldokoth, mutante, and arnaudb: It is that lovely time of the day again! You are hereby commanded to deploy GitLab version upgrades. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260911T1100). [11:02:09] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12311654 (10Marostegui) 05Resolved→03Open Leaving open until we hear back from Dell. [11:05:11] !log jmm@cumin2003 DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts [11:05:57] !log installing Linux 6.1.187 on Bookworm hosts [11:05:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:06:27] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12311679 (10Geertivp) Please, gmail.com [11:08:36] (03PS1) 10Zabe: Add Apache configuration for wikipedia-ar-arbcom.wikimedia.org [puppet] - 10https://gerrit.wikimedia.org/r/1339694 (https://phabricator.wikimedia.org/T437403) [11:12:21] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12311690 (10Ladsgroup) 05Open→03Resolved a:03Ladsgroup I added that account as admin since you are the chair of WMBE. So now you can do whatever you deem needed. [11:16:48] RECOVERY - orchestrator resolve cache non-FQDNs on dborch1002 is OK: OK: all orchestrator resolve cache entries are FQDNs https://wikitech.wikimedia.org/wiki/Orchestrator [11:20:00] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool [11:20:56] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.14 point update - https://phabricator.wikimedia.org/T426759#12311717 (10MoritzMuehlenhoff) [11:21:34] FIRING: DiskSpace: Disk space puppetboard1003:9100:/ 3.578% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=puppetboard1003 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [11:23:10] 06SRE, 10Data-Persistence-Backup, 10database-backups: Put db2201 back into backup production as a backup source - https://phabricator.wikimedia.org/T437411#12311746 (10Marostegui) Change to be reverted whenever ready: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1324672 [11:26:34] RESOLVED: DiskSpace: Disk space puppetboard1003:9100:/ 3.489% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=puppetboard1003 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [11:46:54] (03PS6) 10Atsuko: airflow: pluggable authmanager for K8s tokens [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) [11:46:54] (03PS1) 10Atsuko: airflow: service account projected volume [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339714 [11:47:50] PROBLEM - Druid historical on an-druid1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [11:56:38] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install pki2003 - https://phabricator.wikimedia.org/T436179#12311845 (10MoritzMuehlenhoff) In the mean time pki2003 was created as a VM, as such I'm retitling the task and updating the racking details. [11:56:52] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Q1:rack/setup/install pki2004 - https://phabricator.wikimedia.org/T436179#12311848 (10MoritzMuehlenhoff) [11:57:02] (03PS1) 10Blake: memcached: Improve readability of tls variable creation. [puppet] - 10https://gerrit.wikimedia.org/r/1339718 (https://phabricator.wikimedia.org/T353511) [11:58:22] (03CR) 10Btullis: [C:03+1] "Nice work." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339714 (owner: 10Atsuko) [11:59:16] (03CR) 10Clément Goubert: [C:03+1] memcached: Improve readability of tls variable creation. [puppet] - 10https://gerrit.wikimedia.org/r/1339718 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [11:59:17] (03PS2) 10Blake: memcached: Improve readability of tls variable creation. [puppet] - 10https://gerrit.wikimedia.org/r/1339718 (https://phabricator.wikimedia.org/T353511) [12:01:50] RECOVERY - Druid historical on an-druid1007 is OK: PROCS OK: 1 process with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [12:02:13] (03CR) 10Blake: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1339718 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [12:08:54] (03CR) 10Brouberol: [C:03+1] "Nicely done!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [12:10:24] (03CR) 10Btullis: airflow: pluggable authmanager for K8s tokens (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [12:11:04] (03CR) 10Muehlenhoff: [C:03+1] "Looks good!" [puppet] - 10https://gerrit.wikimedia.org/r/1338224 (owner: 10Elukey) [12:11:19] (03PS3) 10Blake: memcached: Improve readability of tls variable creation. [puppet] - 10https://gerrit.wikimedia.org/r/1339718 (https://phabricator.wikimedia.org/T353511) [12:11:42] (03CR) 10Brouberol: [C:03+1] "Nicely done!!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339714 (owner: 10Atsuko) [12:12:38] (03CR) 10Majavah: [C:03+2] P:memcached::instance: Cleanup deleted options [puppet] - 10https://gerrit.wikimedia.org/r/1338238 (owner: 10Majavah) [12:14:25] (03CR) 10Majavah: [C:03+2] mediawiki::tools: mediawiki-cache-warmup: Remove mobileServer support [puppet] - 10https://gerrit.wikimedia.org/r/1337933 (owner: 10Majavah) [12:25:37] (03Abandoned) 10Blake: memcached: Improve readability of tls variable creation. [puppet] - 10https://gerrit.wikimedia.org/r/1339718 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [12:31:26] (03PS1) 10Muehlenhoff: Remove unused profile::openldap::hostname [puppet] - 10https://gerrit.wikimedia.org/r/1339732 (https://phabricator.wikimedia.org/T331699) [12:33:35] (03CR) 10CI reject: [V:04-1] Remove unused profile::openldap::hostname [puppet] - 10https://gerrit.wikimedia.org/r/1339732 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [12:37:24] (03CR) 10Atsuko: "@btullis@wikimedia.org thanks for the comments, I've addressed them." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338217 (https://phabricator.wikimedia.org/T437236) (owner: 10Atsuko) [12:38:01] (03PS2) 10Muehlenhoff: Remove unused profile::openldap::hostname [puppet] - 10https://gerrit.wikimedia.org/r/1339732 (https://phabricator.wikimedia.org/T331699) [12:41:01] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1339732 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [12:44:10] FIRING: BFDdown: BFD session down between cr1-esams and fe80::6687:8807:df2:7018 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-esams:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:48:34] (03PS1) 10LWatson: Enable ReaderExperiments in eswiki, jawiki, and ptwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339741 (https://phabricator.wikimedia.org/T436199) [12:49:10] RESOLVED: BFDdown: BFD session down between cr1-esams and fe80::6687:8807:df2:7018 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-esams:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [12:49:58] (03CR) 10Elukey: "@taavi@wikimedia.org hi! Do you think this refactoring is good to go? I am a little bit worried about the cloud.yaml's comment "this will " [puppet] - 10https://gerrit.wikimedia.org/r/1338224 (owner: 10Elukey) [12:50:52] (03CR) 10Elukey: [C:03+1] Remove unused profile::openldap::hostname [puppet] - 10https://gerrit.wikimedia.org/r/1339732 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [12:50:53] (03CR) 10Giuseppe Lavagetto: profile::containerd: add gVisor support (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [12:51:16] (03CR) 10Giuseppe Lavagetto: "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1330130 (owner: 10Giuseppe Lavagetto) [12:51:33] (03CR) 10Giuseppe Lavagetto: "Done" [puppet] - 10https://gerrit.wikimedia.org/r/1330130 (owner: 10Giuseppe Lavagetto) [12:53:30] (03PS5) 10Giuseppe Lavagetto: containerd: use hiera parameters to detect if dragonfly is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1330130 [12:53:30] (03PS5) 10Giuseppe Lavagetto: profile::containerd: add gVisor support [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) [12:53:30] (03PS5) 10Giuseppe Lavagetto: kubernetes::staging: make containerd support gVisor runtime in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1330131 (https://phabricator.wikimedia.org/T435796) [12:53:31] (03PS5) 10Giuseppe Lavagetto: profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) [12:53:32] (03PS2) 10Giuseppe Lavagetto: kubernetes::staging: make containerd support gVisor runtime [puppet] - 10https://gerrit.wikimedia.org/r/1338189 [12:55:42] (03CR) 10CI reject: [V:04-1] profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [12:59:33] (03PS4) 10Federico Ceratto: sre.mysql.decommission: optional DC-ops handover [cookbooks] - 10https://gerrit.wikimedia.org/r/1328570 (https://phabricator.wikimedia.org/T435912) [12:59:34] (03CR) 10Federico Ceratto: "This makes the handover optional" [cookbooks] - 10https://gerrit.wikimedia.org/r/1328570 (https://phabricator.wikimedia.org/T435912) (owner: 10Federico Ceratto) [13:03:16] (03CR) 10Majavah: "they are overridden in enc hiera (the puppet panel in horizon), codesearch finds these overrides: https://codesearch.wmcloud.org/puppet/?q" [puppet] - 10https://gerrit.wikimedia.org/r/1338224 (owner: 10Elukey) [13:06:25] RESOLVED: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:09:19] (03CR) 10Giuseppe Lavagetto: admin: add support for gVisor RuntimeClass handlers (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [13:10:41] 10ops-eqiad, 06SRE, 06DC-Ops, 06cloud-services-team (Hardware), 13Patch-For-Review: Q3:rack/setup/install cloudcephosd105[3456] - https://phabricator.wikimedia.org/T419892#12312045 (10elukey) @Andrew my understanding is that cloudcephosd[1048-52] shouldn't be able to reimage with the current settings. If... [13:10:44] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [13:11:19] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [13:15:21] (03CR) 10Elukey: "Nice thanks! I have added the new key to deployment-prep's hiera project:" [puppet] - 10https://gerrit.wikimedia.org/r/1338224 (owner: 10Elukey) [13:15:45] (03CR) 10Giuseppe Lavagetto: shellbox: add gVisor support (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333139 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [13:15:47] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [13:15:57] (03CR) 10Giuseppe Lavagetto: "Done" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333139 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [13:16:05] (03PS5) 10Giuseppe Lavagetto: admin: add support for gVisor RuntimeClass handlers [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) [13:16:06] (03PS3) 10Giuseppe Lavagetto: validating-admission-policies: add gVisor enforcement policy [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333210 (https://phabricator.wikimedia.org/T436655) [13:16:06] (03PS5) 10Giuseppe Lavagetto: shellbox: add gVisor support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333139 (https://phabricator.wikimedia.org/T436649) [13:16:22] (03CR) 10Majavah: [C:03+1] "indeed!" [puppet] - 10https://gerrit.wikimedia.org/r/1338224 (owner: 10Elukey) [13:16:42] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [13:16:50] FIRING: [2x] ProxoidHigh5xxErrorRate: hcaptcha-proxy1001:3903: >1% of requests are 5xx errors in eqiad - https://wikitech.wikimedia.org/wiki/HCaptcha#Runbook - https://alerts.wikimedia.org/?q=alertname%3DProxoidHigh5xxErrorRate [13:19:38] (03PS1) 10Elukey: Default redfish's spicerack attribute to wmfroot [software/spicerack] - 10https://gerrit.wikimedia.org/r/1339749 (https://phabricator.wikimedia.org/T426180) [13:20:51] (03PS2) 10Bking: Druid: Point PSI alerts to the correct prometheus instance [alerts] - 10https://gerrit.wikimedia.org/r/1339209 (https://phabricator.wikimedia.org/T436968) [13:21:35] (03CR) 10Bking: "No, you were right. I forgot to add the correct files to the change. This is now fixed." [alerts] - 10https://gerrit.wikimedia.org/r/1339209 (https://phabricator.wikimedia.org/T436968) (owner: 10Bking) [13:23:01] (03CR) 10Brouberol: [C:03+1] Druid: Point PSI alerts to the correct prometheus instance [alerts] - 10https://gerrit.wikimedia.org/r/1339209 (https://phabricator.wikimedia.org/T436968) (owner: 10Bking) [13:24:56] (03CR) 10Bking: [C:03+2] Druid: Point PSI alerts to the correct prometheus instance [alerts] - 10https://gerrit.wikimedia.org/r/1339209 (https://phabricator.wikimedia.org/T436968) (owner: 10Bking) [13:26:50] FIRING: [2x] ProxoidHigh5xxErrorRate: hcaptcha-proxy1001:3903: >1% of requests are 5xx errors in eqiad - https://wikitech.wikimedia.org/wiki/HCaptcha#Runbook - https://alerts.wikimedia.org/?q=alertname%3DProxoidHigh5xxErrorRate [13:27:53] (03CR) 10Elukey: [C:03+2] Rename docker::registry hiera key to docker_registry_endpoint [puppet] - 10https://gerrit.wikimedia.org/r/1338224 (owner: 10Elukey) [13:28:56] !log bking@cumin2003 START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade. [13:32:34] (03CR) 10Arnaudb: "Both hops on the client leg have been tested. The mesh sidecar answers a websocket upgrade with a 403, before the application sees anythin" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338751 (https://phabricator.wikimedia.org/T436657) (owner: 10Arnaudb) [13:33:22] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [13:33:37] (03CR) 10JMeybohm: [C:03+1] containerd: use hiera parameters to detect if dragonfly is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1330130 (owner: 10Giuseppe Lavagetto) [13:33:46] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [13:33:46] 06SRE, 10SRE-swift-storage, 06Infrastructure-Foundations: Unable to reimage or reprovision ms-be2082 due to redfish connection errors - https://phabricator.wikimedia.org/T433635#12312105 (10MatthewVernon) 05Open→03Resolved a:03MatthewVernon We reimaged this node OK, and all the SM C-J nodes are UEF... [13:35:39] (03PS1) 10Dpogorzelski: liftwing-studio: chart for Open WebUI and LiteLLM [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339756 (https://phabricator.wikimedia.org/T437706) [13:37:01] (03CR) 10JMeybohm: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9405/co" [puppet] - 10https://gerrit.wikimedia.org/r/1330131 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [13:37:11] (03PS4) 10Federico Ceratto: core-mysql.my.cnf.erb: support templating expire_logs_days for core_test [puppet] - 10https://gerrit.wikimedia.org/r/1337601 (https://phabricator.wikimedia.org/T435059) [13:38:17] (03PS1) 10PipelineBot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339757 [13:40:52] !log bking@cumin2003 END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade. [13:41:18] (03CR) 10JMeybohm: [V:03+1] "PCC SUCCESS (CORE_DIFF 13): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9406/c" [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [13:42:37] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.14 point update - https://phabricator.wikimedia.org/T426759#12312140 (10MoritzMuehlenhoff) [13:42:41] (03PS6) 10Giuseppe Lavagetto: profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) [13:42:41] (03PS3) 10Giuseppe Lavagetto: kubernetes::staging: make containerd support gVisor runtime [puppet] - 10https://gerrit.wikimedia.org/r/1338189 [13:43:56] (03CR) 10CI reject: [V:04-1] profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [13:48:47] (03CR) 10Federico Ceratto: "I added max-binlog-size in a way that does not change the conf file for prod hosts at all. I suspect that compressing the binlog would not" [puppet] - 10https://gerrit.wikimedia.org/r/1337601 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [13:49:13] (03CR) 10JMeybohm: [C:03+1] admin: add support for gVisor RuntimeClass handlers (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [13:50:41] (03CR) 10JMeybohm: [C:03+1] "Done" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333210 (https://phabricator.wikimedia.org/T436655) (owner: 10Giuseppe Lavagetto) [13:51:13] (03CR) 10JMeybohm: [C:03+1] shellbox: add gVisor support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333139 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [13:51:50] FIRING: [2x] ProxoidHigh5xxErrorRate: hcaptcha-proxy1001:3903: >1% of requests are 5xx errors in eqiad - https://wikitech.wikimedia.org/wiki/HCaptcha#Runbook - https://alerts.wikimedia.org/?q=alertname%3DProxoidHigh5xxErrorRate [13:52:46] (03PS7) 10Giuseppe Lavagetto: profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) [13:52:46] (03PS4) 10Giuseppe Lavagetto: kubernetes::staging: make containerd support gVisor runtime [puppet] - 10https://gerrit.wikimedia.org/r/1338189 [13:53:45] (03PS2) 10Arnaudb: modules: Prepare mesh.configuration minor version bump [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338959 (https://phabricator.wikimedia.org/T436657) [13:53:46] (03PS7) 10Arnaudb: mesh: add opt-in websocket support in configuration 1.17.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338751 (https://phabricator.wikimedia.org/T436657) [13:53:46] (03PS2) 10Arnaudb: mesh: document the idle timeout key the templates actually read [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338960 (https://phabricator.wikimedia.org/T436657) [13:53:46] (03PS7) 10Arnaudb: scaffold: point new services at mesh 1.17 and istio 1.5 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338752 (https://phabricator.wikimedia.org/T436657) [13:56:50] FIRING: [2x] ProxoidHigh5xxErrorRate: hcaptcha-proxy1001:3903: >1% of requests are 5xx errors in eqiad - https://wikitech.wikimedia.org/wiki/HCaptcha#Runbook - https://alerts.wikimedia.org/?q=alertname%3DProxoidHigh5xxErrorRate [14:06:45] (03PS6) 10Giuseppe Lavagetto: profile::containerd: add gVisor support [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) [14:06:45] (03PS6) 10Giuseppe Lavagetto: kubernetes::staging: make containerd support gVisor runtime in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1330131 (https://phabricator.wikimedia.org/T435796) [14:06:45] (03PS8) 10Giuseppe Lavagetto: profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) [14:06:46] (03PS5) 10Giuseppe Lavagetto: kubernetes::staging: make containerd support gVisor runtime [puppet] - 10https://gerrit.wikimedia.org/r/1338189 [14:07:28] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9407/co" [puppet] - 10https://gerrit.wikimedia.org/r/1330131 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:09:54] !log ssastry@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [14:10:23] (03PS2) 10Dpogorzelski: liftwing-studio: chart for Open WebUI and LiteLLM [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339756 (https://phabricator.wikimedia.org/T437706) [14:10:23] !log ssastry@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [14:10:24] !log ssastry@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [14:10:52] !log ssastry@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [14:11:21] (03CR) 10JMeybohm: [C:03+1] profile::containerd: add gVisor support [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:11:33] (03CR) 10JMeybohm: [C:03+1] kubernetes::staging: make containerd support gVisor runtime in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1330131 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:11:50] RESOLVED: [2x] ProxoidHigh5xxErrorRate: hcaptcha-proxy1001:3903: >1% of requests are 5xx errors in eqiad - https://wikitech.wikimedia.org/wiki/HCaptcha#Runbook - https://alerts.wikimedia.org/?q=alertname%3DProxoidHigh5xxErrorRate [14:14:08] (03CR) 10JMeybohm: [V:03+1] "PCC SUCCESS (CORE_DIFF 13): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9408/c" [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [14:14:58] RECOVERY - Host an-worker1204 is UP: PING OK - Packet loss = 0%, RTA = 0.25 ms [14:14:58] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.7 point update - https://phabricator.wikimedia.org/T437715 (10MoritzMuehlenhoff) 03NEW [14:15:03] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.7 point update - https://phabricator.wikimedia.org/T437715#12312267 (10MoritzMuehlenhoff) p:05Triage→03Medium [14:17:28] (03CR) 10JMeybohm: [V:03+1] profile::kubernetes::node: Add gvisor labels (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [14:18:18] (03PS9) 10Giuseppe Lavagetto: profile::kubernetes::node: Add gvisor labels [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) [14:19:44] (03CR) 10JMeybohm: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9409/co" [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [14:20:47] (03CR) 10JMeybohm: [V:03+1 C:03+1] profile::kubernetes::node: Add gvisor labels (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1330187 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [14:20:53] (03CR) 10Ssingh: Move the pki service entry to lvs_setup (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338937 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [14:20:56] (03CR) 10JMeybohm: [C:03+1] kubernetes::staging: make containerd support gVisor runtime [puppet] - 10https://gerrit.wikimedia.org/r/1338189 (owner: 10Giuseppe Lavagetto) [14:28:47] (03CR) 10JHathaway: [C:03+1] Default redfish's spicerack attribute to wmfroot [software/spicerack] - 10https://gerrit.wikimedia.org/r/1339749 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [14:31:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.35% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:35:16] (03PS1) 10Bking: sre.hadoop.roll-restart-workers: optionally skip hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1339776 (https://phabricator.wikimedia.org/T437713) [14:39:43] !log bking@cumin2003 START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade. [14:41:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 25% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:43:10] (03PS1) 10Muehlenhoff: Enable dbbackups::transfer for cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1339778 (https://phabricator.wikimedia.org/T427897) [14:46:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:46:28] PROBLEM - Host cp3066 is DOWN: PING CRITICAL - Packet loss = 100% [14:46:30] PROBLEM - Host ncredir3005 is DOWN: PING CRITICAL - Packet loss = 100% [14:46:35] (03PS2) 10Muehlenhoff: Enable dbbackups::transfer for cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1339778 (https://phabricator.wikimedia.org/T427897) [14:46:48] RECOVERY - Host ncredir3005 is UP: PING WARNING - Packet loss = 75%, RTA = 78.70 ms [14:46:48] RECOVERY - Host cp3066 is UP: PING OK - Packet loss = 0%, RTA = 78.06 ms [14:47:28] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12312415 (10MoritzMuehlenhoff) [14:48:01] uh? [14:50:44] (03PS1) 10PipelineBot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339781 [14:52:22] (03CR) 10Elukey: [C:03+2] Default redfish's spicerack attribute to wmfroot [software/spicerack] - 10https://gerrit.wikimedia.org/r/1339749 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [14:56:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:05:36] (03CR) 10Elukey: [C:03+1] "I have zero context on the vendored module but the change and new facts occurrences LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1337990 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:06:07] (03CR) 10JHathaway: [C:03+2] rspamd: use new repo fork [puppet] - 10https://gerrit.wikimedia.org/r/1337990 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:30:06] 06SRE, 10corto, 10Incident Tooling: Increase trusted volunteer's visibility into production incidents - https://phabricator.wikimedia.org/T426137#12312543 (10sbassett) >>! In T426137#12278248, @Novem_Linguae wrote: >>>! In T426137#12277753, @SomeRandomDeveloper wrote: >> However, out of the 9 other members i... [15:33:14] (03PS6) 10Btullis: ceph: Make the CephX import guard compare the key [puppet] - 10https://gerrit.wikimedia.org/r/1337876 (https://phabricator.wikimedia.org/T437233) [15:41:42] (03CR) 10JHathaway: "Keith if you could review this, in Cole's absence, that would be great." [puppet] - 10https://gerrit.wikimedia.org/r/1338040 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:45:52] (03CR) 10Muehlenhoff: "I think generally we should have a mechanism were we designate the intended profile of a role in Puppet, so that we can apply the appropri" [cookbooks] - 10https://gerrit.wikimedia.org/r/1333146 (https://phabricator.wikimedia.org/T435537) (owner: 10Elukey) [15:50:00] (03PS1) 10SBassett: Use escaped() for story link parentheses in recent changes [extensions/Wikistories] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1339808 (https://phabricator.wikimedia.org/T182213) [15:56:08] Hey all - I’d like to get some patches out in a minute via spiderpig. I know it’s Friday. Please let me know if you have any objections. [15:58:30] (03PS1) 10Aghirelli: wmf-config: Register content/v2-beta REST module as disabled [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339811 (https://phabricator.wikimedia.org/T432798) [15:58:31] cc claime, Amir1 as oncallers [15:58:33] ^ [15:58:56] shoot [16:04:41] (03CR) 10Dzahn: [C:03+2] "ACK, sounds good to me" [dns] - 10https://gerrit.wikimedia.org/r/1338267 (https://phabricator.wikimedia.org/T437403) (owner: 10Dzahn) [16:08:01] !log bking@cumin2003 END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade. [16:11:42] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:17:21] (03PS1) 10SBassett: Use escaped() for HTML parentheses params in ChangeLineFormatter [extensions/Wikibase] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1339813 (https://phabricator.wikimedia.org/T182213) [16:17:55] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/Wikistories] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1339808 (https://phabricator.wikimedia.org/T182213) (owner: 10SBassett) [16:17:56] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/Wikibase] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1339813 (https://phabricator.wikimedia.org/T182213) (owner: 10SBassett) [16:19:24] (03Merged) 10jenkins-bot: Use escaped() for story link parentheses in recent changes [extensions/Wikistories] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1339808 (https://phabricator.wikimedia.org/T182213) (owner: 10SBassett) [16:24:43] 10ops-eqiad, 06SRE, 06DC-Ops, 10Kafka-Infrastructure, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12312748 (10VRiley-WMF) Okay, so there are a few items that we'll need to consider. Row A currently only have a... [16:25:40] (03CR) 10Kimberly Sarabia: [C:03+1] Enable ReaderExperiments in eswiki, jawiki, and ptwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1339741 (https://phabricator.wikimedia.org/T436199) (owner: 10LWatson) [16:32:31] 06SRE, 10LDAP-Access-Requests: Grant Access to for - https://phabricator.wikimedia.org/T437730 (10Infinity198) 03NEW [16:34:39] 06SRE, 10SRE-Access-Requests: Requesting access to releasers-mobile for LPetty-WMF - https://phabricator.wikimedia.org/T437662#12312785 (10Dzahn) [16:39:14] (03PS1) 10Dzahn: admin: add Lamar Petty and add them to releasers-mobile [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) [16:39:36] (03CR) 10CI reject: [V:04-1] admin: add Lamar Petty and add them to releasers-mobile [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) (owner: 10Dzahn) [16:40:27] !log sbassett@deploy1003 Started scap sync-world: Backport for [[gerrit:1339808|Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813|Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] [16:40:55] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to releasers-mobile for LPetty-WMF - https://phabricator.wikimedia.org/T437662#12312797 (10Dzahn) Hello @thcipriani or @dancy this will require your approval as owners of the `releasers-mobile` group. [16:42:43] !log sbassett@deploy1003 sbassett: Backport for [[gerrit:1339808|Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813|Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [16:43:16] !log sbassett@deploy1003 sbassett: Continuing with deployment [16:47:50] !log sbassett@deploy1003 Finished scap sync-world: Backport for [[gerrit:1339808|Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813|Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s) [17:07:36] (03CR) 10RLazarus: [C:03+1] "Nice refactor!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338227 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [17:18:38] (03CR) 10Dzahn: "broken puppet on CI servers: Server Error: Function lookup() did not find a value for the name 'docker_registry_endpoint'" [puppet] - 10https://gerrit.wikimedia.org/r/1338224 (owner: 10Elukey) [17:20:15] (03PS2) 10Bking: sre.hadoop.roll-restart-workers: optionally skip hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1339776 (https://phabricator.wikimedia.org/T437713) [17:48:08] (03CR) 10RLazarus: [C:03+2] "Thanks for the review!" [software/httpbb] - 10https://gerrit.wikimedia.org/r/1338398 (https://phabricator.wikimedia.org/T434337) (owner: 10RLazarus) [17:49:22] (03Merged) 10jenkins-bot: Drop support for Python 3.9 [software/httpbb] - 10https://gerrit.wikimedia.org/r/1338398 (https://phabricator.wikimedia.org/T434337) (owner: 10RLazarus) [17:56:22] (03PS1) 10RLazarus: mcrouter-wancache: Shift traffic from mc-wf1002 to mc-wf1001 [puppet] - 10https://gerrit.wikimedia.org/r/1339855 (https://phabricator.wikimedia.org/T421711) [17:56:22] (03CR) 10RLazarus: [C:04-1] "Holding for Monday." [puppet] - 10https://gerrit.wikimedia.org/r/1339855 (https://phabricator.wikimedia.org/T421711) (owner: 10RLazarus) [18:21:38] (03PS1) 10Ssingh: admin: add volker-e to analytics-privatedata-users (krb) [puppet] - 10https://gerrit.wikimedia.org/r/1339865 (https://phabricator.wikimedia.org/T437634) [18:25:15] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313153 (10ssingh) @thcipriani / @dancy: This requires your approval please, for both groups (`deployment` and `restricted`). [18:26:02] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313154 (10ssingh) [18:26:48] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313155 (10ssingh) @WRai-WMF: You already have a key on file, for your current shell access, which seems to be different than the one here. Do you want to use this new key? If... [18:27:34] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313158 (10taavi) `deployment` includes everything `restricted` can do, no need to be in both groups. [18:29:26] (03PS1) 10Ssingh: admin: add wrai to deployment [puppet] - 10https://gerrit.wikimedia.org/r/1339866 (https://phabricator.wikimedia.org/T437652) [18:29:53] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313176 (10ssingh) >>! In T437652#12313158, @taavi wrote: > `deployment` includes everything `restricted` can do, no need to be in both groups. Thanks, I... [18:30:48] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and analytics-privatedata-users for Cklimas - https://phabricator.wikimedia.org/T437653#12313178 (10ssingh) @thcipriani / @dancy: This requires your approval please, `deployment` group. [18:34:51] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313198 (10WRai-WMF) > You already have a key on file, for your current shell access, which seems to be different than the one here. Do you want to use th... [18:39:10] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313201 (10ssingh) >>! In T437652#12313198, @WRai-WMF wrote: >> You already have a key on file, for your current shell access, which seems to be different... [18:39:40] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313202 (10ssingh) [18:39:52] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313206 (10ssingh) [18:40:39] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for WRai-WMF - https://phabricator.wikimedia.org/T437652#12313207 (10WRai-WMF) Thank you. [18:45:25] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and analytics-privatedata-users for Cklimas - https://phabricator.wikimedia.org/T437653#12313210 (10ssingh) @Cklimas: Hi. Question about the `analytics-privatedata-users`: it is safe to assume that you do not require access to private data, correct?... [18:46:37] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and analytics-privatedata-users for Cklimas - https://phabricator.wikimedia.org/T437653#12313211 (10Cklimas) Going to tag @Seddon to answer that question because I'm not familiar enough with the use case I'll have to be able to give you a good answer. [18:48:54] (03PS1) 10Ssingh: admin: add shell access for cklimas (deployment) [puppet] - 10https://gerrit.wikimedia.org/r/1339879 (https://phabricator.wikimedia.org/T437653) [18:49:55] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and analytics-privatedata-users for Cklimas - https://phabricator.wikimedia.org/T437653#12313219 (10ssingh) [19:08:02] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and restricted for EAlbizzati-WMF - https://phabricator.wikimedia.org/T437741 (10EAlbizzati-WMF) 03NEW [19:30:48] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and restricted for EAlbizzati-WMF - https://phabricator.wikimedia.org/T437741#12313301 (10Urbanecm) FWIW, `restricted` is a subset of `deployment` in terms of permission (if you are in `deployment`, you can do whatever `restricted` can automatically). [19:32:36] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12313306 (10Geertivp) Thanks, a lot, now I have admin access ! [19:34:54] (03PS1) 10Dzahn: Revert^2 "zuul: add firewall rules for Zookeeper" [puppet] - 10https://gerrit.wikimedia.org/r/1339903 [19:37:07] 06SRE, 10SRE-Access-Requests: Requesting access to deployment and restricted for cooltey - https://phabricator.wikimedia.org/T437658#12313325 (10ssingh) @thcipriani / @dancy: This requires your approval please, for deployment. [19:39:27] (03PS2) 10Ssingh: admin: add Lamar Petty and add them to releasers-mobile [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) (owner: 10Dzahn) [19:40:16] (03CR) 10CI reject: [V:04-1] admin: add Lamar Petty and add them to releasers-mobile [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) (owner: 10Dzahn) [19:41:36] (03PS3) 10Dzahn: admin: add Lamar Petty and add them to releasers-mobile [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) [19:42:36] (03CR) 10Ssingh: [C:03+1] "Looks good, uid and access, thanks for the patch! Will merge after approval." [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) (owner: 10Dzahn) [19:42:51] (03CR) 10Dzahn: ":) cool" [puppet] - 10https://gerrit.wikimedia.org/r/1339821 (https://phabricator.wikimedia.org/T437662) (owner: 10Dzahn) [19:44:40] 10ops-eqiad, 06SRE, 10SRE-swift-storage, 06Data-Persistence, 06DC-Ops: Q4:rack/setup/install ms-be1098, ms-be1099, ms-be1100 - https://phabricator.wikimedia.org/T424895#12313334 (10VRiley-WMF) [19:46:01] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to releasers-mobile for LPetty-WMF - https://phabricator.wikimedia.org/T437662#12313337 (10ssingh) [19:49:04] (03PS1) 10Ssingh: admin: add cooltey to deployment [puppet] - 10https://gerrit.wikimedia.org/r/1339906 (https://phabricator.wikimedia.org/T437658) [19:50:52] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for cooltey - https://phabricator.wikimedia.org/T437658#12313355 (10ssingh) [19:51:42] (03PS2) 10Dzahn: Revert^2 "zuul: add firewall rules for Zookeeper" [puppet] - 10https://gerrit.wikimedia.org/r/1339903 [19:52:13] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment and restricted for cooltey - https://phabricator.wikimedia.org/T437658#12313362 (10ssingh) [19:53:16] (03CR) 10Dzahn: [C:03+2] Revert^2 "zuul: add firewall rules for Zookeeper" [puppet] - 10https://gerrit.wikimedia.org/r/1339903 (owner: 10Dzahn) [19:53:36] (03CR) 10BCornwall: [C:03+1] admin: add cooltey to deployment [puppet] - 10https://gerrit.wikimedia.org/r/1339906 (https://phabricator.wikimedia.org/T437658) (owner: 10Ssingh) [19:58:39] (03PS1) 10Ssingh: admin: add derenrich to analytics-privatedata-users (krb) [puppet] - 10https://gerrit.wikimedia.org/r/1339911 (https://phabricator.wikimedia.org/T437668) [20:07:05] (03PS1) 10Aleksandar Mastilovic: Chart to install Blunderbuss 2.9.11 with GitLab CDN support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339917 [20:07:29] (03PS2) 10Aleksandar Mastilovic: Chart to install Blunderbuss 2.9.11 with GitLab CDN support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339917 (https://phabricator.wikimedia.org/T435292) [20:09:49] (03PS1) 10Dzahn: zuul: fix hiera data structure for zookeeper firewall node lookup [puppet] - 10https://gerrit.wikimedia.org/r/1339918 (https://phabricator.wikimedia.org/T435186) [20:10:25] (03CR) 10CI reject: [V:04-1] zuul: fix hiera data structure for zookeeper firewall node lookup [puppet] - 10https://gerrit.wikimedia.org/r/1339918 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [20:11:17] (03PS2) 10Dzahn: zuul: fix hiera data structure for zookeeper firewall node lookup [puppet] - 10https://gerrit.wikimedia.org/r/1339918 (https://phabricator.wikimedia.org/T435186) [20:13:16] (03CR) 10Dzahn: [C:03+2] zuul: fix hiera data structure for zookeeper firewall node lookup [puppet] - 10https://gerrit.wikimedia.org/r/1339918 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [20:40:26] (03CR) 10Pppery: "Reviewer-bot down?" [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1338245 (owner: 10Pppery) [20:41:38] (03PS1) 10Dzahn: zuul: refactor zookeeper inclusion to fix firewalling [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) [20:43:44] (03PS2) 10Dzahn: zuul: refactor zookeeper inclusion to fix firewalling [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) [20:45:01] (03CR) 10Dzahn: [V:03+1] "https://puppet-compiler.wmflabs.org/output/1339931/9410/zuul1001.eqiad.wmnet/index.html" [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [20:45:43] (03CR) 10Dzahn: [V:03+1] "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [20:48:44] (03CR) 10Dzahn: [V:03+1 C:03+2] zuul: refactor zookeeper inclusion to fix firewalling [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [20:49:19] 06SRE, 10Wikimedia-Mailing-lists: wikimediabe-l mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12313596 (10Aklapper) [20:50:22] 06SRE, 10Wikimedia-Mailing-lists: New WikiLatinos Mailing List - https://phabricator.wikimedia.org/T437742#12313597 (10Aklapper) Hi, do the two email addresses listed above belong to different people? [20:51:30] (03CR) 10Muehlenhoff: "https://gerrit.wikimedia.org/r/c/operations/puppet/+/1272766 will allow for the same" [puppet] - 10https://gerrit.wikimedia.org/r/1339931 (https://phabricator.wikimedia.org/T435186) (owner: 10Dzahn) [21:04:13] (03CR) 10Bking: [C:03+2] Chart to install Blunderbuss 2.9.11 with GitLab CDN support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1339917 (https://phabricator.wikimedia.org/T435292) (owner: 10Aleksandar Mastilovic) [21:09:28] (03CR) 10Dzahn: "oh, I was trying to fix this without touching zookeeper::firewall. wasn't aware you had started this. I think we are good with our own zuu" [puppet] - 10https://gerrit.wikimedia.org/r/1272766 (owner: 10Muehlenhoff) [21:11:42] (03CR) 10Dzahn: [C:03+2] "fixed with https://gerrit.wikimedia.org/r/c/operations/puppet/+/1339931" [puppet] - 10https://gerrit.wikimedia.org/r/1339123 (owner: 10Dzahn) [21:11:52] (03CR) 10Dzahn: [C:03+2] "fixed with https://gerrit.wikimedia.org/r/c/operations/puppet/+/1339931" [puppet] - 10https://gerrit.wikimedia.org/r/1326812 (https://phabricator.wikimedia.org/T435186) (owner: 10Hashar) [21:14:08] (03PS1) 10Aaron Schulz: restbase: make /media/math/ endpoints emit 'Deprecated' header [puppet] - 10https://gerrit.wikimedia.org/r/1339947 (https://phabricator.wikimedia.org/T431372) [21:21:00] (03PS2) 10Dzahn: admin: add anilk to analytics-privatedata-users, remove anil [puppet] - 10https://gerrit.wikimedia.org/r/1339129 (https://phabricator.wikimedia.org/T437611) (owner: 10Ssingh) [21:25:49] (03CR) 10Dzahn: [C:03+1] "would like to see it deployed so we can also get Hashar's patch merged which lives on top of this" [puppet] - 10https://gerrit.wikimedia.org/r/1322828 (https://phabricator.wikimedia.org/T424266) (owner: 10Scott French) [21:50:37] !log amastilovic@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply [21:51:55] !log amastilovic@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply [21:56:20] (03PS1) 10Dzahn: zuul: create location for backups of encryption keys [puppet] - 10https://gerrit.wikimedia.org/r/1339975 (https://phabricator.wikimedia.org/T436131) [22:02:37] (03PS1) 10Dzahn: backup/zuul: add backup set for zuul [puppet] - 10https://gerrit.wikimedia.org/r/1339983 (https://phabricator.wikimedia.org/T436131) [22:04:58] PROBLEM - statsv Varnishkafka log producer on cp5020 is CRITICAL: PROCS CRITICAL: 3 processes with args /usr/bin/varnishkafka -S /etc/varnishkafka/statsv.conf https://wikitech.wikimedia.org/wiki/Analytics/Systems/Varnishkafka [22:05:58] RECOVERY - statsv Varnishkafka log producer on cp5020 is OK: PROCS OK: 1 process with args /usr/bin/varnishkafka -S /etc/varnishkafka/statsv.conf https://wikitech.wikimedia.org/wiki/Analytics/Systems/Varnishkafka [22:07:04] (03PS1) 10Dzahn: create attribution.wikimedia.org for future miscweb static site [dns] - 10https://gerrit.wikimedia.org/r/1339986 (https://phabricator.wikimedia.org/T437635) [22:16:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:42:11] 06SRE, 10Wikimedia-Mailing-lists: New WikiLatinos Mailing List - https://phabricator.wikimedia.org/T437742#12313893 (10Oscar) Both are mine, you can include jelmarie.m.rodriguez@gmail.com as the other person [23:41:02] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1340047 [23:41:02] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1340047 (owner: 10TrainBranchBot) [23:48:50] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1340047 (owner: 10TrainBranchBot)