[00:02:29] (03CR) 10TrainBranchBot: [C:03+2] "Approved by musikanimal@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326373 (https://phabricator.wikimedia.org/T422148) (owner: 10MusikAnimal) [00:03:31] (03Merged) 10jenkins-bot: PersonalDashboard: add newly renamed *ReviewChangesMlModel setting [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326373 (https://phabricator.wikimedia.org/T422148) (owner: 10MusikAnimal) [00:04:24] !log musikanimal@deploy1003 Started scap sync-world: Backport for [[gerrit:1326373|PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] [00:04:29] T422148: Include edits to pages the user has recently edited in the PersonalDashboard Review Changes module - https://phabricator.wikimedia.org/T422148 [00:06:19] !log musikanimal@deploy1003 musikanimal: Backport for [[gerrit:1326373|PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [00:07:11] !log musikanimal@deploy1003 musikanimal: Continuing with deployment [00:09:32] !log eevans@cumin1003 END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-eqiad: Set storage compatability to NONE — T433026 - eevans@cumin1003 [00:09:37] T433026: Upgrade aqs cluster to Cassandra 5.0.8 & JDK 17 - https://phabricator.wikimedia.org/T433026 [00:11:31] !log musikanimal@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326373|PersonalDashboard: add newly renamed *ReviewChangesMlModel setting (T422148)]] (duration: 07m 07s) [00:11:36] T422148: Include edits to pages the user has recently edited in the PersonalDashboard Review Changes module - https://phabricator.wikimedia.org/T422148 [00:18:28] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:19:16] (03CR) 10Samwilson: "Yeah, that change didn't have any effect, but that's the point of *this* change. We could delete the code, but we also want to introduce t" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322507 (https://phabricator.wikimedia.org/T430528) (owner: 10Samwilson) [00:36:54] FIRING: [6x] CirrusSearchTitleSuggestIndexTooOld: Some search indices that power autocomplete have not been updated recently - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - TODO - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchTitleSuggestIndexTooOld [00:38:11] !log eevans@cumin1003 START - Cookbook sre.cassandra.roll-restart for nodes matching A:aqs-codfw: Set storage compatability to NONE — T433026 - eevans@cumin1003 [00:38:16] T433026: Upgrade aqs cluster to Cassandra 5.0.8 & JDK 17 - https://phabricator.wikimedia.org/T433026 [00:38:31] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12224558 (10ssingh) Hi @RobH: The DHL tracking says that the shipment is "on hold" since August 7. Do we have any idea what is happening here? Customs clearance? [00:39:19] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12224559 (10ssingh) Though it says "Clearance processing complete at MARSEILLE - FRANCE " which I am assuming means it went through that... [00:40:14] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12224560 (10Jhancock.wm) @Marostegui i finished with the firmware updates. tech support report sent to dell. waiting for reply. [00:41:11] FIRING: [3x] BFDdown: BFD session down between cr2-eqdfw and fe80::a6e1:1a00:1a6f:d3a3 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:43:37] FIRING: [4x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [00:45:39] FIRING: [4x] CoreBGPDown: Core BGP session down between cr1-magru and cr2-eqiad (195.200.68.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [01:04:48] (03PS1) 10Aaron Schulz: Mark all Mathoid-based endpoints as deprecated [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326410 (https://phabricator.wikimedia.org/T431372) [01:10:10] (03CR) 10RLazarus: [C:03+1] rest-gateway: Use envoy-future in staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326221 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [01:10:35] (03PS1) 10TrainBranchBot: Branch commit for wmf/1.47.0-wmf.16 [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326413 (https://phabricator.wikimedia.org/T430835) [01:10:38] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/1.47.0-wmf.16 [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326413 (https://phabricator.wikimedia.org/T430835) (owner: 10TrainBranchBot) [01:11:28] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1326414 [01:11:28] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1326414 (owner: 10TrainBranchBot) [01:14:59] !log eevans@cumin1003 END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:aqs-codfw: Set storage compatability to NONE — T433026 - eevans@cumin1003 [01:15:03] T433026: Upgrade aqs cluster to Cassandra 5.0.8 & JDK 17 - https://phabricator.wikimedia.org/T433026 [01:18:17] (03CR) 10BCornwall: Varnish: Reject non-thumb requests to thumbs (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [01:19:00] (03CR) 10BCornwall: [C:03+1] "Seems good to go by me, would love another +1 on it though." [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [01:19:39] (03CR) 10BCornwall: [C:03+1] "A few nits, but otherwise looks good with minimal regression capability for current infra." [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [01:21:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d8-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [01:22:03] (03Merged) 10jenkins-bot: Branch commit for wmf/1.47.0-wmf.16 [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326413 (https://phabricator.wikimedia.org/T430835) (owner: 10TrainBranchBot) [01:23:29] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1326414 (owner: 10TrainBranchBot) [01:47:33] (03PS1) 10Aaron Schulz: Add wmf-analytics-commons external module to commonswiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326425 (https://phabricator.wikimedia.org/T434927) [02:00:05] Deploy window Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T0200) [02:00:51] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:04:30] (03PS2) 10Scott French: Profile::Service_listener: Add sni_rewrites_host_header to 'split' [puppet] - 10https://gerrit.wikimedia.org/r/1325987 (https://phabricator.wikimedia.org/T427666) [02:04:31] (03PS3) 10Scott French: P:services_proxy::envoy: Fix split set_sni behavior [puppet] - 10https://gerrit.wikimedia.org/r/1325988 (https://phabricator.wikimedia.org/T427666) [02:04:32] (03PS4) 10Scott French: P:services_proxy::envoy: Add support for split host_regex [puppet] - 10https://gerrit.wikimedia.org/r/1325989 (https://phabricator.wikimedia.org/T427666) [02:04:34] (03PS4) 10Scott French: P:kubernetes::deployment_server::global_config: Extend services_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1325992 (https://phabricator.wikimedia.org/T427666) [02:05:09] (03PS4) 10Scott French: mesh.configuration: Add support for split host_regex in 1.15.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305514 (https://phabricator.wikimedia.org/T427666) [02:05:16] (03PS2) 10Scott French: mesh.configuration: Fix split listener stream idle timeout in 1.15.4 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325996 (https://phabricator.wikimedia.org/T427666) [02:07:37] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 46s) [02:08:13] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:18:13] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:43:20] (03CR) 10Krinkle: "Aye, good point. I assumed there was "stuff" in there besides the thing you recently added bit but it was close to empty, so makes sense." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322507 (https://phabricator.wikimedia.org/T430528) (owner: 10Samwilson) [02:48:49] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1163.eqiad.wmnet with OS bookworm [02:49:12] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1164.eqiad.wmnet with OS bookworm [02:49:33] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1165.eqiad.wmnet with OS bookworm [02:49:46] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1196.eqiad.wmnet with OS bookworm [02:50:12] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1197.eqiad.wmnet with OS bookworm [02:52:10] (03PS1) 10Krinkle: Fix "mathjax_ignore" handling around forcemathmode attribute [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326439 (https://phabricator.wikimedia.org/T434686) [02:53:23] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326439 (https://phabricator.wikimedia.org/T434686) (owner: 10Krinkle) [03:00:04] Deploy window Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T0300) [03:02:00] (03PS1) 10TrainBranchBot: testwikis to 1.47.0-wmf.16 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326440 (https://phabricator.wikimedia.org/T430835) [03:02:03] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by mwpresync@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326440 (https://phabricator.wikimedia.org/T430835) (owner: 10TrainBranchBot) [03:03:01] (03Merged) 10jenkins-bot: testwikis to 1.47.0-wmf.16 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326440 (https://phabricator.wikimedia.org/T430835) (owner: 10TrainBranchBot) [03:03:21] !log mwpresync@deploy1003 Started scap sync-world: testwikis to 1.47.0-wmf.16 refs T430835 [03:03:26] T430835: 1.47.0-wmf.16 deployment blockers - https://phabricator.wikimedia.org/T430835 [03:04:18] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage [03:04:27] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage [03:04:34] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage [03:04:59] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage [03:05:13] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage [03:09:00] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1163.eqiad.wmnet with reason: host reimage [03:12:02] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1164.eqiad.wmnet with reason: host reimage [03:15:11] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1196.eqiad.wmnet with reason: host reimage [03:18:43] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1165.eqiad.wmnet with reason: host reimage [03:22:26] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1197.eqiad.wmnet with reason: host reimage [03:25:16] PROBLEM - Host titan1002 is DOWN: PING CRITICAL - Packet loss = 60%, RTA = 3853.04 ms [03:26:52] (03CR) 10Samwilson: "The idea is that as no video formats are specified in [ALLOWED_CONVERSIONS](https://gerrit.wikimedia.org/r/plugins/gitiles/operations/soft" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322507 (https://phabricator.wikimedia.org/T430528) (owner: 10Samwilson) [03:28:21] FIRING: [4x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [03:28:48] RECOVERY - Host titan1002 is UP: PING OK - Packet loss = 0%, RTA = 0.29 ms [03:30:30] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1163.eqiad.wmnet with OS bookworm [03:33:21] FIRING: [4x] ProbeDown: Service ganeti2046:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [03:35:11] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1196.eqiad.wmnet with OS bookworm [03:36:15] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1164.eqiad.wmnet with OS bookworm [03:37:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:38:03] !log mwpresync@deploy1003 Finished scap sync-world: testwikis to 1.47.0-wmf.16 refs T430835 (duration: 34m 43s) [03:38:09] T430835: 1.47.0-wmf.16 deployment blockers - https://phabricator.wikimedia.org/T430835 [03:41:45] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1165.eqiad.wmnet with OS bookworm [03:42:55] (03PS1) 10Samwilson: InitialiseSettings.php: Enable Bulk OCR on pawikisource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326450 (https://phabricator.wikimedia.org/T434648) [03:45:30] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1197.eqiad.wmnet with OS bookworm [03:47:42] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326450 (https://phabricator.wikimedia.org/T434648) (owner: 10Samwilson) [04:00:04] Deploy window Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T0400) [04:02:29] !log mwpresync@deploy1003 Pruned MediaWiki: 1.47.0-wmf.13 (duration: 02m 23s) [04:03:28] (03CR) 10Tim Starling: [C:03+1] InitialiseSettings.php: Enable Bulk OCR on pawikisource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326450 (https://phabricator.wikimedia.org/T434648) (owner: 10Samwilson) [04:14:20] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1144.eqiad.wmnet with OS bookworm [04:29:15] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage [04:33:26] (03PS5) 10BryanDavis: httpbb: Add test suite for testwiki [puppet] - 10https://gerrit.wikimedia.org/r/1323829 (https://phabricator.wikimedia.org/T428972) [04:33:57] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1144.eqiad.wmnet with reason: host reimage [04:36:54] FIRING: [6x] CirrusSearchTitleSuggestIndexTooOld: Some search indices that power autocomplete have not been updated recently - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - TODO - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchTitleSuggestIndexTooOld [04:41:26] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [04:43:37] FIRING: [4x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [04:45:55] FIRING: [4x] CoreBGPDown: Core BGP session down between cr1-magru and cr2-eqiad (195.200.68.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [04:47:47] (03CR) 10TChin: [C:03+1] alerts: webrequest-pageview webrequest-pageview is a new application that needs alerts. Here are new alerts defined for when the application [alerts] - 10https://gerrit.wikimedia.org/r/1324336 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [04:48:02] (03CR) 10TChin: [C:03+1] pageview-trending-relative: Flink app alerts [alerts] - 10https://gerrit.wikimedia.org/r/1325432 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [04:48:18] (03CR) 10TChin: [C:03+1] relative-trending: Alerts for missing data Adding 3 alerts to check if Kafka topics are receiving data or not. [alerts] - 10https://gerrit.wikimedia.org/r/1326241 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [04:57:55] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1144.eqiad.wmnet with OS bookworm [05:01:11] FIRING: [3x] BFDdown: BFD session down between cr1-esams and fe80::6687:8807:df2:7018 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [05:06:07] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1155.eqiad.wmnet with OS bookworm [05:21:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d8-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [05:23:10] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage [05:28:34] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1155.eqiad.wmnet with reason: host reimage [05:38:04] !log updating prometheusBearerToken on gerrit [05:38:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:43:51] FIRING: Transit, peering or transport OUT traffic above 90% capacity - cr2-eqsin:xe-0/1/6 (Peering: BBIX (01130-SG1-X-I-0001)) #page - https://w.wiki/Gbyf - https://grafana.wikimedia.org/d/d968a627-b6f6-47fc-9316-e058854a4945/throughput-network-device-interfaces?var-site=eqsin+prometheus%2Fops&var-device=cr2-eqsin:9804&var-interface=xe-0%2F1%2F6 - https://alerts.wikimedia.org/?q=alertname%3DTransitPeeringTransportOutSaturation [05:46:35] !incidents [05:46:35] 8251 (UNACKED) TransitPeeringTransportOutSaturation network sre (cr2-eqsin:9804 Peering: BBIX (01130-SG1-X-I-0001) xe-0/1/6 gnmi eqsin) [05:46:35] 8250 (RESOLVED) OutboundMXQueueHigh sre (mx-out1001:9154 eqiad) [05:46:49] !ack 8251 [05:46:49] 8251 (ACKED) TransitPeeringTransportOutSaturation network sre (cr2-eqsin:9804 Peering: BBIX (01130-SG1-X-I-0001) xe-0/1/6 gnmi eqsin) [05:49:51] (03Restored) 10Samwilson: Remove defunct feature flag $wgWikisourceEnableOcr [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) (owner: 10Samwilson) [05:50:01] (03PS2) 10Samwilson: InitialiseSettings: Remove redundant feature flag $wgWikisourceEnableOcr [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) [05:50:17] here too jayme, checking [05:50:44] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1155.eqiad.wmnet with OS bookworm [05:50:48] o/ [05:51:07] any joy so far ? [05:51:37] na, just came here [05:51:44] (03CR) 10Ayounsi: [C:03+1] "I'm not expert neither but that change also looks fine to me!" [software/homer] - 10https://gerrit.wikimedia.org/r/1325453 (https://phabricator.wikimedia.org/T434773) (owner: 10Cathal Mooney) [05:51:54] ack [05:52:59] godog: spike from eqsin and esams https://grafana-rw.wikimedia.org/d/7d07a703-4ea9-40b8-b9ae-ae218e53ff15/cdn-upload-square-one?from=2026-08-17T15:16:37.945Z&to=2026-08-18T05:51:58.536Z&timezone=utc&orgId=1&var-site=eqiad&var-site=codfw&var-deployment=mw-web&var-percentile=50&var-cron=.%2A&var-cluster=upload&var-pop=drmrs&var-pop=esams&var-pop=ulsfo&var-pop=magru&var-pop=eqsin&var-http_codes=4&var-http_codes=5&var-http_codes=2&var- [05:53:01] backend=$__all&viewPanel=panel-56 [05:53:22] hmpf https://grafana-rw.wikimedia.org/goto/efviyc3j2oohsa?orgId=default [05:53:28] (03PS3) 10Samwilson: InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) [05:53:47] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) (owner: 10Samwilson) [05:53:59] indeed [05:56:11] fyi, there seems to be a monitoring but where we don't have gNMI data for cr3-eqsin - https://grafana.wikimedia.org/goto/cfviyjtyeqqrkc?orgId=default [05:56:35] s/but/bug/ [05:57:38] (03PS1) 10Arnaudb: gerrit: meaningless change [puppet] - 10https://gerrit.wikimedia.org/r/1326669 [05:57:40] (03CR) 10Arnaudb: [C:03+2] gerrit: meaningless change [puppet] - 10https://gerrit.wikimedia.org/r/1326669 (owner: 10Arnaudb) [05:58:30] (03PS1) 10Kevin Bazira: ml-services: deploy qwen36-27b and qwen3-14b isvcs with support for structured outputs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326670 (https://phabricator.wikimedia.org/T434059) [05:58:51] (03CR) 10Krinkle: "Got it. I assumed it disallowed anything not specifically set. My bad." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322507 (https://phabricator.wikimedia.org/T430528) (owner: 10Samwilson) [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T0600) [06:00:05] marostegui, Amir1, and federico3: That opportune time for a Primary database switchover deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T0600). [06:00:09] !log arnaudb@cumin1003 START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003 [06:02:12] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003 [06:03:21] 06SRE, 06Infrastructure-Foundations, 10netops, 06Traffic, 13Patch-For-Review: BFD session fails from Anycast hosts over IPv6 on boot - https://phabricator.wikimedia.org/T434806#12224818 (10ayounsi) We should reach out to the BIRD team, even if there is a race condition, BIRD's BFD probably needs to come... [06:04:48] !log arnaudb@cumin1003 START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit1003 [06:06:08] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit1003 [06:06:12] !log arnaudb@cumin1003 START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2002 [06:06:25] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:08:13] FIRING: [5x] JobUnavailable: Reduced availability for job gerrit in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:09:00] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2002 [06:11:36] (03PS1) 10PipelineBot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326687 [06:13:13] FIRING: [9x] JobUnavailable: Reduced availability for job gerrit in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:13:27] it's from the token rotation ↑ [06:13:46] shoudl recover with the next puppet run [06:18:13] FIRING: [9x] JobUnavailable: Reduced availability for job gerrit in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:20:55] (03CR) 10Arthur taylor: [C:03+1] "Makes sense to me - let's do it" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326307 (owner: 10Lucas Werkmeister (WMDE)) [06:23:13] FIRING: [9x] JobUnavailable: Reduced availability for job gerrit in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:23:51] RESOLVED: Transit, peering or transport OUT traffic above 90% capacity - cr2-eqsin:xe-0/1/6 (Peering: BBIX (01130-SG1-X-I-0001)) #page - https://w.wiki/Gbyf - https://grafana.wikimedia.org/d/d968a627-b6f6-47fc-9316-e058854a4945/throughput-network-device-interfaces?var-site=eqsin+prometheus%2Fops&var-device=cr2-eqsin:9804&var-interface=xe-0%2F1%2F6 - https://alerts.wikimedia.org/?q=alertname%3DTransitPeeringTransportOutSaturation [06:25:44] (03CR) 10Ayounsi: [C:03+2] RANCID: ignore Fan Speed in SR-Linux [puppet] - 10https://gerrit.wikimedia.org/r/1322761 (owner: 10Ayounsi) [06:41:29] !log filippo@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1054.eqiad.wmnet with OS trixie [06:41:43] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224869 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt1054.eqiad.wmnet with OS trixie [06:43:53] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1198.eqiad.wmnet with OS bookworm [06:44:52] !log upgrade eqsin gnmic to 0.47.0 [06:44:54] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:44:54] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1194.eqiad.wmnet with OS bookworm [06:48:19] (03CR) 10Tim Starling: [C:03+1] InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) (owner: 10Samwilson) [06:52:18] !log filippo@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1054.eqiad.wmnet with OS trixie [06:52:28] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224896 (10fgiunchedi) >>! In T431682#12223002, @VRiley-WMF wrote: > Hey @fgiunchedi it seems as though that cloudvirt1054, cloudvirt1055, cloudvirt1056, cloudvir... [06:52:31] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224897 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1054.eqiad.wmnet with OS trixie executed with e... [06:53:11] !log filippo@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1055.eqiad.wmnet with OS trixie [06:53:26] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224898 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt1055.eqiad.wmnet with OS trixie [06:58:15] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage [06:59:22] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224919 (10fgiunchedi) cloudvirt1055 can't find a network device to boot from: ` Booting from HTTP Device 1: Embedded NIC 1 Port 1 Partition 1 HTTP: No media det... [06:59:31] !log filippo@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1055.eqiad.wmnet with OS trixie [06:59:44] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224921 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1055.eqiad.wmnet with OS trixie executed with e... [07:00:00] !log filippo@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie [07:00:05] Amir1, urbanecm, and awight: I seem to be stuck in Groundhog week. Sigh. Time for (yet another) UTC morning backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T0700). [07:00:05] Krinkle and samwilson: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:14] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224922 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt1056.eqiad.wmnet with OS trixie [07:00:18] !log filippo@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie [07:00:30] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224923 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1056.eqiad.wmnet with OS trixie executed with e... [07:01:07] !log filippo@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1056.eqiad.wmnet with OS trixie [07:01:23] !log filippo@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1056.eqiad.wmnet with OS trixie [07:01:23] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224926 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt1056.eqiad.wmnet with OS trixie [07:01:32] o/ [07:01:37] hullo [07:01:37] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224927 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1056.eqiad.wmnet with OS trixie executed with e... [07:02:19] samwilson: feel free to go first [07:02:43] ok sure, thanks :) [07:02:54] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1198.eqiad.wmnet with reason: host reimage [07:03:03] (03PS2) 10Elukey: docker_registry: route /v2/releng to its new s3 backend [puppet] - 10https://gerrit.wikimedia.org/r/1326323 (https://phabricator.wikimedia.org/T432829) [07:03:27] (03CR) 10Elukey: docker_registry: route /v2/releng to its new s3 backend (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1326323 (https://phabricator.wikimedia.org/T432829) (owner: 10Elukey) [07:04:06] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samwilson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326450 (https://phabricator.wikimedia.org/T434648) (owner: 10Samwilson) [07:05:00] (03Merged) 10jenkins-bot: InitialiseSettings.php: Enable Bulk OCR on pawikisource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326450 (https://phabricator.wikimedia.org/T434648) (owner: 10Samwilson) [07:05:09] (03CR) 10Bartosz Wójtowicz: [C:03+1] ml-services: deploy qwen36-27b and qwen3-14b isvcs with support for structured outputs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326670 (https://phabricator.wikimedia.org/T434059) (owner: 10Kevin Bazira) [07:05:21] (03CR) 10Muehlenhoff: [C:03+1] "Looks good. There might still be subtle issues, so best to test in Pontoon and alternatively on a single puppetserver first." [puppet] - 10https://gerrit.wikimedia.org/r/1296495 (https://phabricator.wikimedia.org/T420184) (owner: 10Arnaudb) [07:05:43] !log samwilson@deploy1003 Started scap sync-world: Backport for [[gerrit:1326450|InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] [07:05:48] T434648: Deploy Bulk OCR on pawikisource - https://phabricator.wikimedia.org/T434648 [07:06:57] (03PS1) 10Aqu: airflow: release-scope extra_k8s_secrets in devenv [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326691 [07:07:21] (03CR) 10Kevin Bazira: [C:03+2] ml-services: deploy qwen36-27b and qwen3-14b isvcs with support for structured outputs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326670 (https://phabricator.wikimedia.org/T434059) (owner: 10Kevin Bazira) [07:07:47] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224932 (10fgiunchedi) cloudvirt1056 initially failed the reimage when resetting a system that's off: ` Resetting chassis power status for cloudvirt1056 to Force... [07:07:58] !log filippo@cumin1003 START - Cookbook sre.hosts.reimage for host cloudvirt1057.eqiad.wmnet with OS trixie [07:09:37] (03Merged) 10jenkins-bot: ml-services: deploy qwen36-27b and qwen3-14b isvcs with support for structured outputs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326670 (https://phabricator.wikimedia.org/T434059) (owner: 10Kevin Bazira) [07:10:05] !log samwilson@deploy1003 samwilson: Backport for [[gerrit:1326450|InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:11:02] !log jmm@cumin2003 START - Cookbook sre.hosts.reimage for host ganeti2046.codfw.wmnet with OS bookworm [07:11:07] !log samwilson@deploy1003 samwilson: Continuing with deployment [07:11:16] 10ops-codfw, 06SRE, 06DC-Ops: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12224941 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin2003 for host ganeti2046.codfw.wmnet with OS bookworm [07:11:47] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti1048.eqiad.wmnet [07:13:29] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224944 (10fgiunchedi) And cloudvirt1057 also doesn't boot from the network: ` Booting from HTTP Device 1: Embedded NIC 1 Port 1 Partition 1 HTTP: No media detec... [07:14:50] RECOVERY - Host ganeti2046 is UP: PING OK - Packet loss = 0%, RTA = 30.32 ms [07:16:24] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224945 (10fgiunchedi) @VRiley-WMF to recap my findings: * cloudvirt1055 / cloudvirt1056 / cloudvirt1057 all don't find the network to boot * cloudvirt1054 has a... [07:16:55] jmm@cumin2003 drain-node (PID 2686880) is awaiting input [07:17:00] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [07:18:00] !log samwilson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326450|InitialiseSettings.php: Enable Bulk OCR on pawikisource (T434648)]] (duration: 12m 17s) [07:18:05] T434648: Deploy Bulk OCR on pawikisource - https://phabricator.wikimedia.org/T434648 [07:19:28] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samwilson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) (owner: 10Samwilson) [07:19:39] (03CR) 10CI reject: [V:04-1] InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) (owner: 10Samwilson) [07:20:03] (03PS28) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [07:21:33] (03PS4) 10Samwilson: InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) [07:22:32] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samwilson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) (owner: 10Samwilson) [07:23:19] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1198.eqiad.wmnet with OS bookworm [07:23:29] (03Merged) 10jenkins-bot: InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr [mediawiki-config] - 10https://gerrit.wikimedia.org/r/701016 (https://phabricator.wikimedia.org/T285311) (owner: 10Samwilson) [07:23:46] !log samwilson@deploy1003 Started scap sync-world: Backport for [[gerrit:701016|InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] [07:23:53] T285311: Enable OCR improvements on all remaining Wikisources - https://phabricator.wikimedia.org/T285311 [07:24:00] !log jmm@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage [07:24:12] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti1048.eqiad.wmnet [07:24:37] 06SRE, 06Data-Engineering, 10Kafka-Infrastructure, 06serviceops-radar, and 3 others: Configuration Management for Kafka settings - https://phabricator.wikimedia.org/T276088#12224954 (10RKemper) PoC is working end-to-end in pontoon. Working on documentation for spin-up instructions, but in meantime the most... [07:25:00] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T435157 [07:25:54] !log samwilson@deploy1003 samwilson: Backport for [[gerrit:701016|InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:25:55] PROBLEM - Host ml-etcd1002 is DOWN: PING CRITICAL - Packet loss = 100% [07:26:05] PROBLEM - Host dse-k8s-etcd1003 is DOWN: PING CRITICAL - Packet loss = 100% [07:27:18] !log samwilson@deploy1003 samwilson: Continuing with deployment [07:28:32] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ganeti2046.codfw.wmnet with reason: host reimage [07:29:29] !log add gnmic 0.47.0 to bookworm and trixie reprepro [07:29:31] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:29:44] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti1048.eqiad.wmnet [07:30:13] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti1048.eqiad.wmnet [07:30:32] RECOVERY - Host dse-k8s-etcd1003 is UP: PING OK - Packet loss = 0%, RTA = 0.22 ms [07:30:48] RECOVERY - Host ml-etcd1002 is UP: PING OK - Packet loss = 0%, RTA = 0.25 ms [07:31:33] !log samwilson@deploy1003 Finished scap sync-world: Backport for [[gerrit:701016|InitialiseSettings and -labs: Remove redundant feature flag $wgWikisourceEnableOcr (T285311)]] (duration: 07m 47s) [07:31:38] T285311: Enable OCR improvements on all remaining Wikisources - https://phabricator.wikimedia.org/T285311 [07:31:50] Hey I have a patch for backport that I forgot to add in the calendar, can i do it now? [07:32:03] *add it now [07:32:47] (03CR) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [07:33:36] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:34:50] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - T435157 [07:35:44] Krinkle, all yours now [07:35:55] thx [07:35:57] ryankemper@cumin2003 reimage (PID 2681273) is awaiting input [07:36:27] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326439 (https://phabricator.wikimedia.org/T434686) (owner: 10Krinkle) [07:37:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:37:47] (03Merged) 10jenkins-bot: Fix "mathjax_ignore" handling around forcemathmode attribute [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326439 (https://phabricator.wikimedia.org/T434686) (owner: 10Krinkle) [07:38:12] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1326439|Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] [07:38:17] T434686: Fix forcemathmode attribute - https://phabricator.wikimedia.org/T434686 [07:38:22] FIRING: [4x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:38:58] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T435157 [07:40:13] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1326439|Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:42:42] (03PS2) 10AOkoth: phabricator: fix php version in restart script [puppet] - 10https://gerrit.wikimedia.org/r/1326292 (https://phabricator.wikimedia.org/T433988) [07:42:53] (03CR) 10Slyngshede: [C:03+2] varnish: Configure text for thumbnails [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [07:43:22] FIRING: [4x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [07:44:35] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ganeti2046.codfw.wmnet with OS bookworm [07:44:49] 10ops-codfw, 06SRE, 06DC-Ops: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12224986 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin2003 for host ganeti2046.codfw.wmnet with OS bookworm completed: - ganeti2046 (**WARN**) - Downtimed o... [07:45:39] FIRING: [4x] CoreBGPDown: Core BGP session down between cr1-magru and cr2-eqiad (195.200.68.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [07:47:14] !log krinkle@deploy1003 krinkle: Continuing with deployment [07:48:53] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - T435157 [07:50:34] (03CR) 10AOkoth: phabricator: fix php version in restart script (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1326292 (https://phabricator.wikimedia.org/T433988) (owner: 10AOkoth) [07:51:15] (03CR) 10AOkoth: "Looks good now from the output linked below:" [puppet] - 10https://gerrit.wikimedia.org/r/1326292 (https://phabricator.wikimedia.org/T433988) (owner: 10AOkoth) [07:51:27] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326439|Fix "mathjax_ignore" handling around forcemathmode attribute (T434686)]] (duration: 13m 15s) [07:51:35] T434686: Fix forcemathmode attribute - https://phabricator.wikimedia.org/T434686 [07:51:51] (03PS29) 10Slyngshede: Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) [07:52:33] Krinkle: do you have more to deploy ? [07:53:07] (03PS8) 10Dreamy Jazz: Register the mediawiki.wikimedia_antiabuse.content_policy_score stream [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1321579 (https://phabricator.wikimedia.org/T432848) (owner: 10Mpostoronca) [07:54:02] maybe i can fit one more backport before the train rollout? andre brennen ? [07:54:58] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T435157 [07:55:22] (03PS1) 10Marostegui: db1289: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1326697 [07:57:02] (03CR) 10Federico Ceratto: [C:03+2] Add MariaDB test-s8 section VMs [puppet] - 10https://gerrit.wikimedia.org/r/1326316 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [08:00:03] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326232 (https://phabricator.wikimedia.org/T435115) (owner: 10Jgiannelos) [08:00:05] (03CR) 10Marostegui: [C:03+2] db1289: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1326697 (owner: 10Marostegui) [08:00:15] nvm i put it for the next window [08:00:31] federico3: you can merge my patch [08:02:09] looking [08:02:38] federico3: what I mean, feel free to merge my patch when running puppet-merge as your patch is still there [08:03:00] (03PS1) 10Marostegui: instances.yaml: Add db1289 [puppet] - 10https://gerrit.wikimedia.org/r/1326719 (https://phabricator.wikimedia.org/T407942) [08:03:15] yep, I'm merging it with mine [08:03:35] thanks! [08:03:59] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1289 [puppet] - 10https://gerrit.wikimedia.org/r/1326719 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [08:05:32] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1289 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96145 and previous config saved to /var/cache/conftool/dbconfig/20260818-080531-marostegui.json [08:05:38] T407942: Productionize db12[65-90] - https://phabricator.wikimedia.org/T407942 [08:05:53] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1289: Pool back [08:06:47] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2035.codfw.wmnet [08:08:20] (03PS1) 10Marostegui: db1288: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1326773 [08:10:13] nemo-yiannis: hi, sorry, I was a bit late for today. thanks for the ping [08:10:15] (03CR) 10Marostegui: [C:03+2] db1288: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1326773 (owner: 10Marostegui) [08:11:27] jmm@cumin2003 drain-node (PID 2713253) is awaiting input [08:11:53] aokoth@cumin1003 aokoth: The backup on gitlab1004 is complete, ready to proceed with upgrade. [08:13:33] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2035.codfw.wmnet [08:14:03] (03CR) 10Btullis: [C:03+2] airflow: release-scope extra_k8s_secrets in devenv [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326691 (owner: 10Aqu) [08:14:21] (03PS1) 10Marostegui: instances.yaml: Add db1288 [puppet] - 10https://gerrit.wikimedia.org/r/1326774 (https://phabricator.wikimedia.org/T407942) [08:15:18] PROBLEM - Host ml-staging-etcd2003 is DOWN: PING CRITICAL - Packet loss = 100% [08:16:58] (03PS1) 10STran: SI: Instrument case update on first edit [extensions/CheckUser] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326775 (https://phabricator.wikimedia.org/T435048) [08:17:07] (03Merged) 10jenkins-bot: airflow: release-scope extra_k8s_secrets in devenv [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326691 (owner: 10Aqu) [08:17:09] nemo-yiannis: if you're still around you could deploy your backport; train is blocked [08:17:26] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [extensions/CheckUser] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326775 (https://phabricator.wikimedia.org/T435048) (owner: 10STran) [08:18:42] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2035.codfw.wmnet [08:18:53] !log filippo@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudvirt1057.eqiad.wmnet with OS trixie [08:18:53] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2035.codfw.wmnet [08:20:26] RECOVERY - Host ml-staging-etcd2003 is UP: PING WARNING - Packet loss = 75%, RTA = 30.66 ms [08:20:30] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1288 [puppet] - 10https://gerrit.wikimedia.org/r/1326774 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [08:20:49] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2036.codfw.wmnet [08:22:35] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1288 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96148 and previous config saved to /var/cache/conftool/dbconfig/20260818-082234-marostegui.json [08:22:40] T407942: Productionize db12[65-90] - https://phabricator.wikimedia.org/T407942 [08:23:10] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2036.codfw.wmnet [08:23:33] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1181.eqiad.wmnet with OS bookworm [08:23:44] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T435157 [08:25:15] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1288: Pool back [08:25:25] (03CR) 10Ayounsi: [C:03+1] "looks good, thanks!" [cookbooks] - 10https://gerrit.wikimedia.org/r/1324839 (https://phabricator.wikimedia.org/T434737) (owner: 10Ryan Kemper) [08:27:45] (03PS1) 10Marostegui: db1287: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1326776 [08:28:30] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2036.codfw.wmnet [08:28:37] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2036.codfw.wmnet [08:28:52] (03CR) 10Marostegui: [C:03+2] db1287: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1326776 (owner: 10Marostegui) [08:29:00] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet [08:29:49] (03CR) 10Arnaudb: [C:03+1] "lgtm! thanks for the fix :-)" [puppet] - 10https://gerrit.wikimedia.org/r/1326292 (https://phabricator.wikimedia.org/T433988) (owner: 10AOkoth) [08:30:32] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2037.codfw.wmnet [08:30:49] (03PS1) 10Marostegui: instances.yaml: Add db1287 [puppet] - 10https://gerrit.wikimedia.org/r/1326777 (https://phabricator.wikimedia.org/T407942) [08:30:51] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet [08:31:03] !log filippo@cumin1003 START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet [08:31:37] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1287 [puppet] - 10https://gerrit.wikimedia.org/r/1326777 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [08:32:50] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet [08:33:12] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1287 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96150 and previous config saved to /var/cache/conftool/dbconfig/20260818-083311-marostegui.json [08:33:17] T407942: Productionize db12[65-90] - https://phabricator.wikimedia.org/T407942 [08:33:34] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1287: Pool back [08:34:01] (03CR) 10Slyngshede: [C:03+1] "Looks good" [software/bitu] - 10https://gerrit.wikimedia.org/r/1320185 (owner: 10Perryprog) [08:34:46] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet [08:35:05] !log filippo@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet [08:35:13] (03PS1) 10Marostegui: mariadb: Productionize db1283 [puppet] - 10https://gerrit.wikimedia.org/r/1326778 (https://phabricator.wikimedia.org/T407942) [08:36:54] FIRING: [6x] CirrusSearchTitleSuggestIndexTooOld: Some search indices that power autocomplete have not been updated recently - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - TODO - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchTitleSuggestIndexTooOld [08:37:23] jmm@cumin2003 drain-node (PID 2718349) is awaiting input [08:39:09] 06SRE, 10SRE-Access-Requests: Requesting access to for - https://phabricator.wikimedia.org/T435168 (10Nicholusmuwonge_wmde) 03NEW [08:43:22] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [08:47:09] 06SRE, 10SRE-Access-Requests: Requesting access to 'restricted' for nicholusmuwonge - https://phabricator.wikimedia.org/T435168#12225245 (10Nicholusmuwonge_wmde) [08:48:00] andre: I think it's ok to deploy it in the next window [08:48:21] ok :) [08:51:02] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1289: Pool back [08:52:10] 06SRE, 10SRE-Access-Requests: Requesting access to 'restricted' for nicholusmuwonge - https://phabricator.wikimedia.org/T435168#12225280 (10Tobi_WMDE_SW) I'm endorsing this request. [08:52:24] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2037.codfw.wmnet [08:54:04] PROBLEM - Host dse-k8s-ctrl2001 is DOWN: PING CRITICAL - Packet loss = 100% [08:54:04] PROBLEM - Host dse-k8s-etcd2001 is DOWN: PING CRITICAL - Packet loss = 100% [08:54:34] PROBLEM - Host ml-etcd2002 is DOWN: PING CRITICAL - Packet loss = 100% [08:55:30] RECOVERY - Host dse-k8s-ctrl2001 is UP: PING OK - Packet loss = 0%, RTA = 31.79 ms [08:55:30] RECOVERY - Host dse-k8s-etcd2001 is UP: PING OK - Packet loss = 0%, RTA = 30.53 ms [08:55:48] RECOVERY - Host ml-etcd2002 is UP: PING OK - Packet loss = 0%, RTA = 34.97 ms [08:56:09] !log filippo@cumin1003 START - Cookbook sre.mysql.multiinstance_reboot for 10 hosts [08:57:47] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2037.codfw.wmnet [08:57:55] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org [08:57:55] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2037.codfw.wmnet [08:58:36] (03CR) 10JavierMonton: [C:03+2] relative-trending: Alerts for missing data Adding 3 alerts to check if Kafka topics are receiving data or not. [alerts] - 10https://gerrit.wikimedia.org/r/1326241 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [08:58:39] (03CR) 10JavierMonton: [C:03+2] pageview-trending-relative: Flink app alerts [alerts] - 10https://gerrit.wikimedia.org/r/1325432 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [08:58:42] (03CR) 10JavierMonton: [C:03+2] alerts: webrequest-pageview webrequest-pageview is a new application that needs alerts. Here are new alerts defined for when the application [alerts] - 10https://gerrit.wikimedia.org/r/1324336 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [09:01:19] (03Merged) 10jenkins-bot: relative-trending: Alerts for missing data Adding 3 alerts to check if Kafka topics are receiving data or not. [alerts] - 10https://gerrit.wikimedia.org/r/1326241 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [09:01:21] (03Merged) 10jenkins-bot: pageview-trending-relative: Flink app alerts [alerts] - 10https://gerrit.wikimedia.org/r/1325432 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [09:01:25] (03Merged) 10jenkins-bot: alerts: webrequest-pageview webrequest-pageview is a new application that needs alerts. Here are new alerts defined for when the application is failing, lagging, or losing records for any unexpected reason. [alerts] - 10https://gerrit.wikimedia.org/r/1324336 (https://phabricator.wikimedia.org/T431536) (owner: 10JavierMonton) [09:01:26] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:03:23] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2038.codfw.wmnet [09:04:08] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org [09:10:24] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1288: Pool back [09:11:10] jmm@cumin2003 drain-node (PID 2724135) is awaiting input [09:11:20] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org [09:13:15] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet [09:13:17] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [09:13:22] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2038.codfw.wmnet [09:14:16] !log slyngshede@cumin1003 START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org [09:14:50] (03CR) 10Tiziano Fogli: [C:03+2] grafana/image-renderer: remove GIR plugin [debs/grafana-plugins] - 10https://gerrit.wikimedia.org/r/1326295 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [09:14:54] (03CR) 10Tiziano Fogli: [V:03+2 C:03+2] grafana/image-renderer: remove GIR plugin [debs/grafana-plugins] - 10https://gerrit.wikimedia.org/r/1326295 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [09:15:02] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.dns.netbox (exit_code=97) [09:15:09] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db1901.eqiad.wmnet [09:15:54] PROBLEM - Host aux-k8s-etcd2005 is DOWN: PING CRITICAL - Packet loss = 100% [09:17:27] btullis@cumin1003 reimage (PID 661894) is awaiting input [09:17:29] !log fceratto@cumin1003 START - Cookbook sre.hosts.decommission for hosts db1901.eqiad.wmnet [09:18:12] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org [09:18:24] !log slyngshede@cumin1003 START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org [09:18:41] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2038.codfw.wmnet [09:18:44] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1287: Pool back [09:19:22] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2038.codfw.wmnet [09:19:36] (03PS1) 10Slyngshede: IDP: Switch-over [dns] - 10https://gerrit.wikimedia.org/r/1326784 [09:19:44] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2039.codfw.wmnet [09:20:32] RECOVERY - Host aux-k8s-etcd2005 is UP: PING OK - Packet loss = 0%, RTA = 30.58 ms [09:21:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d8-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [09:21:39] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [09:22:21] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org [09:22:42] !log slyngshede@cumin1003 START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org [09:22:52] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2039.codfw.wmnet [09:25:14] PROBLEM - Host kubestagemaster2004 is DOWN: PING CRITICAL - Packet loss = 100% [09:25:46] RECOVERY - Host kubestagemaster2004 is UP: PING OK - Packet loss = 0%, RTA = 32.13 ms [09:26:40] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org [09:26:46] (03CR) 10Slyngshede: [C:03+2] IDP: Switch-over [dns] - 10https://gerrit.wikimedia.org/r/1326784 (owner: 10Slyngshede) [09:27:04] !log slyngshede@dns1004 START - running authdns-update [09:27:10] fceratto@cumin1003 decommission (PID 672933) is awaiting input [09:28:14] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2039.codfw.wmnet [09:28:20] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2039.codfw.wmnet [09:29:08] !log slyngshede@dns1004 END - running authdns-update [09:29:43] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2040.codfw.wmnet [09:31:13] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [09:31:22] PROBLEM - SSH on bast4006 is CRITICAL: connect to address 198.35.26.104 and port 22: Connection refused https://wikitech.wikimedia.org/wiki/SSH/monitoring [09:31:32] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1901.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003" [09:31:32] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:31:33] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1901.eqiad.wmnet [09:31:45] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet [09:31:47] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [09:36:12] jmm@cumin2003 drain-node (PID 2729515) is awaiting input [09:36:53] !log slyngshede@cumin1003 START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org [09:37:20] fceratto@cumin1003 makevm (PID 684667) is awaiting input [09:40:53] !log slyngshede@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org [09:41:09] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2040.codfw.wmnet [09:41:31] i'm planning a no-build deployment in about 50 minutes (during the infra window) to roll out https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1313111 [09:41:48] if there are any concerns, please let me know :) [09:42:49] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" [09:42:54] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" [09:42:54] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:42:54] !log fceratto@cumin1003 START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors [09:42:58] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors [09:46:29] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2040.codfw.wmnet [09:46:35] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2040.codfw.wmnet [09:48:49] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host people1005.eqiad.wmnet [09:49:06] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host people2004.codfw.wmnet [09:49:22] RECOVERY - SSH on bast4006 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [09:49:26] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host releases1003.eqiad.wmnet [09:49:40] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host releases2003.codfw.wmnet [09:50:46] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases2003.codfw.wmnet [09:50:51] (03CR) 10Marostegui: [C:03+2] mariadb: Productionize db1283 [puppet] - 10https://gerrit.wikimedia.org/r/1326778 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [09:52:03] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy7002.magru.wmnet [09:52:42] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people1005.eqiad.wmnet [09:53:06] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host people2004.codfw.wmnet [09:53:15] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy7001.magru.wmnet [09:53:24] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host releases1003.eqiad.wmnet [09:53:39] jouncebot: now [09:53:39] For the next 0 hour(s) and 6 minute(s): MediaWiki train - Utc-0+Utc-7 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T0800) [09:53:43] jouncebot: next [09:53:43] In 0 hour(s) and 6 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1000) [09:53:46] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy6002.drmrs.wmnet [09:53:58] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy6001.drmrs.wmnet [09:54:21] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org [09:54:21] !log filippo@cumin1003 END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for 10 hosts [09:57:47] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6002.drmrs.wmnet [09:57:48] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy6001.drmrs.wmnet [09:57:58] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet [09:58:12] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7002.magru.wmnet [09:59:18] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy5004.eqsin.wmnet [09:59:23] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy7001.magru.wmnet [09:59:27] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy5003.eqsin.wmnet [09:59:34] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy4004.ulsfo.wmnet [09:59:43] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy4003.ulsfo.wmnet [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1000) [10:01:18] (03CR) 10Blake: [C:03+2] mediawiki: Refactor lamp.deployment into component containers. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1313111 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [10:01:19] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5004.eqsin.wmnet [10:01:21] !log marostegui@cumin1003 START - Cookbook sre.mysql.clone of db1169.eqiad.wmnet onto db1283.eqiad.wmnet [10:01:25] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 [10:01:39] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy3002.esams.wmnet [10:01:40] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4004.ulsfo.wmnet [10:01:42] jmm@cumin2003 drain-node (PID 2734916) is awaiting input [10:01:54] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy2002.codfw.wmnet [10:02:00] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1169: Depool db1169.eqiad.wmnet to then clone it to db1283.eqiad.wmnet - marostegui@cumin1003 [10:02:26] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2031.codfw.wmnet [10:03:05] (03PS1) 10Marostegui: db1286: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1326790 [10:03:36] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy5003.eqsin.wmnet [10:03:43] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy4003.ulsfo.wmnet [10:04:04] (03CR) 10Marostegui: [C:03+2] db1286: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1326790 (owner: 10Marostegui) [10:04:04] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy1002.eqiad.wmnet [10:04:22] PROBLEM - Host kubestagemaster2005 is DOWN: PING CRITICAL - Packet loss = 100% [10:05:31] (03Merged) 10jenkins-bot: mediawiki: Refactor lamp.deployment into component containers. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1313111 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [10:05:38] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3002.esams.wmnet [10:05:46] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2002.codfw.wmnet [10:06:40] FIRING: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:07:53] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1002.eqiad.wmnet [10:08:12] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy1001.eqiad.wmnet [10:08:15] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy3001.esams.wmnet [10:08:18] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host tcp-proxy2001.codfw.wmnet [10:08:30] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host zuul2003.codfw.wmnet [10:08:34] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2031.codfw.wmnet [10:08:46] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet [10:10:30] RECOVERY - Host kubestagemaster2005 is UP: PING OK - Packet loss = 0%, RTA = 30.62 ms [10:12:08] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy1001.eqiad.wmnet [10:12:15] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy3001.esams.wmnet [10:12:17] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host tcp-proxy2001.codfw.wmnet [10:12:30] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2003.codfw.wmnet [10:12:42] (03PS1) 10Majavah: Add fake Grafana rendered token for metricsinfra [labs/private] - 10https://gerrit.wikimedia.org/r/1326791 [10:13:21] FIRING: [2x] ProbeDown: Service ganeti2031:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:13:25] (03CR) 10Majavah: [V:03+2 C:03+2] Add fake Grafana rendered token for metricsinfra [labs/private] - 10https://gerrit.wikimedia.org/r/1326791 (owner: 10Majavah) [10:14:34] !log blake@deploy1003 Started scap sync-world: no-build deployment for T417800 [10:14:38] T417800: Implement proper sidecar container support in mediawiki pods - https://phabricator.wikimedia.org/T417800 [10:15:51] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet [10:16:04] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12225607 (10Marostegui) @Jhancock.wm confirm I can start mariadb again then or you'd need the host for something else? [10:16:38] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host zuul2001.codfw.wmnet [10:16:41] !log arnaudb@cumin1003 START - Cookbook sre.hosts.reboot-single for host zuul2002.codfw.wmnet [10:17:13] (03PS1) 10Majavah: O:wmcs::metricsinfra::grafana: Add renderer plugin [puppet] - 10https://gerrit.wikimedia.org/r/1326793 [10:17:49] !log blake@deploy1003 Finished scap sync-world: no-build deployment for T417800 (duration: 04m 40s) [10:18:55] jmm@cumin2003 drain-node (PID 2737160) is awaiting input [10:20:37] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2001.codfw.wmnet [10:20:41] !log arnaudb@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zuul2002.codfw.wmnet [10:21:37] (03PS1) 10Marostegui: instances.yaml: Add db1286 [puppet] - 10https://gerrit.wikimedia.org/r/1326798 (https://phabricator.wikimedia.org/T407942) [10:22:14] (03CR) 10Marostegui: [C:03+2] instances.yaml: Add db1286 [puppet] - 10https://gerrit.wikimedia.org/r/1326798 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [10:22:48] !log installing Django security updates [10:22:51] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:22:51] (03PS1) 10Filippo Giunchedi: Stop mounting dumps on w[qc]ds [puppet] - 10https://gerrit.wikimedia.org/r/1326800 (https://phabricator.wikimedia.org/T432212) [10:23:13] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:24:33] !log marostegui@cumin1003 dbctl commit (dc=all): 'Add db1286 to dbctl T407942', diff saved to https://phabricator.wikimedia.org/P96163 and previous config saved to /var/cache/conftool/dbconfig/20260818-102431-marostegui.json [10:24:39] T407942: Productionize db12[65-90] - https://phabricator.wikimedia.org/T407942 [10:24:43] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1286: Pool back [10:26:17] jmm@cumin2003 drain-node (PID 2737160) is awaiting input [10:27:24] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2032.codfw.wmnet [10:27:39] (03PS1) 10Atsuko: airflow3: updating pyspark [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326804 (https://phabricator.wikimedia.org/T435172) [10:29:23] PROBLEM - Host ml-etcd2001 is DOWN: PING CRITICAL - Packet loss = 100% [10:30:35] (03CR) 10Btullis: [C:03+1] "Looks good to me. We can manually unmount and remove the entries from /etc/fstab on the affected hosts." [puppet] - 10https://gerrit.wikimedia.org/r/1326800 (https://phabricator.wikimedia.org/T432212) (owner: 10Filippo Giunchedi) [10:30:43] (03PS1) 10Tiziano Fogli: grafana/image-renderer: forward host env vars [puppet] - 10https://gerrit.wikimedia.org/r/1326805 (https://phabricator.wikimedia.org/T432970) [10:30:43] (03CR) 10Tiziano Fogli: [C:03+2] "Otherwise the chromium bin won't be found." [puppet] - 10https://gerrit.wikimedia.org/r/1326805 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [10:32:23] !log cgoubert@cumin2003 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad [10:32:50] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet [10:33:31] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2032.codfw.wmnet [10:33:39] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet [10:34:27] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [10:34:51] (03PS4) 10Blake: mediawiki: optionally use initContainers for sidecars. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325883 (https://phabricator.wikimedia.org/T417800) [10:35:55] RECOVERY - Host ml-etcd2001 is UP: PING OK - Packet loss = 0%, RTA = 31.90 ms [10:36:58] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) [10:36:59] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet [10:37:34] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet [10:38:21] FIRING: [2x] ProbeDown: Service ganeti2032:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:39:52] (03PS1) 10Marostegui: mariadb: Productionize db1282 [puppet] - 10https://gerrit.wikimedia.org/r/1326807 (https://phabricator.wikimedia.org/T407942) [10:40:42] (03CR) 10Marostegui: [C:03+2] mariadb: Productionize db1282 [puppet] - 10https://gerrit.wikimedia.org/r/1326807 (https://phabricator.wikimedia.org/T407942) (owner: 10Marostegui) [10:40:57] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet [10:41:07] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2047.codfw.wmnet [10:42:49] (03PS1) 10JMeybohm: Hand PageAssesment ownership to Content-Platform-Team [puppet] - 10https://gerrit.wikimedia.org/r/1326808 (https://phabricator.wikimedia.org/T434959) [10:43:20] (03CR) 10Blake: [C:03+2] rest-gateway: Use envoy-future in staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326221 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:43:21] PROBLEM - Host ml-staging-etcd2002 is DOWN: PING CRITICAL - Packet loss = 100% [10:43:26] 06SRE, 06Content-Platform-Team, 10MediaWiki-extensions-PageAssessments, 13Patch-For-Review: pageassessments-cleanup alerts tagged with defunct team - https://phabricator.wikimedia.org/T434959#12225717 (10JMeybohm) [10:44:25] !log marostegui@cumin1003 START - Cookbook sre.mysql.clone of db1168.eqiad.wmnet onto db1282.eqiad.wmnet [10:44:29] !log marostegui@cumin1003 START - Cookbook sre.mysql.depool depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 [10:44:45] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1168: Depool db1168.eqiad.wmnet to then clone it to db1282.eqiad.wmnet - marostegui@cumin1003 [10:44:59] (03CR) 10CI reject: [V:04-1] Hand PageAssesment ownership to Content-Platform-Team [puppet] - 10https://gerrit.wikimedia.org/r/1326808 (https://phabricator.wikimedia.org/T434959) (owner: 10JMeybohm) [10:45:32] (03PS2) 10JMeybohm: Hand PageAssesment ownership to Content-Platform-Team [puppet] - 10https://gerrit.wikimedia.org/r/1326808 (https://phabricator.wikimedia.org/T434959) [10:45:33] (03PS1) 10Clément Goubert: prometheus: Add redis_lock job [puppet] - 10https://gerrit.wikimedia.org/r/1326792 (https://phabricator.wikimedia.org/T427999) [10:45:45] (03Merged) 10jenkins-bot: rest-gateway: Use envoy-future in staging. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326221 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [10:45:49] RECOVERY - Host ml-staging-etcd2002 is UP: PING OK - Packet loss = 0%, RTA = 32.03 ms [10:45:50] (03CR) 10JMeybohm: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1326808 (https://phabricator.wikimedia.org/T434959) (owner: 10JMeybohm) [10:46:11] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [10:46:26] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db1901.eqiad.wmnet [10:46:28] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2047.codfw.wmnet [10:46:28] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [10:46:34] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet [10:46:44] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [10:47:38] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [10:47:46] 06SRE, 10SRE-Access-Requests, 10SRE-tools, 10Cumin, and 2 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12225729 (10JMeybohm) 05Open→03Resolved >>! In T433409#12216045, @JMeybohm wrote: > From the sudo rule added I would assume you need to c... [10:48:18] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet [10:48:26] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1044].eqiad.wmnet [10:48:55] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet [10:49:23] (03PS1) 10Cathal Mooney: Prometheus: add recording rules for number of gnmic series per device [puppet] - 10https://gerrit.wikimedia.org/r/1326809 (https://phabricator.wikimedia.org/T435184) [10:49:42] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [10:50:04] !log blake@deploy1003 helmfile [staging] START helmfile.d/services/rest-gateway: apply [10:50:30] !log blake@deploy1003 helmfile [staging] DONE helmfile.d/services/rest-gateway: apply [10:50:38] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet [10:51:07] 06SRE, 10SRE-Access-Requests, 06tools-infrastructure-team, 13Patch-For-Review: Requesting access to wmcs-roots for bliviero - https://phabricator.wikimedia.org/T435123#12225753 (10JMeybohm) 05Open→03Stalled [10:51:08] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" [10:51:12] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db1901.eqiad.wmnet - fceratto@cumin1003" [10:51:13] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [10:51:13] !log fceratto@cumin1003 START - Cookbook sre.dns.wipe-cache db1901.eqiad.wmnet on all recursors [10:51:17] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db1901.eqiad.wmnet on all recursors [10:51:33] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12225755 (10JMeybohm) 05In progress→03Stalled [10:52:07] 06SRE, 10SRE-Access-Requests, 06Data-Engineering: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12225767 (10JMeybohm) 05Open→03Stalled [10:53:02] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12225769 (10MoritzMuehlenhoff) [10:53:34] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet [10:53:55] (03CR) 10Fabfur: [C:03+1] Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [10:54:44] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage [10:54:46] (03PS1) 10Cathal Mooney: team-netops: Add GnmiInterfaceCountersDrop alert [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) [10:55:31] jmm@cumin2003 drain-node (PID 2743864) is awaiting input [10:56:46] (03CR) 10CI reject: [V:04-1] team-netops: Add GnmiInterfaceCountersDrop alert [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [10:57:23] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet [10:58:22] (03CR) 10Slyngshede: [C:03+2] Varnish: Reject non-thumb requests to thumbs [puppet] - 10https://gerrit.wikimedia.org/r/1324264 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [10:59:01] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet [10:59:07] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2049.codfw.wmnet [10:59:18] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1181.eqiad.wmnet with reason: host reimage [10:59:39] (03PS2) 10Cathal Mooney: team-netops: Add GnmiInterfaceCountersDrop alert [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) [11:01:05] PROBLEM - Host aux-k8s-etcd2004 is DOWN: PING CRITICAL - Packet loss = 100% [11:01:36] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet [11:01:36] (03CR) 10CI reject: [V:04-1] team-netops: Add GnmiInterfaceCountersDrop alert [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [11:01:39] FIRING: [6x] CirrusSearchTitleSuggestIndexTooOld: Some search indices that power autocomplete have not been updated recently - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - TODO - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchTitleSuggestIndexTooOld [11:02:54] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet [11:04:39] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2049.codfw.wmnet [11:04:40] (03PS3) 10Cathal Mooney: team-netops: Add GnmiInterfaceCountersDrop alert [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) [11:05:26] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet [11:05:31] RECOVERY - Host aux-k8s-etcd2004 is UP: PING OK - Packet loss = 0%, RTA = 31.89 ms [11:06:24] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage [11:06:28] !log btullis@cumin1003 END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on an-worker1194.eqiad.wmnet with reason: host reimage [11:06:39] FIRING: [6x] CirrusSearchTitleSuggestIndexTooOld: Some search indices that power autocomplete have not been updated recently - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - TODO - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchTitleSuggestIndexTooOld [11:07:16] !log installing PHP 8.4 security updates [11:07:18] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:08:22] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet [11:08:34] (03PS2) 10Hnowlan: sre/cdn: create recording rule in advance of moving to ratio for CDN [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) [11:09:17] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet [11:09:25] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1045-1050,1056-1057,1064-1066,1073-1075].eqiad.wmnet [11:09:27] (03PS1) 10Hnowlan: sre/cdn: use traffic ratios in ATSBackendErrorsHigh [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) [11:09:52] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1286: Pool back [11:09:57] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet [11:10:26] 06SRE, 10SRE-Access-Requests: Requesting access to 'restricted' for nicholusmuwonge - https://phabricator.wikimedia.org/T435168#12225801 (10JMeybohm) [11:10:35] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet [11:12:35] 06SRE, 10SRE-Access-Requests: Requesting access to 'restricted' for nicholusmuwonge - https://phabricator.wikimedia.org/T435168#12225818 (10JMeybohm) 05Open→03Stalled The user is member of the `nda` group, so NDA has been signed. @dancy || @thcipriani please sign off as `restricted` group approves. [11:14:25] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet [11:15:19] jmm@cumin2003 drain-node (PID 2748837) is awaiting input [11:15:52] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2050.codfw.wmnet [11:18:12] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet [11:21:13] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2050.codfw.wmnet [11:21:19] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet [11:23:55] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1181.eqiad.wmnet with OS bookworm [11:24:42] (03PS1) 10JavierMonton: stream: mw-page-html-feature-counts-change (staging) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326814 (https://phabricator.wikimedia.org/T429456) [11:24:57] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1194.eqiad.wmnet with OS bookworm [11:25:14] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1168: Pool db1168.eqiad.wmnet in after cloning [11:25:46] (03CR) 10Filippo Giunchedi: [C:03+2] Stop mounting dumps on w[qc]ds [puppet] - 10https://gerrit.wikimedia.org/r/1326800 (https://phabricator.wikimedia.org/T432212) (owner: 10Filippo Giunchedi) [11:26:50] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet [11:26:58] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1076-1081,1084-1087,1093-1095,1113].eqiad.wmnet [11:27:28] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1114-1127].eqiad.wmnet [11:27:40] (03CR) 10Ayounsi: [C:03+1] "looks good!" [puppet] - 10https://gerrit.wikimedia.org/r/1326809 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [11:30:22] (03PS2) 10Majavah: O:wmcs::metricsinfra::grafana: Add renderer plugin [puppet] - 10https://gerrit.wikimedia.org/r/1326793 [11:30:54] !log jmm@cumin2003 START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad [11:30:54] (03CR) 10Ayounsi: [C:03+1] "Thanks, that will help!" [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [11:31:32] (03PS1) 10Majavah: P:simplelamp2: Set memcached_user to modern value [puppet] - 10https://gerrit.wikimedia.org/r/1326816 (https://phabricator.wikimedia.org/T273950) [11:33:17] (03PS4) 10Cathal Mooney: team-netops: Add GnmiInterfaceCountersDrop alert [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) [11:35:02] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [11:35:49] !log marostegui@cumin1003 START - Cookbook sre.mysql.pool pool db1169: Pool db1169.eqiad.wmnet in after cloning [11:35:58] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.ganeti.makevm (exit_code=93) for new host db1901.eqiad.wmnet [11:36:05] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1114-1127].eqiad.wmnet [11:36:44] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2025.codfw.wmnet [11:37:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:39:52] jmm@cumin2003 drain-node (PID 2754684) is awaiting input [11:39:56] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db2901.codfw.wmnet [11:39:57] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.ganeti.makevm (exit_code=99) for new host db2901.codfw.wmnet [11:41:39] FIRING: [3x] CirrusSearchTitleSuggestIndexTooOld: Some search indices that power autocomplete have not been updated recently - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - TODO - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchTitleSuggestIndexTooOld [11:43:00] (03CR) 10Filippo Giunchedi: [C:03+1] P:simplelamp2: Set memcached_user to modern value [puppet] - 10https://gerrit.wikimedia.org/r/1326816 (https://phabricator.wikimedia.org/T273950) (owner: 10Majavah) [11:43:15] (03CR) 10Filippo Giunchedi: [C:03+1] O:wmcs::metricsinfra::grafana: Add renderer plugin [puppet] - 10https://gerrit.wikimedia.org/r/1326793 (owner: 10Majavah) [11:45:29] (03CR) 10Majavah: [C:03+2] P:simplelamp2: Set memcached_user to modern value [puppet] - 10https://gerrit.wikimedia.org/r/1326816 (https://phabricator.wikimedia.org/T273950) (owner: 10Majavah) [11:45:32] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1114-1127].eqiad.wmnet [11:45:40] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1114-1127].eqiad.wmnet [11:45:48] (03CR) 10Majavah: [C:03+2] O:wmcs::metricsinfra::grafana: Add renderer plugin [puppet] - 10https://gerrit.wikimedia.org/r/1326793 (owner: 10Majavah) [11:45:55] FIRING: [2x] CoreBGPDown: Core BGP session down between cr1-magru and cr2-eqiad (195.200.68.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=magru&var-device=cr1-magru:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [11:46:10] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet [11:50:51] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2025.codfw.wmnet [11:51:39] RESOLVED: [3x] CirrusSearchTitleSuggestIndexTooOld: Some search indices that power autocomplete have not been updated recently - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - TODO - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchTitleSuggestIndexTooOld [11:52:30] 06SRE, 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: Modernise memcached systemd unit / sync, and make it presentable - https://phabricator.wikimedia.org/T273950#12225970 (10taavi) As far as I can tell it's just `profile::swift::proxy` and `profile::thanos::swift::f... [11:54:13] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet [11:54:50] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2046.codfw.wmnet [11:55:11] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12225972 (10BTullis) @VRiley-WMF - Yes, please go ahead and try that other port. I tried all sorts of things too, including sending raw IPMI comm... [11:55:39] (03CR) 10Cathal Mooney: [C:03+2] Prometheus: add recording rules for number of gnmic series per device [puppet] - 10https://gerrit.wikimedia.org/r/1326809 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [11:57:03] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2025.codfw.wmnet [11:57:09] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2025.codfw.wmnet [11:58:21] FIRING: [2x] ProbeDown: Service ganeti2025:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:58:33] (03PS1) 10Blake: envoy: New upstream version 1.39.0. [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1326818 (https://phabricator.wikimedia.org/T421418) [12:00:04] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1200) [12:01:59] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet [12:02:00] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2046.codfw.wmnet [12:02:07] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1128-1134,1142-1148].eqiad.wmnet [12:02:14] 07Puppet, 06Collaboration-Services, 10Gerrit, 06Infrastructure-Foundations, 13Patch-For-Review: Change puppet-merge git origin to use gerrit.discovery.wmnet instead of gerrit.wikimedia.org - https://phabricator.wikimedia.org/T420184#12226002 (10ABran-WMF) [12:02:39] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet [12:03:34] jouncebot: nowandnext [12:03:34] For the next 0 hour(s) and 56 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1200) [12:03:34] In 0 hour(s) and 56 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1300) [12:05:15] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2026.codfw.wmnet [12:05:48] (03CR) 10Slyngshede: P:tofurkey enable Tofurkey for MAGRU (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1324689 (https://phabricator.wikimedia.org/T355446) (owner: 10Slyngshede) [12:06:25] RESOLVED: SystemdUnitFailed: dump_proxy_ranges.service on puppetserver1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:06:51] !log jmm@cumin2003 END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad [12:07:31] (03PS5) 10Blake: mediawiki: optionally use initContainers for sidecars. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325883 (https://phabricator.wikimedia.org/T417800) [12:07:57] (03PS6) 10Blake: mediawiki: stop using lamp.deployment in job.yaml.tpl. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325884 (https://phabricator.wikimedia.org/T417800) [12:08:19] jmm@cumin2003 drain-node (PID 2760834) is awaiting input [12:10:29] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1168: Pool db1168.eqiad.wmnet in after cloning [12:10:37] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1168.eqiad.wmnet onto db1282.eqiad.wmnet [12:11:11] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet [12:14:39] (03PS1) 10PipelineBot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326822 [12:14:45] FIRING: CirrusStreamingUpdaterFlinkJobUnstable: cirrus_streaming_updater_consumer_cloudelastic_eqiad in eqiad (k8s) is unstable - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/K9x0c4aVk/flink-app?var-datasource=eqiad+prometheus%2Fk8s&var-namespace=cirrus-streaming-updater&var-helm_release=consumer-cloudelastic - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterFlinkJobUnsta [12:18:24] (03CR) 10Dbrant: [C:03+2] wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326822 (owner: 10PipelineBot) [12:19:19] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet [12:19:27] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1149-1153,1158,1240-1247].eqiad.wmnet [12:19:45] RESOLVED: CirrusStreamingUpdaterFlinkJobUnstable: cirrus_streaming_updater_consumer_cloudelastic_eqiad in eqiad (k8s) is unstable - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/K9x0c4aVk/flink-app?var-datasource=eqiad+prometheus%2Fk8s&var-namespace=cirrus-streaming-updater&var-helm_release=consumer-cloudelastic - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterFlinkJobUns [12:19:56] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1248-1261].eqiad.wmnet [12:20:44] (03Merged) 10jenkins-bot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326822 (owner: 10PipelineBot) [12:20:46] PROBLEM - Host an-worker1147 is DOWN: PING CRITICAL - Packet loss = 100% [12:21:08] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Pool db1169.eqiad.wmnet in after cloning [12:21:16] !log marostegui@cumin1003 END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1169.eqiad.wmnet onto db1283.eqiad.wmnet [12:21:39] !log jmm@cumin2003 START - Cookbook sre.ganeti.addnode for new host ganeti2046.codfw.wmnet to cluster codfw and group A [12:21:51] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12226059 (10VRiley-WMF) 05Open→03In progress [12:22:00] !log dbrant@deploy1003 helmfile [staging] START helmfile.d/services/wikifeeds: apply [12:22:23] !log dbrant@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifeeds: apply [12:22:47] !log dbrant@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifeeds: apply [12:22:57] (03CR) 10Kamila Součková: [V:03+2 C:03+2] "verified by building locally" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1325907 (https://phabricator.wikimedia.org/T432988) (owner: 10Kamila Součková) [12:23:14] !log dbrant@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply [12:23:24] !log dbrant@deploy1003 helmfile [codfw] START helmfile.d/services/wikifeeds: apply [12:23:52] !log dbrant@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply [12:24:27] jouncebot: now [12:24:27] For the next 0 hour(s) and 35 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1200) [12:24:45] I’ll roll out a harmless little query service GUI deploy, hope that’s okay [12:24:51] (03CR) 10Lucas Werkmeister (WMDE): [C:03+2] "let’s try it :)" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326307 (owner: 10Lucas Werkmeister (WMDE)) [12:25:54] (03CR) 10Cathal Mooney: team-netops: Add GnmiInterfaceCountersDrop alert (032 comments) [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [12:25:58] RECOVERY - Host an-worker1147 is UP: PING WARNING - Packet loss = 90%, RTA = 0.31 ms [12:26:07] !log readded ganeti2046 to the codfw cluster following firmware update and reimage T434681 [12:26:12] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:26:14] T434681: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681 [12:27:24] (03Merged) 10jenkins-bot: wikidata-query-gui: update query-next staging custom-config.json [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326307 (owner: 10Lucas Werkmeister (WMDE)) [12:27:24] jmm@cumin2003 drain-node (PID 2760834) is awaiting input [12:27:32] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2046.codfw.wmnet to cluster codfw and group A [12:27:48] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1248-1261].eqiad.wmnet [12:28:44] PROBLEM - Host an-worker1147 is DOWN: PING CRITICAL - Packet loss = 100% [12:28:46] !log lucaswerkmeister-wmde@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [12:28:50] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2026.codfw.wmnet [12:29:24] !log lucaswerkmeister-wmde@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [12:32:12] (03CR) 10Bking: [C:03+1] cirrus: remove orphaned frozen writes command [puppet] - 10https://gerrit.wikimedia.org/r/1326393 (https://phabricator.wikimedia.org/T433306) (owner: 10Ryan Kemper) [12:34:06] !log lucaswerkmeister-wmde@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [12:34:27] !log lucaswerkmeister-wmde@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [12:34:31] !log lucaswerkmeister-wmde@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [12:34:50] !log lucaswerkmeister-wmde@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [12:35:13] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2026.codfw.wmnet [12:36:08] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1248-1261].eqiad.wmnet [12:36:16] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1248-1261].eqiad.wmnet [12:36:43] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet [12:37:27] !log depool again clouddb10[13,14,16,18,20] that were repooled by the cookbook sre.mysql.multiinstance_reboot [12:37:29] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:37:34] !log fnegri@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet [12:37:52] !log fnegri@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet [12:38:09] !log fnegri@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet [12:38:16] !log fnegri@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet [12:38:21] !log fnegri@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet [12:39:19] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org [12:39:27] (03PS1) 10Cathal Mooney: Prometheus recording rules: correct label for 'job' in gnmic count [puppet] - 10https://gerrit.wikimedia.org/r/1326824 (https://phabricator.wikimedia.org/T435184) [12:39:45] !log fnegri@cumin1003 conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet [12:40:20] !log fnegri@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1032 [12:40:26] (03CR) 10Ottomata: [C:03+1] "nit about comment but +1 otherwise!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326814 (https://phabricator.wikimedia.org/T429456) (owner: 10JavierMonton) [12:40:46] !log fnegri@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet [12:40:54] !log fceratto@cumin1003 START - Cookbook sre.ganeti.makevm for new host db2902.codfw.wmnet [12:40:56] (03CR) 10Ayounsi: [C:03+1] Prometheus recording rules: correct label for 'job' in gnmic count [puppet] - 10https://gerrit.wikimedia.org/r/1326824 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [12:40:57] !log fceratto@cumin1003 START - Cookbook sre.dns.netbox [12:41:21] !log fnegri@cumin1003 conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet [12:41:46] !log also depooled clouddb1017 (forgot it in the previous list) [12:41:49] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:41:56] (03CR) 10Cathal Mooney: [C:03+2] Prometheus recording rules: correct label for 'job' in gnmic count [puppet] - 10https://gerrit.wikimedia.org/r/1326824 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [12:42:09] !log repooled clouddb1032 that was currently {"weight": 0, "pooled": "inactive"} for both s4 and s6 [12:42:12] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:43:23] (03PS1) 10MVernon: swift: move ms-be106{4,5,8} to new-style storage, remove from rings [puppet] - 10https://gerrit.wikimedia.org/r/1326826 (https://phabricator.wikimedia.org/T308644) [12:44:37] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1273,1275-1287].eqiad.wmnet [12:45:09] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org [12:46:24] fceratto@cumin1003 makevm (PID 824663) is awaiting input [12:46:52] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org [12:49:18] (03CR) 10Marostegui: [C:03+1] swift: move ms-be106{4,5,8} to new-style storage, remove from rings [puppet] - 10https://gerrit.wikimedia.org/r/1326826 (https://phabricator.wikimedia.org/T308644) (owner: 10MVernon) [12:50:25] (03CR) 10MVernon: [C:03+2] swift: move ms-be106{4,5,8} to new-style storage, remove from rings [puppet] - 10https://gerrit.wikimedia.org/r/1326826 (https://phabricator.wikimedia.org/T308644) (owner: 10MVernon) [12:51:26] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" [12:51:30] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM db2902.codfw.wmnet - fceratto@cumin1003" [12:51:31] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [12:51:31] !log fceratto@cumin1003 START - Cookbook sre.dns.wipe-cache db2902.codfw.wmnet on all recursors [12:51:35] !log fceratto@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) db2902.codfw.wmnet on all recursors [12:52:45] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org [12:53:14] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet [12:53:22] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1273,1275-1287].eqiad.wmnet [12:53:49] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet [12:59:17] RECOVERY - Host an-worker1147 is UP: PING OK - Packet loss = 0%, RTA = 0.31 ms [12:59:53] (03CR) 10Hashar: "The addition of zuul1004 caused Zookeeper to be unresponsive which I filed as T435186. The reason is that firewall rules need to be added " [puppet] - 10https://gerrit.wikimedia.org/r/1319528 (https://phabricator.wikimedia.org/T427353) (owner: 10Dzahn) [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1300). Please do the needful. [13:00:05] nemo-yiannis and Tran: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:11] 👋 [13:00:12] I can’t deploy today, sorry [13:00:21] i can deploy mine [13:00:27] (03PS1) 10Muehlenhoff: Move the insetup role report to cumin2003 [puppet] - 10https://gerrit.wikimedia.org/r/1326828 [13:00:29] o/ I can deploy my own as well [13:00:35] should i go ahead? [13:00:44] sure [13:00:46] thanks [13:01:01] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jgiannelos@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326232 (https://phabricator.wikimedia.org/T435115) (owner: 10Jgiannelos) [13:01:26] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [13:02:05] (03Merged) 10jenkins-bot: prv: Enable parsoid rendering for 5 wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326232 (https://phabricator.wikimedia.org/T435115) (owner: 10Jgiannelos) [13:02:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:02:28] !log jgiannelos@deploy1003 Started scap sync-world: Backport for [[gerrit:1326232|prv: Enable parsoid rendering for 5 wikis (T435115)]] [13:02:28] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet [13:02:34] T435115: Rollout parsoid read views on wikisource - week of 14 Aug - https://phabricator.wikimedia.org/T435115 [13:03:12] (03PS2) 10JavierMonton: stream: mw-page-html-feature-counts-change (staging) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326814 (https://phabricator.wikimedia.org/T429456) [13:04:15] (03CR) 10Elukey: docker_registry: route /v2/releng to its new s3 backend (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1326323 (https://phabricator.wikimedia.org/T432829) (owner: 10Elukey) [13:04:44] !log jgiannelos@deploy1003 jgiannelos: Backport for [[gerrit:1326232|prv: Enable parsoid rendering for 5 wikis (T435115)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:04:47] PROBLEM - Host an-worker1147 is DOWN: PING CRITICAL - Packet loss = 100% [13:05:33] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet [13:06:24] (03CR) 10JavierMonton: stream: mw-page-html-feature-counts-change (staging) (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326814 (https://phabricator.wikimedia.org/T429456) (owner: 10JavierMonton) [13:06:29] !log jgiannelos@deploy1003 jgiannelos: Continuing with deployment [13:06:30] (03CR) 10JavierMonton: [C:03+2] stream: mw-page-html-feature-counts-change (staging) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326814 (https://phabricator.wikimedia.org/T429456) (owner: 10JavierMonton) [13:08:07] (03PS1) 10Muehlenhoff: Remove cluster management role from cumin2002 [puppet] - 10https://gerrit.wikimedia.org/r/1326829 (https://phabricator.wikimedia.org/T427897) [13:08:17] RECOVERY - Host an-worker1147 is UP: PING OK - Packet loss = 0%, RTA = 0.27 ms [13:08:39] (03Merged) 10jenkins-bot: stream: mw-page-html-feature-counts-change (staging) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326814 (https://phabricator.wikimedia.org/T429456) (owner: 10JavierMonton) [13:09:20] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet [13:09:59] PROBLEM - Host an-worker1147 is DOWN: PING CRITICAL - Packet loss = 100% [13:10:15] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply [13:10:24] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet [13:10:28] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply [13:10:32] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1288-1289,1291-1301].eqiad.wmnet [13:10:45] !log jgiannelos@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326232|prv: Enable parsoid rendering for 5 wikis (T435115)]] (duration: 08m 17s) [13:10:50] T435115: Rollout parsoid read views on wikisource - week of 14 Aug - https://phabricator.wikimedia.org/T435115 [13:10:56] ok i am done Tran [13:11:02] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1302-1314].eqiad.wmnet [13:11:03] Thanks, going to start mine [13:11:35] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326775 (https://phabricator.wikimedia.org/T435048) (owner: 10STran) [13:11:49] (03PS1) 10FNegri: clouddb: Revert temp changes to clouddb1025 [puppet] - 10https://gerrit.wikimedia.org/r/1326830 (https://phabricator.wikimedia.org/T409557) [13:12:37] (03PS5) 10Hashar: zuul: add firewall rules for Zookeeper [puppet] - 10https://gerrit.wikimedia.org/r/1326812 (https://phabricator.wikimedia.org/T435186) [13:12:37] (03CR) 10Hashar: "PCC https://puppet-compiler.wmflabs.org/output/1326812/7566/" [puppet] - 10https://gerrit.wikimedia.org/r/1326812 (https://phabricator.wikimedia.org/T435186) (owner: 10Hashar) [13:13:13] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:13:20] !log installing util-linux security updates [13:13:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:13:46] (03Merged) 10jenkins-bot: SI: Instrument case update on first edit [extensions/CheckUser] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326775 (https://phabricator.wikimedia.org/T435048) (owner: 10STran) [13:14:08] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1326775|SI: Instrument case update on first edit (T435048)]] [13:14:13] T435048: Trigger case_updated event on account first editing - https://phabricator.wikimedia.org/T435048 [13:14:15] (03CR) 10FNegri: "s6 must be removed from the host, s4 can stay and it will serve requests for x4, like in clouddb1024." [puppet] - 10https://gerrit.wikimedia.org/r/1326830 (https://phabricator.wikimedia.org/T409557) (owner: 10FNegri) [13:15:11] RECOVERY - Host an-worker1147 is UP: PING WARNING - Packet loss = 90%, RTA = 0.23 ms [13:16:10] !log stran@deploy1003 stran: Backport for [[gerrit:1326775|SI: Instrument case update on first edit (T435048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:17:08] !log stran@deploy1003 stran: Continuing with deployment [13:18:13] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:18:31] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1302-1314].eqiad.wmnet [13:21:21] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326775|SI: Instrument case update on first edit (T435048)]] (duration: 07m 12s) [13:21:27] T435048: Trigger case_updated event on account first editing - https://phabricator.wikimedia.org/T435048 [13:21:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d8-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [13:22:02] I'm done with my deploy but if no one else needs the window, I have a few commands I need to run on prod (table creation) [13:22:56] (03PS1) 10Slyngshede: P:idp enable MFA for production [puppet] - 10https://gerrit.wikimedia.org/r/1326833 (https://phabricator.wikimedia.org/T277841) [13:24:55] (03PS1) 10AikoChou: ml-services: update ores-legacy image tag [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326834 (https://phabricator.wikimedia.org/T429675) [13:26:59] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1302-1314].eqiad.wmnet [13:27:07] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1302-1314].eqiad.wmnet [13:27:34] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1315-1327].eqiad.wmnet [13:32:11] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1326833 (https://phabricator.wikimedia.org/T277841) (owner: 10Slyngshede) [13:32:24] running queries, please ping if you need to do something [13:34:30] (03PS1) 10Elukey: docker_registry: improve registry-homepage-builder.py [puppet] - 10https://gerrit.wikimedia.org/r/1326835 (https://phabricator.wikimedia.org/T427175) [13:34:48] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet [13:35:10] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1315-1327].eqiad.wmnet [13:36:04] (03CR) 10CDanis: [C:03+2] "private patch merged and deployed, helmfile diffs look reasonable" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1320964 (https://phabricator.wikimedia.org/T433982) (owner: 10CDanis) [13:36:46] 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for eamedina - https://phabricator.wikimedia.org/T435200 (10eamedina) 03NEW [13:37:08] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for eamedina - https://phabricator.wikimedia.org/T435200#12226337 (10eamedina) [13:37:38] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for eamedina - https://phabricator.wikimedia.org/T435200#12226338 (10eamedina) [13:37:49] PROBLEM - Host an-worker1147 is DOWN: PING CRITICAL - Packet loss = 100% [13:39:31] (03CR) 10Kevin Bazira: [C:03+1] ml-services: update ores-legacy image tag [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326834 (https://phabricator.wikimedia.org/T429675) (owner: 10AikoChou) [13:40:14] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet [13:41:55] RECOVERY - Host an-worker1147 is UP: PING OK - Packet loss = 0%, RTA = 0.25 ms [13:43:00] (03PS1) 10STran: Enable Suggested Investigations on hewiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326838 (https://phabricator.wikimedia.org/T435146) [13:43:13] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:43:21] (03CR) 10Cathal Mooney: [C:03+2] team-netops: Add GnmiInterfaceCountersDrop alert [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [13:43:38] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1315-1327].eqiad.wmnet [13:43:45] !log cdanis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply [13:43:46] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1315-1327].eqiad.wmnet [13:43:49] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet [13:44:13] PROBLEM - Host an-worker1147 is DOWN: PING CRITICAL - Packet loss = 100% [13:44:17] !log cdanis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply [13:44:21] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for eamedina - https://phabricator.wikimedia.org/T435200#12226349 (10eamedina) Tagging my manager for approval: @Nikerabbit [13:44:21] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet [13:44:42] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326838 (https://phabricator.wikimedia.org/T435146) (owner: 10STran) [13:45:04] (03CR) 10Dreamy Jazz: [C:03+1] Enable Suggested Investigations on hewiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326838 (https://phabricator.wikimedia.org/T435146) (owner: 10STran) [13:45:14] queries done, running an extra config deploy [13:45:34] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326838 (https://phabricator.wikimedia.org/T435146) (owner: 10STran) [13:45:57] (03Merged) 10jenkins-bot: team-netops: Add GnmiInterfaceCountersDrop alert [alerts] - 10https://gerrit.wikimedia.org/r/1326810 (https://phabricator.wikimedia.org/T435184) (owner: 10Cathal Mooney) [13:46:48] (03Merged) 10jenkins-bot: Enable Suggested Investigations on hewiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326838 (https://phabricator.wikimedia.org/T435146) (owner: 10STran) [13:47:08] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1326838|Enable Suggested Investigations on hewiki (T435146)]] [13:47:12] T435146: Enable suggested investigations on hewiki - https://phabricator.wikimedia.org/T435146 [13:47:52] !log phuedx@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-analytics: apply [13:47:56] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet [13:48:03] !log phuedx@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-analytics: apply [13:48:38] !log phuedx@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-analytics: apply [13:49:15] !log stran@deploy1003 stran: Backport for [[gerrit:1326838|Enable Suggested Investigations on hewiki (T435146)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:49:23] !log phuedx@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: apply [13:50:11] !log phuedx@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-analytics: apply [13:50:31] !log cgoubert@cumin2003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet [13:50:33] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet [13:50:33] !log cgoubert@cumin2003 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad [13:50:53] !log phuedx@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: apply [13:51:34] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet [13:51:43] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2026.codfw.wmnet [13:51:48] !log fceratto@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" [13:51:52] !log fceratto@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM db2902.codfw.wmnet - fceratto@cumin1003" [13:52:59] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2041.codfw.wmnet [13:53:13] FIRING: [2x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [13:53:17] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet [13:53:51] !log phuedx@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply [13:53:58] 10ops-eqiad, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T435201 (10phaultfinder) 03NEW [13:54:01] !log phuedx@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-logging-external: apply [13:54:10] !log phuedx@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-logging-external: apply [13:54:20] !log phuedx@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-logging-external: apply [13:54:39] !log stran@deploy1003 stran: Continuing with deployment [13:54:48] !log phuedx@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-logging-external: apply [13:54:53] fceratto@cumin1003 makevm (PID 824663) is awaiting input [13:55:41] !log phuedx@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-logging-external: apply [13:55:47] PROBLEM - ganeti-noded running on ganeti2041 is CRITICAL: PROCS CRITICAL: 3 processes with UID = 0 (root), command name ganeti-noded https://wikitech.wikimedia.org/wiki/Ganeti [13:55:52] !log phuedx@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-logging-external: apply [13:56:14] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for eamedina - https://phabricator.wikimedia.org/T435200#12226389 (10Aklapper) Hi, https://www.mediawiki.org/wiki/Product_Analytics/Superset_Access#Requesting_access links to https://phabricator.wikimedia.org/maniphest/task/edit/for... [13:56:23] (03PS1) 10Elukey: CHANGELOG: add changelogs for release v13.2.0 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1326840 [13:56:47] RECOVERY - ganeti-noded running on ganeti2041 is OK: PROCS OK: 1 process with UID = 0 (root), command name ganeti-noded https://wikitech.wikimedia.org/wiki/Ganeti [13:57:04] !log phuedx@deploy1003 helmfile [staging] START helmfile.d/services/eventgate-main: apply [13:57:06] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet [13:57:13] !log phuedx@deploy1003 helmfile [staging] DONE helmfile.d/services/eventgate-main: apply [13:57:48] !log phuedx@deploy1003 helmfile [codfw] START helmfile.d/services/eventgate-main: apply [13:57:56] 06SRE, 13Patch-For-Review: Add alerting for gnmic total series counters - https://phabricator.wikimedia.org/T435184#12226407 (10Aklapper) [13:58:35] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2041.codfw.wmnet [13:58:36] !log phuedx@deploy1003 helmfile [codfw] DONE helmfile.d/services/eventgate-main: apply [13:58:58] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326838|Enable Suggested Investigations on hewiki (T435146)]] (duration: 11m 50s) [13:59:03] T435146: Enable suggested investigations on hewiki - https://phabricator.wikimedia.org/T435146 [13:59:18] !log phuedx@deploy1003 helmfile [eqiad] START helmfile.d/services/eventgate-main: apply [13:59:25] RECOVERY - Host an-worker1147 is UP: PING OK - Packet loss = 0%, RTA = 0.29 ms [13:59:56] !log phuedx@deploy1003 helmfile [eqiad] DONE helmfile.d/services/eventgate-main: apply [14:00:04] Deploy window Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1400) [14:00:58] done [14:01:03] (03PS1) 10Muehlenhoff: Remove cumin2002 from RAPI access rules [puppet] - 10https://gerrit.wikimedia.org/r/1326842 (https://phabricator.wikimedia.org/T427897) [14:02:49] (03CR) 10Elukey: [C:03+2] CHANGELOG: add changelogs for release v13.2.0 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1326840 (owner: 10Elukey) [14:03:18] (03PS1) 10Muehlenhoff: Remove cumin2002 as mariadb root client [puppet] - 10https://gerrit.wikimedia.org/r/1326843 (https://phabricator.wikimedia.org/T427897) [14:04:01] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2041.codfw.wmnet [14:04:29] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2041.codfw.wmnet [14:04:55] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2042.codfw.wmnet [14:05:39] (03PS1) 10Santiago Faci: Deploy GrowthBook 5.0.0 to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326844 (https://phabricator.wikimedia.org/T434098) [14:06:45] 10ops-codfw, 06SRE, 06DC-Ops: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681#12226473 (10MoritzMuehlenhoff) 05Open→03Resolved a:03MoritzMuehlenhoff The server has been reimaged after being updated to fixed firmware and has been re-added to the codfw Ganeti clu... [14:07:30] (03PS1) 10Muehlenhoff: Remove DB grant for cumin2002 [puppet] - 10https://gerrit.wikimedia.org/r/1326845 (https://phabricator.wikimedia.org/T427897) [14:07:58] (03PS1) 10Elukey: Upstream release v13.2.0 [software/spicerack] (debian) - 10https://gerrit.wikimedia.org/r/1326846 [14:08:21] 06SRE, 10fundraising-tech-ops, 06Infrastructure-Foundations, 10netops: Upgrade JunOS on pfw1-eqiad and pfw1-codfw - https://phabricator.wikimedia.org/T434865#12226495 (10cmooney) >>! In T434865#12221432, @Jgreen wrote: > @cmooney this happens to be a FR maintenance week, is it feasible to plan it sometime... [14:08:54] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12226497 (10VRiley-WMF) I made some configureations on the server and set it to LOM 3 as the iDRAC. I haven't been able to get into the iDRAC, bu... [14:09:07] (03CR) 10Elukey: [V:03+2 C:03+2] Upstream release v13.2.0 [software/spicerack] (debian) - 10https://gerrit.wikimedia.org/r/1326846 (owner: 10Elukey) [14:09:15] (03PS1) 10Muehlenhoff: Remove cumin2002 from the homer peer list of cumin1003 [puppet] - 10https://gerrit.wikimedia.org/r/1326847 (https://phabricator.wikimedia.org/T427897) [14:09:27] 10ops-eqiad, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T435201#12226502 (10VRiley-WMF) a:03VRiley-WMF [14:09:36] 10ops-eqiad, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T435201#12226503 (10VRiley-WMF) 05Open→03Resolved Duplicate ticket [14:11:17] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for eamedina - https://phabricator.wikimedia.org/T435200#12226523 (10eamedina) [14:11:23] RESOLVED: CertAlmostExpired: gNMI TLS certificate for lsw1-d8-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [14:11:46] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for eamedina - https://phabricator.wikimedia.org/T435200#12226539 (10eamedina) Thanks! Was following an older example, task description updated now [14:12:27] (03PS1) 10Sadiya.mohammed13: Disable Wikidata Bridge on Catalan Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326848 (https://phabricator.wikimedia.org/T433713) [14:13:53] jmm@cumin2003 drain-node (PID 2790757) is awaiting input [14:13:54] (03CR) 10Blake: "Done!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325883 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [14:14:56] (03CR) 10Aghirelli: [C:03+1] Add configurable RestModuleOverrides [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324819 (https://phabricator.wikimedia.org/T434267) (owner: 10Milazg) [14:15:24] (03PS3) 10Elukey: docker_registry: route /v2/releng to its new s3 backend [puppet] - 10https://gerrit.wikimedia.org/r/1326323 (https://phabricator.wikimedia.org/T432829) [14:15:24] (03PS2) 10Elukey: docker_registry: improve registry-homepage-builder.py [puppet] - 10https://gerrit.wikimedia.org/r/1326835 (https://phabricator.wikimedia.org/T427175) [14:15:25] (03PS1) 10Elukey: profile::docker_registry: scan the Releng's endpoint too [puppet] - 10https://gerrit.wikimedia.org/r/1326849 (https://phabricator.wikimedia.org/T427175) [14:16:42] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2042.codfw.wmnet [14:16:51] !log uploaded spicerack_13.2.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia [14:16:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:18:55] PROBLEM - Host aux-k8s-etcd2003 is DOWN: PING CRITICAL - Packet loss = 100% [14:20:45] RECOVERY - Host aux-k8s-etcd2003 is UP: PING OK - Packet loss = 0%, RTA = 30.43 ms [14:21:26] !log fceratto@cumin1003 START - Cookbook sre.hosts.reimage for host db2902.codfw.wmnet with OS trixie [14:22:02] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2042.codfw.wmnet [14:22:08] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2042.codfw.wmnet [14:22:55] 06SRE, 06Infrastructure-Foundations, 10Puppet-Infrastructure: Create PCC Puppet 8 nodes - https://phabricator.wikimedia.org/T374495#12226664 (10LSobanski) a:05jhathaway→03None [14:23:36] 06SRE, 06Infrastructure-Foundations, 10Puppet-Core, 07Puppet (Puppet 7.0): puppet7: drop instances of :undef in erb files - https://phabricator.wikimedia.org/T341071#12226668 (10LSobanski) [14:24:38] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Migrate remaining container build/report steps from build2001 to build2004 - https://phabricator.wikimedia.org/T417389#12226677 (10MoritzMuehlenhoff) p:05Medium→03High [14:24:58] (03CR) 10Atsuko: [C:03+2] airflow3: updating pyspark [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326804 (https://phabricator.wikimedia.org/T435172) (owner: 10Atsuko) [14:27:02] 06SRE, 06Infrastructure-Foundations, 10Puppet-Core, 07Puppet (Puppet 7.0): puppet7: drop instances of :undef in erb files - https://phabricator.wikimedia.org/T341071#12226687 (10LSobanski) p:05Medium→03Low [14:27:12] (03Merged) 10jenkins-bot: airflow3: updating pyspark [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326804 (https://phabricator.wikimedia.org/T435172) (owner: 10Atsuko) [14:29:33] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for eamedina - https://phabricator.wikimedia.org/T435200#12226719 (10Nikerabbit) Approved. [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1430) [14:30:42] 06SRE, 06Infrastructure-Foundations, 10Mail, 10Wikimedia-Mailing-lists, 07Upstream: lists.wikimedia.org - adhere to RFC8048 (one-click unsubscribe) dkim guidelines - https://phabricator.wikimedia.org/T355802#12226738 (10LSobanski) a:05jhathaway→03None [14:32:33] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12226752 (10VRiley-WMF) hey @fgiunchedi so the reason I think it may have to do with some of the scripts is because 1048, 1049, and 1051 are Dell R640's the res... [14:33:21] 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T433348#12226772 (10VRiley-WMF) a:05VRiley-WMF→03BTullis [14:33:47] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2043.codfw.wmnet [14:37:48] (03PS1) 10Jforrester: CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler [extensions/CategoryTree] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326852 (https://phabricator.wikimedia.org/T435161) [14:37:57] (03PS1) 10Jforrester: CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326853 (https://phabricator.wikimedia.org/T435161) [14:38:19] andre: Cherry-picks good to deploy ^. [14:38:52] (03CR) 10Scott French: [C:03+1] docker_registry: route /v2/releng to its new s3 backend [puppet] - 10https://gerrit.wikimedia.org/r/1326323 (https://phabricator.wikimedia.org/T432829) (owner: 10Elukey) [14:39:53] (03CR) 10Clare Ming: [C:03+2] Deploy GrowthBook 5.0.0 to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326844 (https://phabricator.wikimedia.org/T434098) (owner: 10Santiago Faci) [14:40:09] (03CR) 10Ayounsi: [C:03+1] Remove cumin2002 from RAPI access rules [puppet] - 10https://gerrit.wikimedia.org/r/1326842 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [14:40:49] jmm@cumin2003 drain-node (PID 2796305) is awaiting input [14:42:24] (03Merged) 10jenkins-bot: Deploy GrowthBook 5.0.0 to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326844 (https://phabricator.wikimedia.org/T434098) (owner: 10Santiago Faci) [14:42:32] (03CR) 10Lucas Werkmeister (WMDE): [C:03+1] Disable Wikidata Bridge on Catalan Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326848 (https://phabricator.wikimedia.org/T433713) (owner: 10Sadiya.mohammed13) [14:45:40] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2043.codfw.wmnet [14:47:45] PROBLEM - Host dse-k8s-etcd2003 is DOWN: PING CRITICAL - Packet loss = 100% [14:48:22] James_F: thanks a lot! gonna take a look after the Phab deploy, heh [14:48:35] Yeah, definitely focus on that. :-) [14:49:08] (03CR) 10Elukey: [C:03+2] profile::docker_registry: scan the Releng's endpoint too [puppet] - 10https://gerrit.wikimedia.org/r/1326849 (https://phabricator.wikimedia.org/T427175) (owner: 10Elukey) [14:50:42] (03CR) 10Aklapper: [C:04-1] "Do not merge. Blocked by older Wikibase release branches depending on this, see https://phabricator.wikimedia.org/T405596#12220781 for det" [puppet] - 10https://gerrit.wikimedia.org/r/1298763 (https://phabricator.wikimedia.org/T405596) (owner: 10Aklapper) [14:51:01] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2043.codfw.wmnet [14:51:09] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2043.codfw.wmnet [14:51:15] RECOVERY - Host dse-k8s-etcd2003 is UP: PING OK - Packet loss = 0%, RTA = 31.97 ms [14:51:38] (03CR) 10Scott French: [C:03+1] "Thanks, Luca! Seems reasonable to me." [puppet] - 10https://gerrit.wikimedia.org/r/1326835 (https://phabricator.wikimedia.org/T427175) (owner: 10Elukey) [14:53:57] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1236.eqiad.wmnet with OS bookworm [14:55:50] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply [14:56:51] !log jmm@cumin2003 START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2044.codfw.wmnet [14:57:01] !log arnaudb@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance [14:57:03] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply [15:00:05] jelto, arnoldokoth, mutante, and arnaudb: SRE Collaboration Services office hours (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1500). Please do the needful. [15:02:07] !log brennen@deploy1003 Started deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for T435213 [15:02:13] T435213: Deploy Phab/Phorge 2026-08-18 - https://phabricator.wikimedia.org/T435213 [15:02:39] jmm@cumin2003 drain-node (PID 2801147) is awaiting input [15:03:04] !log brennen@deploy1003 Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab2003 for T435213 (duration: 00m 57s) [15:03:20] !log brennen@deploy1003 Started deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for T435213 [15:03:28] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host ganeti2044.codfw.wmnet [15:04:21] !log brennen@deploy1003 Finished deploy [phabricator/deployment@6b9b6ff]: deploy phab1004 for T435213 (duration: 01m 01s) [15:04:27] (03PS1) 10Elukey: setup.py: add setuptools to tests env [software/spicerack] - 10https://gerrit.wikimedia.org/r/1326858 [15:05:24] PROBLEM - Host kubestagemaster2003 is DOWN: PING CRITICAL - Packet loss = 100% [15:05:36] PROBLEM - Host ml-etcd2003 is DOWN: PING CRITICAL - Packet loss = 100% [15:05:49] (03CR) 10AikoChou: [C:03+2] ml-services: update ores-legacy image tag [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326834 (https://phabricator.wikimedia.org/T429675) (owner: 10AikoChou) [15:05:58] PROBLEM - Host dse-k8s-ctrl2002 is DOWN: PING CRITICAL - Packet loss = 100% [15:08:10] (03Merged) 10jenkins-bot: ml-services: update ores-legacy image tag [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326834 (https://phabricator.wikimedia.org/T429675) (owner: 10AikoChou) [15:08:48] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ganeti2044.codfw.wmnet [15:09:07] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet [15:09:14] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2044.codfw.wmnet [15:09:16] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for BGerdemann_(WMF) - https://phabricator.wikimedia.org/T433313#12226955 (10Bethany) Hello. Sorry for the confusion and delayed response. I was OOO. I had changed by request to level 2 after initially requesting level 3. The July... [15:09:32] FIRING: [2x] KubernetesCalicoDown: dse-k8s-ctrl2002.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [15:09:36] (03PS3) 10Elukey: docker_registry: improve registry-homepage-builder.py [puppet] - 10https://gerrit.wikimedia.org/r/1326835 (https://phabricator.wikimedia.org/T427175) [15:09:36] (03PS4) 10Elukey: docker_registry: route /v2/releng to its new s3 backend [puppet] - 10https://gerrit.wikimedia.org/r/1326323 (https://phabricator.wikimedia.org/T432829) [15:09:50] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage [15:09:56] (03CR) 10Elukey: docker_registry: improve registry-homepage-builder.py (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1326835 (https://phabricator.wikimedia.org/T427175) (owner: 10Elukey) [15:10:28] RECOVERY - Host dse-k8s-ctrl2002 is UP: PING OK - Packet loss = 0%, RTA = 30.52 ms [15:10:30] RECOVERY - Host kubestagemaster2003 is UP: PING WARNING - Packet loss = 66%, RTA = 38.81 ms [15:10:38] RECOVERY - Host ml-etcd2003 is UP: PING OK - Packet loss = 0%, RTA = 32.13 ms [15:10:38] 07sre-alert-triage, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Alert in need of triage: Dell PowerEdge or Supermicro Broadcom RAID Controller (instance an-worker1208) - https://phabricator.wikimedia.org/T430138#12226957 (10BTullis) 05Open→03Resolved a:03BTullis This has been fixed as part of {T42... [15:12:40] !log failover ganeti master in codfw to ganeti2047 [15:12:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:13:21] FIRING: [3x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:13:23] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1236.eqiad.wmnet with reason: host reimage [15:13:40] FIRING: KubernetesRsyslogDown: rsyslog on dse-k8s-worker2001:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://grafana.wikimedia.org/d/OagQjQmnk?var-server=dse-k8s-worker2001 - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [15:14:17] fceratto@cumin1003 makevm (PID 824663) is awaiting input [15:14:32] RESOLVED: [2x] KubernetesCalicoDown: dse-k8s-ctrl2002.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [15:14:58] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet [15:15:51] PROBLEM - ganeti-wconfd running on ganeti2048 is CRITICAL: PROCS CRITICAL: 0 processes with UID = 110 (gnt-masterd), command name ganeti-wconfd https://wikitech.wikimedia.org/wiki/Ganeti [15:16:47] (03PS5) 10Dduvall: drivers: Driver registration and factory [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1326398 (https://phabricator.wikimedia.org/T434957) [15:18:40] FIRING: [2x] KubernetesRsyslogDown: rsyslog on dse-k8s-worker2001:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [15:18:41] (03CR) 10Marostegui: [C:03+1] "I'd suggest you stop mariadb first, and then merge this. And then clean up /srv/sqldadata.s6" [puppet] - 10https://gerrit.wikimedia.org/r/1326830 (https://phabricator.wikimedia.org/T409557) (owner: 10FNegri) [15:20:15] (03CR) 10Clément Goubert: "I think it's fine, just keep an eye on `mw-cron` and maybe run a few `mw-script` tests. Adding @rlazarus@wikimedia.org for `mw-script` eye" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325884 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:20:24] (03CR) 10Clément Goubert: [C:03+1] mediawiki: stop using lamp.deployment in job.yaml.tpl. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325884 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:20:44] !log aikochou@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . [15:20:48] (03CR) 10Elukey: [C:03+2] docker_registry: improve registry-homepage-builder.py [puppet] - 10https://gerrit.wikimedia.org/r/1326835 (https://phabricator.wikimedia.org/T427175) (owner: 10Elukey) [15:21:46] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326848 (https://phabricator.wikimedia.org/T433713) (owner: 10Sadiya.mohammed13) [15:21:52] (03CR) 10Clément Goubert: "Hmm actually one thing to check is that the mesh will not fail if it doesn't have apache backing on the public port if there is one." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325884 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:23:40] RESOLVED: [2x] KubernetesRsyslogDown: rsyslog on dse-k8s-worker2001:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [15:24:09] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12227027 (10Jhancock.wm) you can start mariadb. We can probably close this ticket for now too. i got a reply from dell to open a new ticket when if it crashes again. They don't have anythi... [15:24:30] (03PS1) 10Federico Ceratto: preseed.yaml: Reorder MariaDB test-s8 section VMs Bug: T435059 [puppet] - 10https://gerrit.wikimedia.org/r/1326861 (https://phabricator.wikimedia.org/T435059) [15:24:52] (03PS2) 10Federico Ceratto: preseed.yaml: Reorder MariaDB test-s8 section VMs [puppet] - 10https://gerrit.wikimedia.org/r/1326861 (https://phabricator.wikimedia.org/T435059) [15:24:55] (03CR) 10Marostegui: [C:03+1] "Copied votes on follow-up patch sets have been updated:" [puppet] - 10https://gerrit.wikimedia.org/r/1326861 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [15:26:39] !log aikochou@deploy1003 helmfile [ml-serve-codfw] 'sync' command on namespace 'ores-legacy' for release 'main' . [15:26:39] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2209 firmware upgrade after a crash - https://phabricator.wikimedia.org/T431952#12227031 (10Marostegui) 05Open→03Resolved Thanks - I will start mariadb and repool the host. The fact that they say we can open a new ticket if it crashes again is probably cause... [15:27:43] (03CR) 10CI reject: [V:04-1] preseed.yaml: Reorder MariaDB test-s8 section VMs [puppet] - 10https://gerrit.wikimedia.org/r/1326861 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [15:28:52] (03CR) 10FNegri: "sounds good, I will do this tomorrow and ping you in IRC!" [puppet] - 10https://gerrit.wikimedia.org/r/1326830 (https://phabricator.wikimedia.org/T409557) (owner: 10FNegri) [15:29:26] !log aikochou@deploy1003 helmfile [ml-serve-eqiad] 'sync' command on namespace 'ores-legacy' for release 'main' . [15:31:52] (03CR) 10Blake: "Is there a way I might go about testing this for a single job?" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325884 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [15:32:51] (03CR) 10Ayounsi: [C:03+1] Remove cumin2002 from the homer peer list of cumin1003 [puppet] - 10https://gerrit.wikimedia.org/r/1326847 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [15:33:10] (03CR) 10BCornwall: [C:04-1] "I do not agree with utilizing mtail for analytical purposes - mtail is in need of retirement longer-term: It is particularly bad for perfo" [puppet] - 10https://gerrit.wikimedia.org/r/1324778 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [15:37:14] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1236.eqiad.wmnet with OS bookworm [15:38:23] 06SRE, 10SRE-Access-Requests: Requesting access to 'restricted' for nicholusmuwonge - https://phabricator.wikimedia.org/T435168#12227098 (10dancy) >>! In T435168#12225818, @JMeybohm wrote: > The user is member of the `nda` group, so NDA has been signed. > > @dancy || @thcipriani please sign off as `restricted... [15:40:56] (03CR) 10Marostegui: [C:03+1] "Feel free to merge anytime, this is a noop. I created https://phabricator.wikimedia.org/T435219 so we can remove them from prod." [puppet] - 10https://gerrit.wikimedia.org/r/1326845 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [15:41:05] !log bounce PIC 0/0 on cr1-magru to set port to 40G [15:41:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:43:57] PROBLEM - OSPF status on cr2-magru is CRITICAL: OSPFv2: 1/3 UP : OSPFv3: 1/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:44:06] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-magru - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [15:44:18] PROBLEM - Host cr1-magru is DOWN: PING CRITICAL - Packet loss = 100% [15:44:18] PROBLEM - Host cr1-magru IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [15:44:32] !ack [15:44:33] 8253 (ACKED) Host cr1-magru [15:44:51] FIRING: [2x] SwitchCoreInterfaceDown: Switch core interface down - asw1-b3-magru:et-0/0/48 (Core: cr1-magru:et-0/0/1) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:45:24] (03CR) 10Federico Ceratto: [C:03+2] preseed.yaml: Reorder MariaDB test-s8 section VMs [puppet] - 10https://gerrit.wikimedia.org/r/1326861 (https://phabricator.wikimedia.org/T435059) (owner: 10Federico Ceratto) [15:45:40] FIRING: [6x] CoreBGPDown: Core BGP session down between asw1-b3-magru and cr1-magru (195.200.68.142) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:45:57] RECOVERY - OSPF status on cr2-magru is OK: OSPFv2: 3/3 UP : OSPFv3: 3/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:46:11] FIRING: [4x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:47:55] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1145.eqiad.wmnet with OS bookworm [15:47:57] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1199.eqiad.wmnet with OS bookworm [15:48:03] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1200.eqiad.wmnet with OS bookworm [15:48:29] (03CR) 10Marostegui: [C:03+1] Remove cumin2002 as mariadb root client [puppet] - 10https://gerrit.wikimedia.org/r/1326843 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [15:49:06] RESOLVED: NetworkDeviceAlarmActive: Alarm active on cr1-magru - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [15:49:20] RECOVERY - Host cr1-magru is UP: PING OK - Packet loss = 0%, RTA = 167.39 ms [15:49:21] RECOVERY - Host cr1-magru IPv6 is UP: PING OK - Packet loss = 0%, RTA = 167.70 ms [15:49:51] RESOLVED: [2x] SwitchCoreInterfaceDown: Switch core interface down - asw1-b3-magru:et-0/0/48 (Core: cr1-magru:et-0/0/1) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:50:03] PROBLEM - OSPF status on cr1-magru is CRITICAL: OSPFv2: 2/3 UP : OSPFv3: 2/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:50:25] !log installing zip security updates [15:50:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:50:27] (03PS1) 10Jforrester: Flow: Allow null $html in the CategoryViewerGenerateLink handler [extensions/Flow] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326865 (https://phabricator.wikimedia.org/T435161) [15:50:40] RESOLVED: [8x] CoreBGPDown: Core BGP session down between asw1-b3-magru and cr1-magru (195.200.68.142) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:51:11] FIRING: [4x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:51:13] (03CR) 10BCornwall: [V:03+1] "From my read, `prometheus::blackbox::check::tcp` also includes a cert expiration check when `$use_tls` is set (which `$force_tls` will act" [puppet] - 10https://gerrit.wikimedia.org/r/1307419 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [15:51:42] (03CR) 10Ryan Kemper: [C:03+2] move-vlan: only query target host for inplace [cookbooks] - 10https://gerrit.wikimedia.org/r/1324839 (https://phabricator.wikimedia.org/T434737) (owner: 10Ryan Kemper) [15:51:48] jouncebot nowandnext [15:51:48] For the next 0 hour(s) and 8 minute(s): SRE Collaboration Services office hours (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1500) [15:51:49] In 0 hour(s) and 8 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1600) [15:52:47] 10ops-eqiad, 06SRE, 10Cloud-VPS, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12227216 (10Andrew) Any reason we can't just flip these over to uefi? [15:53:24] FIRING: [8x] CoreBGPDown: Core BGP session down between asw1-b3-magru and cr1-magru (195.200.68.142) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:53:28] (03CR) 10CI reject: [V:04-1] Flow: Allow null $html in the CategoryViewerGenerateLink handler [extensions/Flow] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326865 (https://phabricator.wikimedia.org/T435161) (owner: 10Jforrester) [15:55:55] (03Merged) 10jenkins-bot: move-vlan: only query target host for inplace [cookbooks] - 10https://gerrit.wikimedia.org/r/1324839 (https://phabricator.wikimedia.org/T434737) (owner: 10Ryan Kemper) [15:57:43] (03CR) 10BCornwall: [V:03+1 C:03+1] "From my read, `prometheus::blackbox::check::tcp` also includes a cert expiration check when `$use_tls` is set (which `$force_tls` will act" [puppet] - 10https://gerrit.wikimedia.org/r/1307419 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [16:00:05] jhathaway and rzl: Your horoscope predicts another Puppet request window deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1600). [16:00:05] cjming and dancy: A patch you scheduled for Puppet request window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [16:01:02] o/ [16:02:08] hello hello [16:02:10] o/ [16:02:26] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage [16:02:35] (03CR) 10RLazarus: [C:03+2] tk_constructive_edits: use foreachwikiindblist instead and ignore errors [puppet] - 10https://gerrit.wikimedia.org/r/1325555 (https://phabricator.wikimedia.org/T431493) (owner: 10DLynch) [16:02:37] (03CR) 10RLazarus: [C:03+2] ci: add ReadingLists to gitcache [puppet] - 10https://gerrit.wikimedia.org/r/1324278 (https://phabricator.wikimedia.org/T403560) (owner: 10Phedenskog) [16:02:55] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage [16:03:43] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage [16:03:47] cjming: will you want to kick off a test run, or just wait for the next scheduled one? [16:03:55] dancy: and anything for you to test? [16:04:07] rzl: No testing needed. [16:04:14] cool [16:04:28] i can't recall off the top of my head how to initiate a test run - do you have the cmd handy? [16:05:02] also fine to wait for next scheduled one [16:05:04] yep, it's https://wikitech.wikimedia.org/wiki/Mw-cron_jobs#Manually_running_a_CronJob - but let me run puppet on the deployment server first [16:05:27] !log sudo -i cookbook sre.cdn.roll-upgrade-ats --query 'A:cp-esams' --task-id T434478 --reason '9.2.15 upgrade' [16:05:30] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:05:35] !log cdobbins@cumin1003 START - Cookbook sre.cdn.roll-upgrade-ats Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade (T434478) [16:05:42] that's running now - by the time it finishes, we might be pretty close to the 18th minute of the hour anyway :) but I'll let you know [16:06:03] (03CR) 10Jforrester: "recheck" [extensions/Flow] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326865 (https://phabricator.wikimedia.org/T435161) (owner: 10Jforrester) [16:06:09] RESOLVED: [2x] CoreBGPDown: Core BGP session down between cr1-magru and cr2-eqiad (195.200.68.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=magru&var-device=cr1-magru:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [16:06:16] thanks - i do it infrequently enough that i always have to look it up [16:06:32] cjming: Don't worry about it, I do it regularly and I still have to look it up [16:06:36] me too [16:06:42] and I don't do it nearly as much as claime [16:06:55] gtk! [16:07:11] (03PS1) 10DLynch: LLMSuggestionsEditCheck: don't over-cache the description [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326870 (https://phabricator.wikimedia.org/T428641) [16:07:34] (03PS1) 10DLynch: LLMSuggestionsEditCheck: don't over-cache the description [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326871 (https://phabricator.wikimedia.org/T428641) [16:07:41] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326870 (https://phabricator.wikimedia.org/T428641) (owner: 10DLynch) [16:07:48] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326871 (https://phabricator.wikimedia.org/T428641) (owner: 10DLynch) [16:08:01] (03CR) 10Hnowlan: [C:03+1] prometheus: Add redis_lock job [puppet] - 10https://gerrit.wikimedia.org/r/1326792 (https://phabricator.wikimedia.org/T427999) (owner: 10Clément Goubert) [16:08:13] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:09:01] (03CR) 10Clément Goubert: [C:03+2] prometheus: Add redis_lock job [puppet] - 10https://gerrit.wikimedia.org/r/1326792 (https://phabricator.wikimedia.org/T427999) (owner: 10Clément Goubert) [16:09:43] (03PS1) 10Pppery: Update source strings [phabricator/translations] (wmf/stable) - 10https://gerrit.wikimedia.org/r/1326873 [16:09:45] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1199.eqiad.wmnet with reason: host reimage [16:11:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr1-magru and cr2-eqiad (195.200.68.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=magru&var-device=cr1-magru:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [16:11:53] cjming: puppet's finished -- you're welcome to either start a run now or just eyeball the next one [16:12:13] if it turns out there's a problem and you need a followup change merged, just ping me, no need to wait for the next window [16:12:33] i'll eyeball since it's 6 minutes away -- thanks so much rzl! [16:12:51] of course! [16:13:13] FIRING: [3x] JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:13:44] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1200.eqiad.wmnet with reason: host reimage [16:18:20] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1145.eqiad.wmnet with reason: host reimage [16:19:52] rzl: Thanks! [16:22:25] PROBLEM - Host titan1002 is DOWN: PING CRITICAL - Packet loss = 100% [16:22:39] (03PS1) 10Reedy: Add banner notifying of upcoming 2FA enforcement [extensions/OATHAuth] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326875 (https://phabricator.wikimedia.org/T420792) [16:23:15] RECOVERY - Host titan1002 is UP: PING OK - Packet loss = 0%, RTA = 0.32 ms [16:23:21] FIRING: [4x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:24:07] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12227382 (10MoritzMuehlenhoff) [16:24:55] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12227388 (10MoritzMuehlenhoff) [16:28:21] FIRING: [4x] ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:29:43] (03CR) 10Ladsgroup: svg: use rsvg-convert's language parameter (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1042203 (https://phabricator.wikimedia.org/T261192) (owner: 10Hnowlan) [16:32:25] (03PS11) 10Hnowlan: svg: use rsvg-convert's language parameter [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1042203 (https://phabricator.wikimedia.org/T261192) [16:32:50] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1199.eqiad.wmnet with OS bookworm [16:34:13] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1200.eqiad.wmnet with OS bookworm [16:34:17] PROBLEM - SSH on an-launcher1003 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [16:35:28] (03PS1) 10Urbanecm: Echo: Start using virtual domains [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) [16:35:47] 06SRE, 06Infrastructure-Foundations, 10Puppet-Infrastructure: Fix remaining scoped legacy fact usage - https://phabricator.wikimedia.org/T435225 (10jhathaway) 03NEW [16:35:59] 06SRE, 06Infrastructure-Foundations, 10Puppet-Infrastructure: Fix remaining scoped legacy fact usage - https://phabricator.wikimedia.org/T435225#12227444 (10jhathaway) a:03jhathaway [16:36:07] RECOVERY - SSH on an-launcher1003 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [16:36:31] (03CR) 10Urbanecm: [C:04-2] "not until the Echo code is in prod" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) (owner: 10Urbanecm) [16:41:20] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1145.eqiad.wmnet with OS bookworm [16:41:37] !log dancy@deploy1003 Installing scap version "4.282.0" for 3 host(s) [16:43:32] !log dancy@deploy1003 Installation of scap version "4.282.0" completed for 3 hosts [16:43:51] !log dancy@deploy1003 Started scap sync-world: Testing T434726 [16:43:56] T434726: Where is my change running tool - https://phabricator.wikimedia.org/T434726 [16:50:32] !log dancy@deploy1003 Finished scap sync-world: Testing T434726 (duration: 06m 40s) [16:50:36] T434726: Where is my change running tool - https://phabricator.wikimedia.org/T434726 [16:53:29] (03PS1) 10Ahmon Dancy: w/deployment-info.php: Handle new file format [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326885 (https://phabricator.wikimedia.org/T434726) [16:55:46] jouncebot nowandnext [16:55:46] For the next 0 hour(s) and 4 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1600) [16:55:46] In 0 hour(s) and 4 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1700) [16:57:12] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dancy@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326885 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1700) [17:01:51] (03Merged) 10jenkins-bot: w/deployment-info.php: Handle new file format [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326885 (https://phabricator.wikimedia.org/T434726) (owner: 10Ahmon Dancy) [17:02:19] !log dancy@deploy1003 Started scap sync-world: Backport for [[gerrit:1326885|w/deployment-info.php: Handle new file format (T434726)]] [17:02:24] T434726: Where is my change running tool - https://phabricator.wikimedia.org/T434726 [17:04:24] (03CR) 10Hnowlan: svg: use rsvg-convert's language parameter (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1042203 (https://phabricator.wikimedia.org/T261192) (owner: 10Hnowlan) [17:04:35] !log dancy@deploy1003 dancy: Backport for [[gerrit:1326885|w/deployment-info.php: Handle new file format (T434726)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:05:22] !log dancy@deploy1003 dancy: Continuing with deployment [17:08:37] !log jasmine@cumin2002 START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie [17:08:58] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes, 13Patch-For-Review: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12227635 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-ctrl2006.cod... [17:09:34] !log dancy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326885|w/deployment-info.php: Handle new file format (T434726)]] (duration: 07m 15s) [17:09:39] T434726: Where is my change running tool - https://phabricator.wikimedia.org/T434726 [17:14:03] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1146.eqiad.wmnet with OS bookworm [17:14:04] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1156.eqiad.wmnet with OS bookworm [17:14:07] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1177.eqiad.wmnet with OS bookworm [17:14:08] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1201.eqiad.wmnet with OS bookworm [17:14:11] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1202.eqiad.wmnet with OS bookworm [17:21:58] !log jasmine@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage [17:23:51] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/0 (Transit: EdgeUno (E1-SER-7853-IP)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:24:39] FIRING: [2x] TransitBGPDown: Transit BGP session down between cr2-magru and EdgeUno (200.25.58.212) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [17:24:59] andre: Do you need help deploying the UBN train fix? [17:25:51] James_F: I'm basically waiting for the usual train window starting in 35min to backport, not to potentially interfere with anything else as I didn't smell a sense of urgency yet [17:25:59] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage [17:26:05] Ack. [17:26:24] No rush from my end, just wanted to check if you were expecting me to do it. :-) [17:26:29] thanks :) [17:27:55] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage [17:28:01] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage [17:28:21] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage [17:28:51] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/0 (Transit: EdgeUno (E1-SER-7853-IP)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:29:32] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage [17:30:02] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage [17:31:51] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/0 (Transit: EdgeUno (E1-SER-7853-IP)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:34:10] jouncebot: nowandnext [17:34:10] For the next 0 hour(s) and 25 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1700) [17:34:10] In 0 hour(s) and 25 minute(s): MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1800) [17:34:39] RESOLVED: [2x] TransitBGPDown: Transit BGP session down between cr2-magru and EdgeUno (200.25.58.212) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DTransitBGPDown [17:34:47] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1201.eqiad.wmnet with reason: host reimage [17:36:51] RESOLVED: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/0 (Transit: EdgeUno (E1-SER-7853-IP)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:38:03] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1202.eqiad.wmnet with reason: host reimage [17:39:45] (03CR) 10Ladsgroup: "Random suggestion, just introduce the virtual domain now. Wait a couple of weeks and then remove old stuff at your leisure 😄" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) (owner: 10Urbanecm) [17:41:55] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1177.eqiad.wmnet with reason: host reimage [17:42:33] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie [17:42:50] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes, 13Patch-For-Review: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12227844 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-ctrl2006.codfw.w... [17:43:52] (03CR) 10Ladsgroup: [C:03+2] "tested on mw-experimental" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326380 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [17:44:16] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326380 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [17:45:32] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.cdn.roll-upgrade-ats (exit_code=0) Rolling upgrade of ATS on A:cp-esams and A:cp - 9.2.15 upgrade (T434478) [17:45:56] (03PS2) 10Andrew Bogott: backy2: avoid parallel read/write options [puppet] - 10https://gerrit.wikimedia.org/r/1326378 (https://phabricator.wikimedia.org/T434014) [17:46:05] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1146.eqiad.wmnet with reason: host reimage [17:48:13] (03CR) 10Andrew Bogott: [C:03+2] backy2: avoid parallel read/write options [puppet] - 10https://gerrit.wikimedia.org/r/1326378 (https://phabricator.wikimedia.org/T434014) (owner: 10Andrew Bogott) [17:48:30] (03Merged) 10jenkins-bot: Introduce main lock manager service [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326380 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [17:48:49] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1156.eqiad.wmnet with reason: host reimage [17:48:51] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1326380|Introduce main lock manager service (T366938 T427999)]] [17:49:03] T366938: Reduce relying on database locks - https://phabricator.wikimedia.org/T366938 [17:49:05] T427999: Redis solution for LockManager - https://phabricator.wikimedia.org/T427999 [17:49:06] (03PS1) 10Lerickson: Enable posting to eventgate from the proxy. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326893 (https://phabricator.wikimedia.org/T433375) [17:49:34] (03PS2) 10Lerickson: WIP: Enable posting to eventgate from the proxy. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326893 (https://phabricator.wikimedia.org/T433375) [17:50:55] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1326380|Introduce main lock manager service (T366938 T427999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:53:18] (03PS2) 10Urbanecm: Echo: Start using virtual domains [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) [17:53:23] (03CR) 10Urbanecm: "Fair enough, done." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) (owner: 10Urbanecm) [17:54:09] 10ops-eqiad, 06DC-Ops: Unresponsive management for an-worker1147.mgmt:22 - https://phabricator.wikimedia.org/T435240 (10phaultfinder) 03NEW [17:54:49] 06SRE, 10SRE-Access-Requests, 06Data-Platform-SRE: Requesting access to Analytics Data Lake for mkrolik/mkrolik-wmf - https://phabricator.wikimedia.org/T434877#12227958 (10Ottomata) [17:55:26] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1201.eqiad.wmnet with OS bookworm [17:57:15] (03PS1) 10Ladsgroup: Disable redis lock manager on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326895 (https://phabricator.wikimedia.org/T366938) [17:57:24] (03CR) 10CI reject: [V:04-1] Disable redis lock manager on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326895 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [17:57:58] (03PS2) 10Ladsgroup: Disable redis lock manager on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326895 (https://phabricator.wikimedia.org/T366938) [17:58:37] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1202.eqiad.wmnet with OS bookworm [17:58:44] !log ladsgroup@deploy1003 ladsgroup: Rolling back deployment [18:00:05] andre and brennen: #bothumor I � Unicode. All rise for MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1800). [18:00:08] o/ [18:00:15] o/ [18:00:17] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326380|Introduce main lock manager service (T366938 T427999)]] (duration: 11m 25s) [18:00:21] yell if i can be of assistance, andre [18:00:26] T366938: Reduce relying on database locks - https://phabricator.wikimedia.org/T366938 [18:00:26] T427999: Redis solution for LockManager - https://phabricator.wikimedia.org/T427999 [18:00:29] brennen, thanks, I will [18:00:48] (03CR) 10Ladsgroup: [C:03+2] Disable redis lock manager on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326895 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [18:00:53] I won't deploy this [18:00:57] Amir1: seems that you're done? [18:01:02] ah :) [18:01:03] just rebase [18:01:35] one sec and I'll be done [18:01:36] sorry [18:01:39] Amir1: np [18:02:06] (03Merged) 10jenkins-bot: Disable redis lock manager on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326895 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [18:03:50] andre: okay, I have rebased it. Feel free to push but if you see a lot of errors related to redis, let me know [18:04:08] this really shouldn't cause issues but you never know [18:04:23] Amir1: I'll have to do some backports and then run the delayed train, let's see what else happens :D [18:04:28] <3 [18:05:38] * andre going to scap backport 1326852 1326853 1326865, then run the train [18:07:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aklapper@deploy1003 using scap backport" [extensions/CategoryTree] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326852 (https://phabricator.wikimedia.org/T435161) (owner: 10Jforrester) [18:07:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aklapper@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326853 (https://phabricator.wikimedia.org/T435161) (owner: 10Jforrester) [18:07:04] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aklapper@deploy1003 using scap backport" [extensions/Flow] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326865 (https://phabricator.wikimedia.org/T435161) (owner: 10Jforrester) [18:07:55] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1177.eqiad.wmnet with OS bookworm [18:08:13] (03Merged) 10jenkins-bot: CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler [extensions/CategoryTree] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326852 (https://phabricator.wikimedia.org/T435161) (owner: 10Jforrester) [18:08:41] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1146.eqiad.wmnet with OS bookworm [18:09:26] (03PS1) 10Krinkle: ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326896 (https://phabricator.wikimedia.org/T434469) [18:09:34] (03PS1) 10Krinkle: ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326897 (https://phabricator.wikimedia.org/T434469) [18:10:23] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326896 (https://phabricator.wikimedia.org/T434469) (owner: 10Krinkle) [18:10:30] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326897 (https://phabricator.wikimedia.org/T434469) (owner: 10Krinkle) [18:11:57] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-experimental: apply [18:11:59] (03Merged) 10jenkins-bot: CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326853 (https://phabricator.wikimedia.org/T435161) (owner: 10Jforrester) [18:12:03] (03Merged) 10jenkins-bot: Flow: Allow null $html in the CategoryViewerGenerateLink handler [extensions/Flow] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326865 (https://phabricator.wikimedia.org/T435161) (owner: 10Jforrester) [18:12:29] !log aklapper@deploy1003 Started scap sync-world: Backport for [[gerrit:1326852|CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853|CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865|Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] [18:12:34] T435161: TypeError: MediaWiki\HookContainer\HookRunner::onCategoryViewerGenerateLink(): Argument #4 ($html) must be of type string, null given, called in /srv/mediawiki/php-1.47.0-wmf.16/includes/Category/CategoryViewer.php on line 235 - https://phabricator.wikimedia.org/T435161 [18:12:38] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply [18:12:42] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts beta module [puppet] - 10https://gerrit.wikimedia.org/r/1326898 (https://phabricator.wikimedia.org/T435225) [18:13:26] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1156.eqiad.wmnet with OS bookworm [18:14:35] !log aklapper@deploy1003 jforrester, aklapper: Backport for [[gerrit:1326852|CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853|CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865|Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki [18:14:35] /Mwdebug). Changes can now be verified there. [18:17:49] !log jasmine@cumin2002 START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie [18:18:07] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes, 13Patch-For-Review: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12228089 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-ctrl2006.cod... [18:18:11] !log aklapper@deploy1003 jforrester, aklapper: Continuing with deployment [18:22:26] !log aklapper@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326852|CategoryTree: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]], [[gerrit:1326853|CategoryViewer: Allow null $html in the CategoryViewerGenerateLink hook (T435161)]], [[gerrit:1326865|Flow: Allow null $html in the CategoryViewerGenerateLink handler (T435161)]] (duration: 09m 57s) [18:22:31] T435161: TypeError: MediaWiki\HookContainer\HookRunner::onCategoryViewerGenerateLink(): Argument #4 ($html) must be of type string, null given, called in /srv/mediawiki/php-1.47.0-wmf.16/includes/Category/CategoryViewer.php on line 235 - https://phabricator.wikimedia.org/T435161 [18:25:19] 06SRE, 10fundraising-tech-ops, 06Infrastructure-Foundations, 10netops: Upgrade JunOS on pfw1-eqiad and pfw1-codfw - https://phabricator.wikimedia.org/T434865#12228113 (10cmooney) As discussed on irc I'll kick off the upgrade on pfw1-codfw at 09:00am UTC tomrorow (Wed Aug 19th). [18:29:19] (03PS1) 10JHathaway: admin::hashuser: fix gid check [puppet] - 10https://gerrit.wikimedia.org/r/1326900 [18:30:31] (03PS1) 10Eevans: sessionstore: Upgrade to Cassandra 5.0.8 (canary) [puppet] - 10https://gerrit.wikimedia.org/r/1326901 (https://phabricator.wikimedia.org/T435154) [18:30:32] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1326900 (owner: 10JHathaway) [18:31:02] !log jasmine@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage [18:34:38] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl2006.codfw.wmnet with reason: host reimage [18:36:33] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.16 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326903 (https://phabricator.wikimedia.org/T430835) [18:36:36] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by aklapper@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326903 (https://phabricator.wikimedia.org/T430835) (owner: 10TrainBranchBot) [18:37:06] (03PS1) 10Eevans: sessionstore: Upgrade to Cassandra 5.0.8 (complete) [puppet] - 10https://gerrit.wikimedia.org/r/1326904 (https://phabricator.wikimedia.org/T435154) [18:37:09] (03PS1) 10Eevans: sessionstore: set storage compatability to `UPGRADING` [puppet] - 10https://gerrit.wikimedia.org/r/1326905 (https://phabricator.wikimedia.org/T435154) [18:37:11] (03PS1) 10Eevans: sessionstore: set storage compatability to `NONE` [puppet] - 10https://gerrit.wikimedia.org/r/1326906 (https://phabricator.wikimedia.org/T435154) [18:37:22] (03CR) 10Eevans: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1326901 (https://phabricator.wikimedia.org/T435154) (owner: 10Eevans) [18:37:38] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.16 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326903 (https://phabricator.wikimedia.org/T430835) (owner: 10TrainBranchBot) [18:40:47] (03PS1) 10Ladsgroup: Revert "Disable redis lock manager on testwiki" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326907 [18:43:22] (03CR) 10Ladsgroup: [C:03+2] svg: use rsvg-convert's language parameter (031 comment) [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1042203 (https://phabricator.wikimedia.org/T261192) (owner: 10Hnowlan) [18:43:38] (03PS1) 10JHathaway: Puppet 8: Replace legacy facts in module interface [puppet] - 10https://gerrit.wikimedia.org/r/1326909 (https://phabricator.wikimedia.org/T372666) [18:44:02] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1326909 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [18:44:23] !log aklapper@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs T430835 [18:44:29] T430835: 1.47.0-wmf.16 deployment blockers - https://phabricator.wikimedia.org/T430835 [18:44:57] (03CR) 10Reedy: [C:03+2] Add banner notifying of upcoming 2FA enforcement [extensions/OATHAuth] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326875 (https://phabricator.wikimedia.org/T420792) (owner: 10Reedy) [18:45:27] (03CR) 10Ladsgroup: Echo: Start using virtual domains (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) (owner: 10Urbanecm) [18:48:09] (03Merged) 10jenkins-bot: Add banner notifying of upcoming 2FA enforcement [extensions/OATHAuth] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326875 (https://phabricator.wikimedia.org/T420792) (owner: 10Reedy) [18:49:37] (03CR) 10Kamila Součková: [C:03+1] wmnet: Reduce _etcd._tcp.conftool (R/W) SRV TTL to 10s [dns] - 10https://gerrit.wikimedia.org/r/1326363 (https://phabricator.wikimedia.org/T435103) (owner: 10Scott French) [18:50:16] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie [18:50:31] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware, 07Kubernetes, 13Patch-For-Review: wikikube-ctrl2006 implementation tracking - https://phabricator.wikimedia.org/T406596#12228191 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-ctrl2006.codfw.w... [18:50:46] (03Merged) 10jenkins-bot: svg: use rsvg-convert's language parameter [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1042203 (https://phabricator.wikimedia.org/T261192) (owner: 10Hnowlan) [18:50:53] (03CR) 10Kamila Součková: [C:03+1] wmnet: Restore _etcd._tcp.conftool (R/W) SRV TTL to 5M [dns] - 10https://gerrit.wikimedia.org/r/1326365 (https://phabricator.wikimedia.org/T435103) (owner: 10Scott French) [18:51:16] (03CR) 10Kamila Součková: [C:03+1] wmnet: Switch _etcd._tcp.conftool (R/W) SRV hosts to eqiad [dns] - 10https://gerrit.wikimedia.org/r/1326364 (https://phabricator.wikimedia.org/T435103) (owner: 10Scott French) [18:51:19] (03PS2) 10Reedy: InitialiseSettings: Enable 2FA enforcement on more private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324751 (https://phabricator.wikimedia.org/T428103) [18:51:57] (03CR) 10Reedy: [C:03+2] InitialiseSettings: Enable 2FA enforcement on more private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324751 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [18:52:47] (03CR) 10Kamila Součková: [C:03+1] hieradata: etcd read-only in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1326355 (https://phabricator.wikimedia.org/T435103) (owner: 10Scott French) [18:52:55] (03Merged) 10jenkins-bot: InitialiseSettings: Enable 2FA enforcement on more private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324751 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [18:53:39] (03CR) 10Kamila Součková: [C:03+1] hieradata: switch etcd replication from eqiad to codfw [puppet] - 10https://gerrit.wikimedia.org/r/1326356 (https://phabricator.wikimedia.org/T435103) (owner: 10Scott French) [18:54:39] (03CR) 10Kamila Součková: [C:03+1] hieradata: etcd read-write in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1326357 (https://phabricator.wikimedia.org/T435103) (owner: 10Scott French) [18:54:40] !log reedy@deploy1003 Started scap sync-world: Backport for [[gerrit:1324751|InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875|Add banner notifying of upcoming 2FA enforcement (T420792)]] [18:54:47] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [18:54:47] T420792: Allow 2FA to be enforced for all accounts on a private wiki - https://phabricator.wikimedia.org/T420792 [19:06:56] (03CR) 10RLazarus: "mwscript-k8s has (shhhhh) a secret flag --helmfile. (https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/product" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1325884 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [19:07:18] PROBLEM - SSH on an-launcher1003 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [19:08:18] RECOVERY - SSH on an-launcher1003 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [19:11:48] (03PS3) 10Urbanecm: Echo: Start using virtual domains [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) [19:11:56] (03CR) 10Urbanecm: Echo: Start using virtual domains (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) (owner: 10Urbanecm) [19:12:36] (03PS1) 10Ladsgroup: thumbor: use rsvg-convert's language parameter [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326912 (https://phabricator.wikimedia.org/T261192) [19:12:53] !log reedy@deploy1003 reedy: Backport for [[gerrit:1324751|InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875|Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [19:13:00] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [19:13:01] T420792: Allow 2FA to be enforced for all accounts on a private wiki - https://phabricator.wikimedia.org/T420792 [19:13:15] !log reedy@deploy1003 reedy: Continuing with deployment [19:14:22] (03CR) 10Ladsgroup: [C:03+1] "Thank you <3" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326881 (https://phabricator.wikimedia.org/T380385) (owner: 10Urbanecm) [19:14:27] !log rebooting kafkamon1003.eqiad.wmnet T435162 [19:14:30] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:15:41] !log rebooting kafkamon2003.codfw.wmnet - T435162 [19:15:44] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:16:43] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm [19:16:47] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm [19:16:52] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm [19:16:55] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm [19:16:58] !log btullis@cumin1003 START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm [19:18:23] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [19:21:55] 10ops-magru: Too low optic power on - cr2-magru:xe-0/1/0 - https://phabricator.wikimedia.org/T434275#12228295 (10cmooney) 05Open→03Resolved a:03cmooney EdgeUno took a little bit of pressing but after coming back to say they seen no issue they sent someone on site and resolved the issue (we observed the... [19:26:26] !log reedy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324751|InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875|Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) [19:26:35] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [19:26:36] T420792: Allow 2FA to be enforced for all accounts on a private wiki - https://phabricator.wikimedia.org/T420792 [19:27:22] (03CR) 10Ahmon Dancy: [V:03+1 C:03+1] "Tested in beta, verified no changes to the wmf-beta-update-all.service definition on deployment-deploy04." [puppet] - 10https://gerrit.wikimedia.org/r/1326898 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:27:37] (03CR) 10Ahmon Dancy: [V:03+1 C:03+1] Puppet 8: Replace legacy facts beta module (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1326898 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [19:29:28] (03PS1) 10Ahmon Dancy: Merge remote-tracking branch 'origin/master' into train-dev [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1326919 [19:30:17] PROBLEM - SSH on an-launcher1003 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [19:31:00] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage [19:31:14] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage [19:32:11] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage [19:32:54] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage [19:32:57] !log btullis@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage [19:33:02] (03CR) 10Ahmon Dancy: [C:03+2] Merge remote-tracking branch 'origin/master' into train-dev [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1326919 (owner: 10Ahmon Dancy) [19:34:05] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage [19:34:13] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326907 (owner: 10Ladsgroup) [19:34:18] (03Merged) 10jenkins-bot: Merge remote-tracking branch 'origin/master' into train-dev [mediawiki-config] (train-dev) - 10https://gerrit.wikimedia.org/r/1326919 (owner: 10Ahmon Dancy) [19:34:24] ugh [19:34:25] sorry [19:34:56] I appreciate the extra approval! [19:35:00] (03CR) 10Ladsgroup: [C:04-2] "one sec" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326907 (owner: 10Ladsgroup) [19:35:16] Hmm. . Did spiderpig accept that change as input? [19:35:35] and/or scap backport. It should have rejected it for being a weird branch. [19:36:00] why, the branch looks correct to me? :D [19:36:17] oh sorry. I was confusing our two outputs. [19:36:21] Disregard. :-) [19:36:46] yeah, it was my fault. I started a scap backport, then you +2'ed the patch, it got confused [19:36:57] please let me know when I can push this change [19:37:37] I'm working on the `train-dev` branch so you're good to go. [19:37:47] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage [19:39:30] Amir1: ^^ [19:39:36] ah okay [19:39:36] thanks [19:39:51] (03CR) 10Ladsgroup: "good to go" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326907 (owner: 10Ladsgroup) [19:41:08] RECOVERY - SSH on an-launcher1003 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [19:41:16] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326907 (owner: 10Ladsgroup) [19:41:20] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage [19:42:10] (03Merged) 10jenkins-bot: Revert "Disable redis lock manager on testwiki" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326907 (owner: 10Ladsgroup) [19:42:32] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1326907|Revert "Disable redis lock manager on testwiki"]] [19:44:41] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1326907|Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [19:46:06] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [19:47:11] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage [19:50:33] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage [19:51:11] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [19:53:38] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326907|Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) [19:54:57] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm [19:57:04] (03PS1) 10Ladsgroup: Enable redis lock manager on s6 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326923 (https://phabricator.wikimedia.org/T366938) [19:58:52] jouncebot: nowandnext [19:58:52] For the next 0 hour(s) and 1 minute(s): MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T1800) [19:58:53] In 0 hour(s) and 1 minute(s): UTC late backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T2000) [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: How many deployers does it take to do UTC late backport window deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T2000). [20:00:05] sadiya_wmde28, kemayo, and Krinkle: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:01:59] o/ [20:02:07] I can deploy for myself. [20:02:32] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm [20:02:58] ...which I guess I will do, since nobody else has said they're here yet. [20:05:45] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm [20:08:11] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm [20:08:49] !log zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # T428406 [20:08:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:08:54] T428406: old file revisions missing of File:A_Warm_Shade_of_Ivory_-_Henry_Mancini_album_cover.jpg - https://phabricator.wikimedia.org/T428406 [20:09:03] o/ [20:09:11] Kemayo: happy to go after you, self service too [20:09:49] (03CR) 10Ladsgroup: [C:03+2] thumbor: use rsvg-convert's language parameter [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326912 (https://phabricator.wikimedia.org/T261192) (owner: 10Ladsgroup) [20:10:18] PROBLEM - SSH on an-launcher1003 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [20:10:54] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326870 (https://phabricator.wikimedia.org/T428641) (owner: 10DLynch) [20:10:55] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kemayo@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326871 (https://phabricator.wikimedia.org/T428641) (owner: 10DLynch) [20:11:39] FIRING: [2x] CoreBGPDown: Core BGP session down between cr1-magru and cr2-eqiad (195.200.68.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=magru&var-device=cr1-magru:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [20:12:08] RECOVERY - SSH on an-launcher1003 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [20:12:16] (03Merged) 10jenkins-bot: thumbor: use rsvg-convert's language parameter [deployment-charts] - 10https://gerrit.wikimedia.org/r/1326912 (https://phabricator.wikimedia.org/T261192) (owner: 10Ladsgroup) [20:12:22] (03Merged) 10jenkins-bot: LLMSuggestionsEditCheck: don't over-cache the description [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326870 (https://phabricator.wikimedia.org/T428641) (owner: 10DLynch) [20:13:28] FIRING: JobUnavailable: Reduced availability for job liberica in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:14:22] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm [20:20:43] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [20:21:43] Sorry, I have a quibble CI job just sitting there saying it's 100% done and not actually handing off. Whenever it decides it's really done, the rest will be quick. [20:23:14] (03Merged) 10jenkins-bot: LLMSuggestionsEditCheck: don't over-cache the description [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326871 (https://phabricator.wikimedia.org/T428641) (owner: 10DLynch) [20:23:17] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [20:23:37] !log kemayo@deploy1003 Started scap sync-world: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] [20:23:42] T428641: Build PoC for integrating initial set of MoS LLM-generated suggestions into Suggestion Mode as `experimental` - https://phabricator.wikimedia.org/T428641 [20:25:44] !log kemayo@deploy1003 kemayo: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:26:44] !log `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash [20:26:44] !log kemayo@deploy1003 kemayo: Continuing with deployment [20:26:46] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:28:37] FIRING: ProbeDown: Service ganeti3005:1811 has failed probes (tcp_ganeti_noded_ip4) - https://wikitech.wikimedia.org/wiki/Ganeti - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:28:55] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [20:31:12] !log kemayo@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) [20:31:13] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [20:31:15] Krinkle: Okay, you're up. [20:31:17] T428641: Build PoC for integrating initial set of MoS LLM-generated suggestions into Suggestion Mode as `experimental` - https://phabricator.wikimedia.org/T428641 [20:32:33] Kemayo: thx [20:33:45] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [20:34:38] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326896 (https://phabricator.wikimedia.org/T434469) (owner: 10Krinkle) [20:34:39] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326897 (https://phabricator.wikimedia.org/T434469) (owner: 10Krinkle) [20:35:49] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [20:41:06] (03PS1) 10Medelius: Exclude NPOV-type LLM-generated suggestions [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326925 (https://phabricator.wikimedia.org/T435253) [20:41:26] (03PS1) 10Medelius: Exclude NPOV-type LLM-generated suggestions [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326926 (https://phabricator.wikimedia.org/T435253) [20:41:53] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326925 (https://phabricator.wikimedia.org/T435253) (owner: 10Medelius) [20:42:12] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, August 18 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326926 (https://phabricator.wikimedia.org/T435253) (owner: 10Medelius) [20:45:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:47:54] (03Merged) 10jenkins-bot: ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook [extensions/Math] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326896 (https://phabricator.wikimedia.org/T434469) (owner: 10Krinkle) [20:48:06] (03Merged) 10jenkins-bot: ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook [extensions/Math] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326897 (https://phabricator.wikimedia.org/T434469) (owner: 10Krinkle) [20:48:29] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] [20:48:40] T434469: Client side MathJax not applied when using Live Preview - https://phabricator.wikimedia.org/T434469 [20:48:40] T419356: MathJax rendering missing after saving edit with VisualEditor - https://phabricator.wikimedia.org/T419356 [20:48:41] T422077: MathJax missing from DiscussionTools reply preview - https://phabricator.wikimedia.org/T422077 [20:49:37] oops never pressed enter on my log line [20:49:38] !log `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory [20:49:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [20:50:32] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:52:10] (03PS6) 10BryanDavis: httpbb: Add test suite for testwiki [puppet] - 10https://gerrit.wikimedia.org/r/1323829 (https://phabricator.wikimedia.org/T428972) [20:52:41] (03PS1) 10Medelius: LLMSuggestionsEditCheck: final comparison should also have the object-replacements [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326928 [20:52:52] (03PS1) 10Medelius: LLMSuggestionsEditCheck: final comparison should also have the object-replacements [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326929 [20:58:08] !log krinkle@deploy1003 krinkle: Continuing with deployment [20:58:52] i know it's near the end, but would i be able to backport after Krinkle? can do it myself [20:59:22] (03CR) 10BryanDavis: "Tests exercised from deploy1003.eqiad.wmnet via an scp'ed copy:" [puppet] - 10https://gerrit.wikimedia.org/r/1323829 (https://phabricator.wikimedia.org/T428972) (owner: 10BryanDavis) [20:59:24] I’d like to get a security patch out during the Readers window, if that’s ok. [21:00:04] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260818T2100) [21:00:22] I also have a patch I want to push :D [21:00:38] I can go later though (go dinner, and come back) [21:01:58] (03PS6) 10Scott French: trafficserver: Support testwiki pretrain routing in XWD [puppet] - 10https://gerrit.wikimedia.org/r/1304189 (https://phabricator.wikimedia.org/T427666) [21:02:00] (03PS2) 10Scott French: hieradata: Divert testwiki traffic to mw-pretrain at ATS [puppet] - 10https://gerrit.wikimedia.org/r/1326930 (https://phabricator.wikimedia.org/T427666) [21:02:26] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) [21:02:38] T434469: Client side MathJax not applied when using Live Preview - https://phabricator.wikimedia.org/T434469 [21:02:39] T419356: MathJax rendering missing after saving edit with VisualEditor - https://phabricator.wikimedia.org/T419356 [21:02:39] T422077: MathJax missing from DiscussionTools reply preview - https://phabricator.wikimedia.org/T422077 [21:03:58] cmede: all yours I suppose :) [21:04:56] alright, thanks [21:05:15] (03CR) 10TrainBranchBot: [C:03+2] "Approved by caro@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326925 (https://phabricator.wikimedia.org/T435253) (owner: 10Medelius) [21:05:15] (03CR) 10TrainBranchBot: [C:03+2] "Approved by caro@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326926 (https://phabricator.wikimedia.org/T435253) (owner: 10Medelius) [21:05:16] (03CR) 10TrainBranchBot: [C:03+2] "Approved by caro@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326929 (owner: 10Medelius) [21:05:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by caro@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326928 (owner: 10Medelius) [21:05:25] 🚢 [21:06:00] (03CR) 10Subramanya Sastry: [C:03+1] "Adding Mateus & sergio & Scott as FYI, but +1ing on their behalf." [puppet] - 10https://gerrit.wikimedia.org/r/1326808 (https://phabricator.wikimedia.org/T434959) (owner: 10JMeybohm) [21:09:33] (03PS2) 10Krinkle: Remove unused/redundant wgMFNoindexPages=true setting [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1264845 (https://phabricator.wikimedia.org/T255458) [21:13:11] cmede: if you can let me know when you’re finished, that would be great, thanks. [21:13:25] sure thing! [21:13:55] sbassett: and please let me know when you're done :D [21:15:13] Amir1: of course! [21:15:26] thank you ^_^ [21:17:35] (03Merged) 10jenkins-bot: Exclude NPOV-type LLM-generated suggestions [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326925 (https://phabricator.wikimedia.org/T435253) (owner: 10Medelius) [21:17:39] (03Merged) 10jenkins-bot: LLMSuggestionsEditCheck: final comparison should also have the object-replacements [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326929 (owner: 10Medelius) [21:17:41] (03Merged) 10jenkins-bot: LLMSuggestionsEditCheck: final comparison should also have the object-replacements [extensions/VisualEditor] (wmf/1.47.0-wmf.15) - 10https://gerrit.wikimedia.org/r/1326928 (owner: 10Medelius) [21:18:39] (03CR) 10CI reject: [V:04-1] Exclude NPOV-type LLM-generated suggestions [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326926 (https://phabricator.wikimedia.org/T435253) (owner: 10Medelius) [21:21:20] (03CR) 10TrainBranchBot: [C:03+2] "Approved by caro@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326926 (https://phabricator.wikimedia.org/T435253) (owner: 10Medelius) [21:24:23] !log ryankemper@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair T434494 [21:24:28] T434494: Migrate production hadoop cluster to bookworm - https://phabricator.wikimedia.org/T434494 [21:26:37] (03CR) 10RLazarus: [C:03+2] Hand PageAssesment ownership to Content-Platform-Team [puppet] - 10https://gerrit.wikimedia.org/r/1326808 (https://phabricator.wikimedia.org/T434959) (owner: 10JMeybohm) [21:30:36] (03Merged) 10jenkins-bot: Exclude NPOV-type LLM-generated suggestions [extensions/VisualEditor] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1326926 (https://phabricator.wikimedia.org/T435253) (owner: 10Medelius) [21:31:03] !log caro@deploy1003 Started scap sync-world: Backport for [[gerrit:1326925|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] [21:31:08] T435253: Exclude NPOV suggestions from LLM-generated MoS suggestions - https://phabricator.wikimedia.org/T435253 [21:33:06] !log caro@deploy1003 caro: Backport for [[gerrit:1326925|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h [21:33:06] ttps://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:33:11] checking [21:34:14] !log caro@deploy1003 caro: Continuing with deployment [21:34:38] cmede: Looks good to me. [21:34:54] thanks! [21:35:58] (No NPOV on the page I know had one, and the page I was testing where I know one overran a citation has started working.) [21:38:29] !log caro@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326925|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 [21:38:29] 7m 26s) [21:38:35] T435253: Exclude NPOV suggestions from LLM-generated MoS suggestions - https://phabricator.wikimedia.org/T435253 [21:38:46] sbassett, you're good to go, sorry about that! [21:44:09] cmede: thanks! [21:49:14] (03PS1) 10RLazarus: maintenance: Change owner of purge_loginnotify to PSI [puppet] - 10https://gerrit.wikimedia.org/r/1326937 [21:52:35] (03CR) 10RLazarus: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1326937 (owner: 10RLazarus) [21:54:56] !log Deployed security fix for T435234 (wmf.15) [21:54:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:04:13] !log Deployed security fix for T435234 (wmf.16) [22:04:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:04:28] Amir1: ok, I should be done now, thanks. [22:08:04] Thanks! [22:08:54] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326923 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [22:09:29] (03CR) 10Dreamy Jazz: [C:03+1] maintenance: Change owner of purge_loginnotify to PSI [puppet] - 10https://gerrit.wikimedia.org/r/1326937 (owner: 10RLazarus) [22:09:39] (03CR) 10Scott French: "Thank you very much, Bryan!" [puppet] - 10https://gerrit.wikimedia.org/r/1323829 (https://phabricator.wikimedia.org/T428972) (owner: 10BryanDavis) [22:09:47] (03CR) 10Dreamy Jazz: [C:03+1] "LGTM from a PSI team standpoint" [puppet] - 10https://gerrit.wikimedia.org/r/1326937 (owner: 10RLazarus) [22:10:33] (03Merged) 10jenkins-bot: Enable redis lock manager on s6 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326923 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [22:10:55] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1326923|Enable redis lock manager on s6 (T366938)]] [22:11:00] T366938: Reduce relying on database locks - https://phabricator.wikimedia.org/T366938 [22:12:59] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1326923|Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [22:18:30] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [22:22:47] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326923|Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) [22:22:54] T366938: Reduce relying on database locks - https://phabricator.wikimedia.org/T366938 [22:33:12] !log ryankemper@cumin2003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm [22:35:36] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm [22:35:42] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm [22:35:49] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm [22:35:56] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm [22:36:07] !log ryankemper@cumin2003 START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm [22:41:44] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [22:50:24] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage [22:51:11] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage [22:51:13] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage [22:53:39] !log ryankemper@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage [22:54:29] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage [22:54:49] (03PS7) 10BryanDavis: httpbb: Add test suite for testwiki [puppet] - 10https://gerrit.wikimedia.org/r/1323829 (https://phabricator.wikimedia.org/T428972) [22:54:49] (03PS1) 10BryanDavis: httpbb: sort managed directory list [puppet] - 10https://gerrit.wikimedia.org/r/1326939 [22:57:09] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage [22:58:37] (03CR) 10BryanDavis: httpbb: Add test suite for testwiki (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1323829 (https://phabricator.wikimedia.org/T428972) (owner: 10BryanDavis) [23:00:23] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage [23:05:01] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage [23:07:46] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12229083 (10andrea.denisse) >>! In T434268#12221369, @elukey wrote: >> I think that along with the upgrade host maybe we... [23:09:36] (03PS2) 10Reedy: InitialiseSettings: Enable 2FA banners on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324752 (https://phabricator.wikimedia.org/T428103) [23:09:45] (03CR) 10CI reject: [V:04-1] InitialiseSettings: Enable 2FA banners on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324752 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [23:09:57] (03PS3) 10Reedy: InitialiseSettings: Enable 2FA banners on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324752 (https://phabricator.wikimedia.org/T428103) [23:15:14] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm [23:15:14] (03PS1) 10Ladsgroup: Retire filebackend lock manager in favour of the default one [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326944 (https://phabricator.wikimedia.org/T366938) [23:16:04] (03CR) 10CI reject: [V:04-1] Retire filebackend lock manager in favour of the default one [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326944 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [23:16:28] (03PS2) 10Ladsgroup: Retire filebackend lock manager in favour of the default one [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326944 (https://phabricator.wikimedia.org/T366938) [23:16:33] (03CR) 10RLazarus: [C:03+1] envoy: New upstream version 1.39.0. [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1326818 (https://phabricator.wikimedia.org/T421418) (owner: 10Blake) [23:17:17] (03PS3) 10Ladsgroup: Retire filebackend lock manager in favour of the default one [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326944 (https://phabricator.wikimedia.org/T366938) [23:18:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [23:19:47] (03PS4) 10Ladsgroup: Retire filebackend lock manager in favour of the default one [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326944 (https://phabricator.wikimedia.org/T366938) [23:20:42] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm [23:21:10] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm [23:27:42] ryankemper@cumin2003 reimage (PID 2886867) is awaiting input [23:27:52] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm [23:28:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ladsgroup@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326944 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [23:29:15] (03Merged) 10jenkins-bot: Retire filebackend lock manager in favour of the default one [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326944 (https://phabricator.wikimedia.org/T366938) (owner: 10Ladsgroup) [23:29:38] !log ladsgroup@deploy1003 Started scap sync-world: Backport for [[gerrit:1326944|Retire filebackend lock manager in favour of the default one (T366938)]] [23:29:43] T366938: Reduce relying on database locks - https://phabricator.wikimedia.org/T366938 [23:31:43] !log ladsgroup@deploy1003 ladsgroup: Backport for [[gerrit:1326944|Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [23:31:44] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv6 ping to magru RIPE Atlas anchor: failures over threshold for measurement 95133216 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [23:34:19] !log ladsgroup@deploy1003 ladsgroup: Continuing with deployment [23:38:33] !log ladsgroup@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326944|Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) [23:38:38] T366938: Reduce relying on database locks - https://phabricator.wikimedia.org/T366938 [23:41:22] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1326947 [23:41:22] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1326947 (owner: 10TrainBranchBot) [23:44:04] Amir1: done? [23:44:13] si [23:44:33] (03CR) 10TrainBranchBot: [C:03+2] "Approved by krinkle@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1264845 (https://phabricator.wikimedia.org/T255458) (owner: 10Krinkle) [23:45:33] (03Merged) 10jenkins-bot: Remove unused/redundant wgMFNoindexPages=true setting [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1264845 (https://phabricator.wikimedia.org/T255458) (owner: 10Krinkle) [23:45:52] !log krinkle@deploy1003 Started scap sync-world: Backport for [[gerrit:1264845|Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] [23:45:58] T255458: Enable $wgMFNoindexPages for all wikis - https://phabricator.wikimedia.org/T255458 [23:47:59] !log krinkle@deploy1003 krinkle: Backport for [[gerrit:1264845|Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [23:48:52] (03CR) 10CI reject: [V:04-1] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1326947 (owner: 10TrainBranchBot) [23:49:34] (03CR) 10Scott French: httpbb: Add test suite for testwiki (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1323829 (https://phabricator.wikimedia.org/T428972) (owner: 10BryanDavis) [23:50:47] (03CR) 10Scott French: [C:03+1] "Thanks, Bryan! I'll merge this along with the patch earlier in the series." [puppet] - 10https://gerrit.wikimedia.org/r/1326939 (owner: 10BryanDavis) [23:51:23] !log krinkle@deploy1003 krinkle: Continuing with deployment [23:51:26] FIRING: [2x] BFDdown: BFD session down between cr2-eqiad and 195.200.68.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [23:55:35] !log krinkle@deploy1003 Finished scap sync-world: Backport for [[gerrit:1264845|Remove unused/redundant wgMFNoindexPages=true setting (T255458)]] (duration: 09m 42s) [23:55:43] T255458: Enable $wgMFNoindexPages for all wikis - https://phabricator.wikimedia.org/T255458