[00:05:20] !log DEPLOYED Refinery at 4e7a2b32 for changes: pageview allowlist 1305158 (+min.wikiquote) 1305162 (+bol.wikipedia), 1305156 (+isv.wikipedia); 1305980 (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop 1295064 (+globalimagelinks) 1295069 (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally) [00:05:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [00:11:08] (03PS1) 10Ladsgroup: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1306797 (https://phabricator.wikimedia.org/T372666) [00:11:21] PROBLEM - Host cirrussearch2103 is DOWN: PING CRITICAL - Packet loss = 100% [00:11:42] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1306797 (https://phabricator.wikimedia.org/T372666) (owner: 10Ladsgroup) [00:11:47] RECOVERY - Host cirrussearch2103 is UP: PING OK - Packet loss = 0%, RTA = 33.07 ms [00:13:17] PROBLEM - Host cirrussearch2102 is DOWN: PING CRITICAL - Packet loss = 100% [00:14:47] RECOVERY - Host cirrussearch2102 is UP: PING OK - Packet loss = 0%, RTA = 31.67 ms [00:15:01] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie [00:18:37] RECOVERY - SSH on urldownloader2005 is OK: SSH OK - OpenSSH_10.0p2 Debian-7+deb13u4 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [00:20:11] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie [00:21:42] RESOLVED: [2x] ProbeDown: Service urldownloader2005:8080 has failed probes (http_url_downloader_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/Url-downloader - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [00:21:45] (03PS2) 10Ladsgroup: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1306797 (https://phabricator.wikimedia.org/T372666) [00:22:17] (03CR) 10CI reject: [V:04-1] Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1306797 (https://phabricator.wikimedia.org/T372666) (owner: 10Ladsgroup) [00:24:23] (03PS3) 10Ladsgroup: Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1306797 (https://phabricator.wikimedia.org/T372666) [00:27:40] (03CR) 10Ladsgroup: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306797 (https://phabricator.wikimedia.org/T372666) (owner: 10Ladsgroup) [00:27:43] PROBLEM - Host es1039 #page is DOWN: PING CRITICAL - Packet loss = 100% [00:29:31] PROBLEM - MariaDB Replica IO: es7 #page on es1035 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@es1039.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on es1039.eqiad.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:29:33] PROBLEM - MariaDB Replica IO: es7 #page on es1040 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@es1039.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on es1039.eqiad.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:29:34] PROBLEM - MariaDB Replica IO: es7 #page on es1048 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@es1039.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on es1039.eqiad.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:29:37] PROBLEM - MariaDB Replica IO: es7 #page on es2039 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@es1039.eqiad.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on es1039.eqiad.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:29:54] sirenbot: stfu [00:30:00] sirenbot: !ack [00:30:21] FIRING: MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?panelId=18&fullscreen&orgId=1&var-datasource=codfw%20prometheus/ops - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [00:30:35] RECOVERY - Host es1039 #page is UP: PING OK - Packet loss = 0%, RTA = 0.35 ms [00:31:33] PROBLEM - MariaDB read only es7 #page on es1039 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [00:32:31] PROBLEM - MariaDB Event Scheduler es7 on es1039 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [00:32:33] PROBLEM - mysqld processes #page on es1039 is CRITICAL: PROCS CRITICAL: 0 processes with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [00:33:31] PROBLEM - MariaDB Events es7 on es1039 is CRITICAL: CRITICAL - Failed to query events: ERROR 2002 (HY000): Cant connect to local server through socket /run/mysqld/mysqld.sock (2) https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [00:33:33] PROBLEM - pt-heartbeat-wikimedia process on es1039 is CRITICAL: PROCS CRITICAL: 0 processes with args pt-heartbeat-wikimedia https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23pt-heartbeat [00:34:31] PROBLEM - MariaDB Replica Lag: es7 #page on es1035 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 604.12 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:34:33] PROBLEM - MariaDB Replica Lag: es7 #page on es1040 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 605.52 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:34:34] PROBLEM - MariaDB Replica Lag: es7 #page on es1048 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 605.94 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:34:37] PROBLEM - MariaDB Replica Lag: es7 #page on es2039 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 609.53 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:35:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [00:39:07] jesus [00:39:24] thing just rebooted, and I think the clock is wrong in SEL [00:39:33] it's been up since [00:40:11] first let me remove it from writes [00:41:06] let me know what I can do, if you need more hands [00:41:36] (03PS1) 10Gerrit maintenance bot: mariadb: Promote es1035 to es7 master [puppet] - 10https://gerrit.wikimedia.org/r/1306798 (https://phabricator.wikimedia.org/T430765) [00:41:42] (03PS1) 10Gerrit maintenance bot: wmnet: Update es7-master alias [dns] - 10https://gerrit.wikimedia.org/r/1306799 (https://phabricator.wikimedia.org/T430765) [00:42:22] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - T430765', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json [00:42:26] T430765: Switchover es7 master (es1039 -> es1035) - https://phabricator.wikimedia.org/T430765 [00:42:28] heya folks. I am also around if you need an extra pair of hands. I am assuming we are doing a switchover? [00:42:30] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074209 (10CDanis) p:05Triage→03High [00:42:33] ok thanks Amir1 [00:42:36] okay, the user impact should be gone now [00:42:40] thanks sukhe [00:42:54] cdanis: sorry you caught a bad one <3 [00:43:01] and thanks Amir1 -- I went to look at the External storage wikitech page and was met with a diagram that mentioned PMTPA [00:43:11] it's removed from writes [00:43:20] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074213 (10CDanis) [00:43:24] Amir1: can you later also tell us what steps you performed for posterity? [00:43:26] cdanis: yeah, it took me like five minutes to remember how to remove ES from the write pool [00:43:27] https://wikitech.wikimedia.org/wiki/Primary_database_switchover ? [00:43:37] later [00:44:03] it's a change I made that makes it removed from the RW pool of ES clusters and moves it to RO ones so replag wouldn't matter [00:44:03] (03CR) 10CI reject: [V:04-1] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1306795 (owner: 10TrainBranchBot) [00:44:09] sudo dbctl --scope eqiad section es7 ro "Maintenance - T430765" [00:44:09] sudo dbctl --scope codfw section es7 ro "Maintenance - T430765" [00:44:09] sudo dbctl config commit -m "Set es7 eqiad as read-only for maintenance - T430765" [00:44:24] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074215 (10CDanis) 00:42:36 okay, the user impact should be gone now 00:43:11 it's removed from writes Followups: documentation, documentation, documentation. [00:44:57] okay, now I need to bring it back online [00:45:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [00:45:18] the switchover would be much easier once the host is online [00:46:29] RECOVERY - mysqld processes #page on es1039 is OK: PROCS OK: 1 process with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [00:47:29] RECOVERY - MariaDB Events es7 on es1039 is OK: OK - All 2 events in ops database are ENABLED https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [00:47:31] RECOVERY - MariaDB Replica IO: es7 #page on es1035 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:47:32] RECOVERY - MariaDB Event Scheduler es7 on es1039 is OK: Version 10.11.16-MariaDB-log, Uptime 69s, read_only: True, event_scheduler: True, 14.53 QPS, connection latency: 0.015102s, query latency: 0.001073s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [00:47:33] RECOVERY - MariaDB Replica IO: es7 #page on es1040 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:47:34] RECOVERY - MariaDB Replica IO: es7 #page on es1048 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:47:37] RECOVERY - MariaDB Replica IO: es7 #page on es2039 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:49:29] RECOVERY - pt-heartbeat-wikimedia process on es1039 is OK: PROCS OK: 3 processes with args pt-heartbeat-wikimedia https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23pt-heartbeat [00:49:37] RECOVERY - MariaDB Replica Lag: es7 #page on es2039 is OK: OK slave_sql_lag Replication lag: 0.27 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:50:21] RESOLVED: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [00:50:31] RECOVERY - MariaDB Replica Lag: es7 #page on es1035 is OK: OK slave_sql_lag Replication lag: 0.07 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:50:33] RECOVERY - MariaDB Replica Lag: es7 #page on es1040 is OK: OK slave_sql_lag Replication lag: 0.00 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:50:34] RECOVERY - MariaDB Replica Lag: es7 #page on es1048 is OK: OK slave_sql_lag Replication lag: 0.00 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [00:51:03] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074230 (10Ladsgroup) For later: ` Jul 01 00:46:43 es1039 mysqld[5895]: 2026-07-01 0:46:43 6 [Warning] Detected table cache mutex contention at instance 1: 26% waits. Additional table cache instance cannot be a... [00:51:16] okay, now that everything is normal. I do the switchover [00:53:16] !log ladsgroup@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 T430765 [00:53:19] T430765: Switchover es7 master (es1039 -> es1035) - https://phabricator.wikimedia.org/T430765 [00:53:30] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Set es1035 with weight 0 T430765', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json [00:54:31] RECOVERY - MariaDB read only es7 #page on es1039 is OK: Version 10.11.16-MariaDB-log, Uptime 489s, read_only: False, event_scheduler: True, 31.83 QPS, connection latency: 0.028759s, query latency: 0.000833s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [00:55:01] 😌 [00:55:33] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074250 (10ssingh) ` < Amir1> it's a change I made that makes it removed from the RW pool of ES clusters and moves it to RO ones so replag wouldn't matter chBot) < Amir1> sudo dbctl --scope eqiad section es7 ro... [00:57:55] (03CR) 10Ladsgroup: [C:03+2] mariadb: Promote es1035 to es7 master [puppet] - 10https://gerrit.wikimedia.org/r/1306798 (https://phabricator.wikimedia.org/T430765) (owner: 10Gerrit maintenance bot) [00:58:26] !log Starting es7 eqiad failover from es1039 to es1035 - T430765 [00:58:28] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [00:58:29] T430765: Switchover es7 master (es1039 -> es1035) - https://phabricator.wikimedia.org/T430765 [01:00:03] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Promote es1035 to es7 primary T430765', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json [01:00:37] Amir1: for later, last question sorry -- are you following the steps at https://wikitech.wikimedia.org/wiki/Primary_database_switchover or is there something else we should reference? [01:00:46] thinking of time when you or another db won't be around [01:01:04] *dba [01:01:13] we follow the checklist outlined in the ticket [01:01:14] https://phabricator.wikimedia.org/T430765 [01:01:28] the checklist is produced by switchmaster (https://switchmaster.toolforge.org/ [01:01:50] thank you, updating the current task [01:03:26] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074260 (10ssingh) For the checklist on the switchover steps: ` < Amir1> we follow the checklist outlined in the ticket < Amir1> https://phabricator.wikimedia.org/T430765 < Amir1> the checklist is produced by s... [01:03:36] (03CR) 10Ladsgroup: [C:03+2] wmnet: Update es7-master alias [dns] - 10https://gerrit.wikimedia.org/r/1306799 (https://phabricator.wikimedia.org/T430765) (owner: 10Gerrit maintenance bot) [01:03:55] !log ladsgroup@dns1004 START - running authdns-update [01:05:52] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Depool es1039 T430765', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json [01:05:55] T430765: Switchover es7 master (es1039 -> es1035) - https://phabricator.wikimedia.org/T430765 [01:05:58] !log ladsgroup@dns1004 END - running authdns-update [01:06:53] cdanis: sukhe: Pooling es7 back for writes [01:07:00] thank you! [01:07:09] Amir1: <3 please make sure you take time off in lieu of this [01:07:17] !log ladsgroup@cumin1003 dbctl commit (dc=all): 'Set es7 eqiad back to read-write - T430765', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json [01:12:21] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1306800 [01:12:21] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1306800 (owner: 10TrainBranchBot) [01:14:24] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074275 (10Ladsgroup) The switchover is done, the cluster is RW now: https://grafana.wikimedia.org/d/000000278/mysql-aggregated?orgId=1&from=2026-07-01T00:12:34.145Z&to=2026-07-01T01:09:38.332Z&timezone=utc&var-... [01:16:02] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074276 (10Ladsgroup) Also it's important to check mariadb logs (the systemd service logs) to make sure things are not firework-y. The crash recovery mechanism of MariaDB is quite robust these days but you never... [01:16:28] I'm still around for a bit to finish my tiff clean up work. Ping me if there are issues [01:20:24] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1306800 (owner: 10TrainBranchBot) [01:24:58] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074277 (10Ladsgroup) And heartbeat needs a restart after crash (the `pt-heartbeat-wikimedia` systemd service) [01:25:49] that pmpta graph is even funnier: https://wikitech.wikimedia.org/wiki/File:External_storage_single_cluster.png [01:25:56] uploaded by 127.0.0.1 [01:45:42] I'm pretty sure I know who did that [02:00:40] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:03:19] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie [02:07:35] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 06m 54s) [02:09:42] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie [02:09:42] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:14:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:22:14] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage [02:26:55] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage [02:27:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:30:22] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage [02:31:45] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie [02:35:26] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage [02:44:42] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqsin:et-0/0/0 (Transport: Hurricane Electric (dc4841.sin1) {#}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:51:27] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage [02:51:41] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie [02:55:21] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie [02:59:25] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage [03:02:04] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [03:21:24] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie [03:29:29] PROBLEM - Host es1039 #page is DOWN: PING CRITICAL - Packet loss = 100% [03:30:37] ! incidents [03:30:44] !incidents [03:30:45] 8118 (UNACKED) Host es1039 (paged) [03:30:45] 8112 (RESOLVED) es1039 (paged)/MariaDB read only es7 (paged) [03:30:45] 8116 (RESOLVED) es1048 (paged)/MariaDB Replica Lag: es7 (paged) [03:30:45] 8114 (RESOLVED) es1040 (paged)/MariaDB Replica Lag: es7 (paged) [03:30:45] 8115 (RESOLVED) es1035 (paged)/MariaDB Replica Lag: es7 (paged) [03:30:46] 8117 (RESOLVED) es2039 (paged)/MariaDB Replica Lag: es7 (paged) [03:30:46] 8111 (RESOLVED) es2039 (paged)/MariaDB Replica IO: es7 (paged) [03:30:46] 8108 (RESOLVED) es1035 (paged)/MariaDB Replica IO: es7 (paged) [03:30:46] 8109 (RESOLVED) es1048 (paged)/MariaDB Replica IO: es7 (paged) [03:30:47] 8110 (RESOLVED) es1040 (paged)/MariaDB Replica IO: es7 (paged) [03:30:47] 8113 (RESOLVED) es1039 (paged)/mysqld processes (paged) [03:30:48] 8107 (RESOLVED) Host es1039 (paged) [03:30:59] !ack 8118 [03:30:59] 8118 (ACKED) Host es1039 (paged) [03:33:40] RECOVERY - Host es1039 #page is UP: PING OK - Packet loss = 0%, RTA = 0.34 ms [03:33:49] wow great [03:33:54] Nice [03:34:38] PROBLEM - mysqld processes #page on es1039 is CRITICAL: PROCS CRITICAL: 0 processes with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [03:34:38] PROBLEM - MariaDB Event Scheduler es7 on es1039 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [03:34:49] !incidents [03:34:50] 8119 (UNACKED) es1039 (paged)/mysqld processes (paged) [03:34:50] 8118 (RESOLVED) Host es1039 (paged) [03:34:50] 8112 (RESOLVED) es1039 (paged)/MariaDB read only es7 (paged) [03:34:50] 8116 (RESOLVED) es1048 (paged)/MariaDB Replica Lag: es7 (paged) [03:34:51] 8114 (RESOLVED) es1040 (paged)/MariaDB Replica Lag: es7 (paged) [03:34:51] 8115 (RESOLVED) es1035 (paged)/MariaDB Replica Lag: es7 (paged) [03:34:51] 8117 (RESOLVED) es2039 (paged)/MariaDB Replica Lag: es7 (paged) [03:34:51] 8111 (RESOLVED) es2039 (paged)/MariaDB Replica IO: es7 (paged) [03:34:51] 8108 (RESOLVED) es1035 (paged)/MariaDB Replica IO: es7 (paged) [03:34:52] 8109 (RESOLVED) es1048 (paged)/MariaDB Replica IO: es7 (paged) [03:34:52] 8110 (RESOLVED) es1040 (paged)/MariaDB Replica IO: es7 (paged) [03:34:53] 8113 (RESOLVED) es1039 (paged)/mysqld processes (paged) [03:34:53] 8107 (RESOLVED) Host es1039 (paged) [03:34:58] !ack 8119 [03:34:59] 8119 (ACKED) es1039 (paged)/mysqld processes (paged) [03:35:15] yeah es1039 has 2 min uptime [03:35:18] PROBLEM - MariaDB Replica SQL: es7 #page on es1039 is CRITICAL: CRITICAL slave_sql_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:35:19] PROBLEM - MariaDB Replica IO: es7 #page on es1039 is CRITICAL: CRITICAL slave_io_state could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:35:24] PROBLEM - MariaDB read only es7 on es1039 is CRITICAL: Could not connect to localhost:3306 https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [03:35:31] !ack [03:35:32] 8120 (ACKED) es1039 (paged)/MariaDB Replica SQL: es7 (paged) [03:35:32] 8121 (ACKED) es1039 (paged)/MariaDB Replica IO: es7 (paged) [03:35:38] PROBLEM - MariaDB Events es7 on es1039 is CRITICAL: CRITICAL - Failed to query events: ERROR 2002 (HY000): Cant connect to local server through socket /run/mysqld/mysqld.sock (2) https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [03:36:43] should we try depooling the host? [03:38:09] I think so, probably best to let a dba check it before using it, if it crashed [03:39:50] Is someone working on it. I'll just check the last puppet commit [03:39:52] I think the same thing happened a few hours ago, see https://phabricator.wikimedia.org/T430764 and cortobot [03:41:21] Okay, let's depool it [03:41:58] The ticket states "I leave es1039 depooled for HW inspection and what is wrong. Repool when needed." [03:42:05] let me check if the host is still depooled [03:42:18] PROBLEM - MariaDB Replica Lag: es7 #page on es1039 is CRITICAL: CRITICAL slave_sql_lag could not connect https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [03:42:26] There's also this: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1306798 [03:42:27] !incidents [03:42:28] 8119 (ACKED) es1039 (paged)/mysqld processes (paged) [03:42:28] 8120 (ACKED) es1039 (paged)/MariaDB Replica SQL: es7 (paged) [03:42:28] 8121 (ACKED) es1039 (paged)/MariaDB Replica IO: es7 (paged) [03:42:28] 8122 (UNACKED) es1039 (paged)/MariaDB Replica Lag: es7 (paged) [03:42:29] 8118 (RESOLVED) Host es1039 (paged) [03:42:29] 8112 (RESOLVED) es1039 (paged)/MariaDB read only es7 (paged) [03:42:29] 8116 (RESOLVED) es1048 (paged)/MariaDB Replica Lag: es7 (paged) [03:42:29] 8114 (RESOLVED) es1040 (paged)/MariaDB Replica Lag: es7 (paged) [03:42:29] 8115 (RESOLVED) es1035 (paged)/MariaDB Replica Lag: es7 (paged) [03:42:30] 8117 (RESOLVED) es2039 (paged)/MariaDB Replica Lag: es7 (paged) [03:42:30] 8111 (RESOLVED) es2039 (paged)/MariaDB Replica IO: es7 (paged) [03:42:31] 8108 (RESOLVED) es1035 (paged)/MariaDB Replica IO: es7 (paged) [03:42:31] 8109 (RESOLVED) es1048 (paged)/MariaDB Replica IO: es7 (paged) [03:42:32] 8110 (RESOLVED) es1040 (paged)/MariaDB Replica IO: es7 (paged) [03:42:32] 8113 (RESOLVED) es1039 (paged)/mysqld processes (paged) [03:42:33] 8107 (RESOLVED) Host es1039 (paged) [03:42:40] !ack 8118 [03:42:40] Attempt to ack incident 8118 failed. [03:42:53] !8119 [03:43:00] !8122 [03:43:13] !ack 8122 [03:43:14] 8122 (ACKED) es1039 (paged)/MariaDB Replica Lag: es7 (paged) [03:43:21] Weee :-) [03:43:57] yyeah that was probably the switch to another master ? es1039 -> es1035 ? [03:44:44] (03CR) 10Hashar: [C:04-1] "**Thank you so much Jaime!** I have replied on T411583#12074313 (Phabricator makes it easier to discover the info). I am keeping this co" [puppet] - 10https://gerrit.wikimedia.org/r/1306166 (https://phabricator.wikimedia.org/T257744) (owner: 10Arnaudb) [03:45:23] sudo dbctl instance es1039 get on cumin1003 returns [03:45:23] "pooled": false [03:45:39] Okay, then let's downtime it for a day or so [03:47:36] I can downtime the host and add a comment in the existing task + corto [03:47:36] !log slyngshede@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash [03:47:43] ah thank you [03:47:47] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074327 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=f965c095-744c-4fe7-997d-d637cbb7a225) set by slyngshede@cumin1003 for 1 day, 0:00:00 on 1 host(s) and their services with reason: Hardw... [03:48:03] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074329 (10SLyngshede-WMF) Host is stilled depooled. Adding downtime, for 1 day. [03:49:17] And the task is already there and tagged, so we can leave the dirty work for a DBA :-) [03:49:56] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074331 (10Jelto) for the record: the host crashed again and triggered multiple p.ages. ` jelto@es1039:~$ uptime 03:35:02 up 2 min, 1 user, load average: 0.13, 0.05, 0.01 ` [03:50:20] yep, I'll try to go to sleep again [03:50:31] !incidents [03:50:31] 8119 (ACKED) es1039 (paged)/mysqld processes (paged) [03:50:31] 8120 (ACKED) es1039 (paged)/MariaDB Replica SQL: es7 (paged) [03:50:32] 8121 (ACKED) es1039 (paged)/MariaDB Replica IO: es7 (paged) [03:50:32] 8122 (ACKED) es1039 (paged)/MariaDB Replica Lag: es7 (paged) [03:50:32] 8118 (RESOLVED) Host es1039 (paged) [03:50:32] 8112 (RESOLVED) es1039 (paged)/MariaDB read only es7 (paged) [03:50:32] 8116 (RESOLVED) es1048 (paged)/MariaDB Replica Lag: es7 (paged) [03:50:33] 8114 (RESOLVED) es1040 (paged)/MariaDB Replica Lag: es7 (paged) [03:50:33] 8115 (RESOLVED) es1035 (paged)/MariaDB Replica Lag: es7 (paged) [03:50:33] 8117 (RESOLVED) es2039 (paged)/MariaDB Replica Lag: es7 (paged) [03:50:34] 8111 (RESOLVED) es2039 (paged)/MariaDB Replica IO: es7 (paged) [03:50:34] 8108 (RESOLVED) es1035 (paged)/MariaDB Replica IO: es7 (paged) [03:50:35] 8109 (RESOLVED) es1048 (paged)/MariaDB Replica IO: es7 (paged) [03:50:35] 8110 (RESOLVED) es1040 (paged)/MariaDB Replica IO: es7 (paged) [03:50:36] 8113 (RESOLVED) es1039 (paged)/mysqld processes (paged) [03:50:36] 8107 (RESOLVED) Host es1039 (paged) [04:27:29] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie [04:45:40] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage [04:49:28] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage [04:54:13] (03PS3) 10Giuseppe Lavagetto: hiddenparma: add default ratelimits file [puppet] - 10https://gerrit.wikimedia.org/r/1306501 (https://phabricator.wikimedia.org/T422249) [04:54:14] (03PS3) 10Giuseppe Lavagetto: cache::varnish: add rate-limit file generated from hiddenparma [puppet] - 10https://gerrit.wikimedia.org/r/1306502 (https://phabricator.wikimedia.org/T422249) [04:54:14] (03PS3) 10Giuseppe Lavagetto: cache::varnish: switch known client rate limits to hp-generated data [puppet] - 10https://gerrit.wikimedia.org/r/1306503 (https://phabricator.wikimedia.org/T422249) [04:56:48] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie [05:00:51] (03CR) 10Giuseppe Lavagetto: [V:03+2 C:03+2] Adding post-deploy step [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1265349 (owner: 10Giuseppe Lavagetto) [05:03:22] (03PS4) 10Giuseppe Lavagetto: hiddenparma: add default ratelimits file [puppet] - 10https://gerrit.wikimedia.org/r/1306501 (https://phabricator.wikimedia.org/T422249) [05:03:22] (03PS4) 10Giuseppe Lavagetto: cache::varnish: add rate-limit file generated from hiddenparma [puppet] - 10https://gerrit.wikimedia.org/r/1306502 (https://phabricator.wikimedia.org/T422249) [05:03:22] (03PS4) 10Giuseppe Lavagetto: cache::varnish: switch known client rate limits to hp-generated data [puppet] - 10https://gerrit.wikimedia.org/r/1306503 (https://phabricator.wikimedia.org/T422249) [05:05:19] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8840/co" [puppet] - 10https://gerrit.wikimedia.org/r/1306501 (https://phabricator.wikimedia.org/T422249) (owner: 10Giuseppe Lavagetto) [05:09:23] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie [05:15:00] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage [05:16:01] !log ryankemper@cumin2002 START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie [05:19:46] (03CR) 10Jelto: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1305985 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [05:20:02] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage [05:32:22] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074368 (10Marostegui) Thank you all, I will follow up from here. [05:34:03] (03CR) 10Jelto: [C:03+1] "lgtm" [puppet] - 10https://gerrit.wikimedia.org/r/1305985 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [05:35:06] (03PS1) 10Medelius: EditCheck: fix pre-save focusedAction error [extensions/VisualEditor] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306842 (https://phabricator.wikimedia.org/T430741) [05:35:34] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [extensions/VisualEditor] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306842 (https://phabricator.wikimedia.org/T430741) (owner: 10Medelius) [05:36:06] !log ryankemper@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage [05:36:24] (03CR) 10Jelto: [C:03+2] Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1305985 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [05:36:39] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074377 (10Marostegui) >>! In T430764#12074209, @CDanis wrote: > 00:42:36 okay, the user impact should be gone now > 00:43:11 it's removed from writes > > Followups: [[ https://wikitech.wikimedi... [05:38:14] (03CR) 10Marostegui: [C:03+2] eqiad.yaml: Add clouddb1027 [puppet] - 10https://gerrit.wikimedia.org/r/1306677 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [05:40:17] !log marostegui@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7 [05:40:20] !log marostegui@cumin1003 conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2 [05:40:47] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage [05:40:49] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie [05:41:07] !log marostegui@cumin1003 conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet [05:43:00] 06SRE, 06DBA: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074381 (10Marostegui) Nothing on HW logs that I can see, last entry from 2025. [05:43:26] (03PS1) 10Marostegui: es1039: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1306847 (https://phabricator.wikimedia.org/T430764) [05:44:41] (03CR) 10Marostegui: [C:03+2] es1039: Disable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1306847 (https://phabricator.wikimedia.org/T430764) (owner: 10Marostegui) [05:45:31] RECOVERY - MariaDB read only es7 on es1039 is OK: Version 10.11.16-MariaDB-log, Uptime 24s, read_only: True, event_scheduler: True, 10.34 QPS, connection latency: 1.143868s, query latency: 0.126464s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Master_comes_back_in_read_only [05:45:41] RECOVERY - MariaDB Events es7 on es1039 is OK: OK - All 4 events in ops database are ENABLED https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [05:45:41] RECOVERY - mysqld processes #page on es1039 is OK: PROCS OK: 1 process with command name mysqld https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting [05:45:43] RECOVERY - MariaDB Event Scheduler es7 on es1039 is OK: Version 10.11.16-MariaDB-log, Uptime 35s, read_only: True, event_scheduler: True, 25.52 QPS, connection latency: 0.011944s, query latency: 0.000557s https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting%23Event_Scheduler [05:45:51] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues [05:49:02] (03CR) 10Jelto: [V:03+1 C:03+2] "puppet runs on the affected hosts were noops" [puppet] - 10https://gerrit.wikimedia.org/r/1305985 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [05:58:49] 06SRE, 06DBA, 13Patch-For-Review: es7 primary (es1039.eqiad.wmnet) crashed - https://phabricator.wikimedia.org/T430764#12074418 (10Marostegui) As soon as I put some load on es1039 it crashed again unfortunately no HW logs produced again. @VRiley-WMF @Jclark-ctr can you check from your side if you are able to... [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T0600) [06:01:19] !log ryankemper@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie [06:10:37] (03CR) 10Giuseppe Lavagetto: [V:03+1 C:03+2] hiddenparma: add default ratelimits file [puppet] - 10https://gerrit.wikimedia.org/r/1306501 (https://phabricator.wikimedia.org/T422249) (owner: 10Giuseppe Lavagetto) [06:19:20] (03PS5) 10Giuseppe Lavagetto: cache::varnish: add rate-limit file generated from hiddenparma [puppet] - 10https://gerrit.wikimedia.org/r/1306502 (https://phabricator.wikimedia.org/T422249) [06:19:20] (03PS5) 10Giuseppe Lavagetto: cache::varnish: switch known client rate limits to hp-generated data [puppet] - 10https://gerrit.wikimedia.org/r/1306503 (https://phabricator.wikimedia.org/T422249) [06:19:20] (03PS1) 10Giuseppe Lavagetto: hiddenparma: brown paper bag fix [puppet] - 10https://gerrit.wikimedia.org/r/1306848 [06:19:52] (03PS2) 10Giuseppe Lavagetto: hiddenparma: brown paper bag fix [puppet] - 10https://gerrit.wikimedia.org/r/1306848 [06:20:25] (03CR) 10Giuseppe Lavagetto: [V:03+2 C:03+2] hiddenparma: brown paper bag fix [puppet] - 10https://gerrit.wikimedia.org/r/1306848 (owner: 10Giuseppe Lavagetto) [06:22:00] (03PS1) 10Muehlenhoff: Remove access for niharika29 [puppet] - 10https://gerrit.wikimedia.org/r/1306849 [06:27:23] (03PS1) 10Abijeet Patro: ULS rewrite: change description key in EmptySearchEntrypoint [extensions/UniversalLanguageSelector] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306850 (https://phabricator.wikimedia.org/T429882) [06:27:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:28:43] (03CR) 10Muehlenhoff: [C:03+2] Remove access for niharika29 [puppet] - 10https://gerrit.wikimedia.org/r/1306849 (owner: 10Muehlenhoff) [06:30:30] !log oblivian@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" [06:30:32] !log oblivian@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 [06:30:52] !log oblivian@cumin1003 END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 [06:30:53] !log oblivian@cumin1003 END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" [06:30:54] (03CR) 10Elukey: redfish: add find_accounts (031 comment) [software/spicerack] - 10https://gerrit.wikimedia.org/r/1303559 (https://phabricator.wikimedia.org/T426180) (owner: 10JHathaway) [06:31:03] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [extensions/UniversalLanguageSelector] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306850 (https://phabricator.wikimedia.org/T429882) (owner: 10Abijeet Patro) [06:31:51] !log jmm@cumin2003 DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts [06:33:57] (03PS1) 10Giuseppe Lavagetto: Use tabs, not spaces [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1306851 [06:34:06] (03CR) 10Giuseppe Lavagetto: [V:03+2 C:03+2] Use tabs, not spaces [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1306851 (owner: 10Giuseppe Lavagetto) [06:34:27] !log oblivian@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" [06:34:28] !log oblivian@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 [06:35:17] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003 [06:35:19] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003" [06:36:47] (03CR) 10Elukey: [C:03+1] rack depool: use build in reason fuction [cookbooks] - 10https://gerrit.wikimedia.org/r/1306637 (owner: 10Ayounsi) [06:37:39] (03CR) 10Elukey: [C:03+1] depool-rack: run the k8s cookbook with relevant alias [cookbooks] - 10https://gerrit.wikimedia.org/r/1306551 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [06:41:32] (03PS1) 10Giuseppe Lavagetto: template fix [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1306854 [06:41:43] (03CR) 10Giuseppe Lavagetto: [V:03+2 C:03+2] template fix [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1306854 (owner: 10Giuseppe Lavagetto) [06:41:55] !log oblivian@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" [06:41:56] !log oblivian@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 [06:42:41] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 [06:42:42] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" [06:43:02] !log oblivian@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" [06:43:03] !log oblivian@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 [06:43:52] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003 [06:43:54] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003" [06:44:42] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqsin:et-0/0/0 (Transport: Hurricane Electric (dc4841.sin1) {#}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [06:45:55] (03CR) 10Ayounsi: [C:03+2] depool-rack: run the k8s cookbook with relevant alias [cookbooks] - 10https://gerrit.wikimedia.org/r/1306551 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [06:49:12] (03Merged) 10jenkins-bot: depool-rack: run the k8s cookbook with relevant alias [cookbooks] - 10https://gerrit.wikimedia.org/r/1306551 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [06:49:49] (03PS1) 10Muehlenhoff: Remove LDAP access for wangombe [puppet] - 10https://gerrit.wikimedia.org/r/1306855 [06:50:43] PROBLEM - OSPF status on cr2-eqiad is CRITICAL: OSPFv2: 6/7 UP : OSPFv3: 6/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:50:45] PROBLEM - OSPF status on cr1-drmrs is CRITICAL: OSPFv2: 3/4 UP : OSPFv3: 3/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [06:51:10] FIRING: [2x] BFDdown: BFD session down between cr1-drmrs and 185.15.58.138 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-drmrs:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:51:39] FIRING: CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.138) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=drmrs&var-device=cr1-drmrs:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [06:51:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [06:53:10] (03CR) 10Muehlenhoff: [C:03+2] Remove LDAP access for wangombe [puppet] - 10https://gerrit.wikimedia.org/r/1306855 (owner: 10Muehlenhoff) [06:55:45] !log upgrade all trixie hosts to pywmflib 3.0 - T430552 [06:55:47] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [06:55:48] T430552: Deploy wmflib 3.0.0 to production - https://phabricator.wikimedia.org/T430552 [06:56:10] FIRING: [4x] BFDdown: BFD session down between cr1-drmrs and 185.15.58.138 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [06:56:39] FIRING: [4x] CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.138) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [07:00:05] Amir1, urbanecm, and awight: #bothumor Q:Why did functions stop calling each other? A:They had arguments. Rise for UTC morning backport window . (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T0700). [07:00:05] chlod, Msz2001, revi, WMDE-Fisch, and abijeet: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:08] o/ [07:00:09] d'oh [07:00:12] o/ [07:00:32] I'm a deployer, who needs help with deploying? [07:00:39] `/me` [07:00:49] \o [07:00:49] * chlod needs help as well [07:01:12] I could selfe serve but don't mind the help [07:01:12] Okay, I'll start with all three config patches, then [07:02:08] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306304 (https://phabricator.wikimedia.org/T430409) (owner: 10Chlod Alejandro) [07:02:09] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306221 (https://phabricator.wikimedia.org/T430512) (owner: 10Mszwarc) [07:02:09] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306649 (https://phabricator.wikimedia.org/T430641) (owner: 10Revi) [07:03:18] (03Merged) 10jenkins-bot: frwiki: change to Wikipedia 25 logo [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306304 (https://phabricator.wikimedia.org/T430409) (owner: 10Chlod Alejandro) [07:03:22] (03Merged) 10jenkins-bot: Temporarily change plwiki tagline for 1.7M articles [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306221 (https://phabricator.wikimedia.org/T430512) (owner: 10Mszwarc) [07:03:25] (03Merged) 10jenkins-bot: CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306649 (https://phabricator.wikimedia.org/T430641) (owner: 10Revi) [07:04:25] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1306304|frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221|Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649|CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] [07:04:32] T430409: Requesting temporary logo change for fr.wikipedia.org - https://phabricator.wikimedia.org/T430409 [07:04:32] T430512: Temporarily change plwiki tagline to reflect 1.7M articles - https://phabricator.wikimedia.org/T430512 [07:04:33] T430641: Add ombuds to highlimits-user of wgWMCGlobalGroupToRateLimitClass - https://phabricator.wikimedia.org/T430641 [07:06:43] !log mszwarc@deploy1003 mszwarc, chlod, revi: Backport for [[gerrit:1306304|frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221|Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649|CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:06:53] checking [07:07:20] revi: is there anything to check for yours? I guess no? [07:07:24] all good for me :) [07:09:18] !log mszwarc@deploy1003 mszwarc, chlod, revi: Continuing with deployment [07:11:01] WMDE-Fisch: I can do yours as well, once these configs are done [07:11:11] sorry, missed the ping [07:11:22] probably nothing I can do because I can't really see the CDN logs [07:11:48] Msz2001: Sure thanks, nothing to test there so you could just go ahead with it [07:12:15] revi: That's what I assumed (I don't know either where these logs are). Anyway, that shouldn't break anything either [07:12:24] WMDE-Fisch: ack [07:12:30] abijeet: You here? I wonder a bit if it makes sense to backport that patch. [07:12:45] (03CR) 10Mszwarc: [C:03+2] "Ahead of deployment" [extensions/Cite] (wmf/1.47.0-wmf.8) - 10https://gerrit.wikimedia.org/r/1306710 (https://phabricator.wikimedia.org/T415904) (owner: 10WMDE-Fisch) [07:12:54] (03CR) 10Mszwarc: [C:03+2] "Ahead o deployment" [extensions/Cite] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306711 (https://phabricator.wikimedia.org/T415904) (owner: 10WMDE-Fisch) [07:13:31] Yeah, and it's rather hard to verify the change, since I am already enjoying the updated configuration because of `global-rollback` I already have (in addition to `ombuds`), and it is rather hard to find people with `ombuds` at all :P [07:13:38] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1306304|frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221|Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649|CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s) [07:13:40] so yeah, let's see if I am to be blamed later ;P [07:13:45] T430409: Requesting temporary logo change for fr.wikipedia.org - https://phabricator.wikimedia.org/T430409 [07:13:45] T430512: Temporarily change plwiki tagline to reflect 1.7M articles - https://phabricator.wikimedia.org/T430512 [07:13:46] T430641: Add ombuds to highlimits-user of wgWMCGlobalGroupToRateLimitClass - https://phabricator.wikimedia.org/T430641 [07:14:32] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/Cite] (wmf/1.47.0-wmf.8) - 10https://gerrit.wikimedia.org/r/1306710 (https://phabricator.wikimedia.org/T415904) (owner: 10WMDE-Fisch) [07:14:32] thanks for the deploy! :D [07:14:33] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/Cite] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306711 (https://phabricator.wikimedia.org/T415904) (owner: 10WMDE-Fisch) [07:15:24] could I squeeze in 1306842 if the rest of these go quickly? I have it scheduled for the next window, but... I'm still awake now [07:15:34] I can deploy [07:15:39] WMDE-Fisch, I'm here. [07:16:13] I'm not sure if this should be backported like that. The renamed key will only be renamed for en. Other languages would need to wait for an update from translate wiki so the backport would lead to missing messages in all other translations [07:17:20] Good point. [07:17:25] cmede: From my POV it's okay, abijeet is still in the queue (but let's see what's the outcome of the above discussion) and Fisch's patches are in deployment right now [07:17:48] ok :) [07:19:05] WMDE-Fisch, the situation is worse without it. I say we go ahead with the backport. See: https://phabricator.wikimedia.org/F91044792 [07:19:05] !log aqu@deploy1003 Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] [07:19:31] What about those not using english as their language? [07:19:39] Yeah it's unforunate, you should have made a patch where you "just" use the other key first [07:19:41] they will continue to see the same key [07:19:42] abijeet: [07:19:48] for other languages [07:20:23] (03Merged) 10jenkins-bot: Fix async loading in footnote click interaction experiment [extensions/Cite] (wmf/1.47.0-wmf.8) - 10https://gerrit.wikimedia.org/r/1306710 (https://phabricator.wikimedia.org/T415904) (owner: 10WMDE-Fisch) [07:20:25] (03Merged) 10jenkins-bot: Fix async loading in footnote click interaction experiment [extensions/Cite] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306711 (https://phabricator.wikimedia.org/T415904) (owner: 10WMDE-Fisch) [07:20:28] and do the rename later [07:20:56] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1306710|Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711|Fix async loading in footnote click interaction experiment (T415904)]] [07:20:59] T415904: [Epic] Experiment Reader Footnote Click Intent - https://phabricator.wikimedia.org/T415904 [07:21:07] !log aqu@deploy1003 Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s) [07:21:07] anyway I have to concur with WMDE-Fisch here because fixing en.json won't fix any and all other langauges [07:21:40] !log aqu@deploy1003 Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] [07:21:49] all other languages will see the same error message, so either fix all language at once if that's that really important or just… let the TWN catch up? [07:22:11] WMDE-Fisch, yea that would have been the cleaner approach. OK lets skip it for now. [07:23:01] !log mszwarc@deploy1003 wmde-fisch, mszwarc: Backport for [[gerrit:1306710|Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711|Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:23:10] revi: abijeet: The key used in the UI will change with this patch. The the code will look for that "new" key in i18n. For that "new" key it will only find the en version and will fall back to it. [07:23:27] So it will be the "new" message but in all languages but english it will fallback to english [07:23:34] oh right [07:23:42] then yes, it will fallback [07:23:53] but not really ideal because missing l10n [07:24:02] !log mszwarc@deploy1003 wmde-fisch, mszwarc: Continuing with deployment [07:24:07] * revi is at IRL work and skimmed the patch [07:24:16] ;-) [07:25:07] But I guess if the "new" en key is better than what was used before it's still fine :-) [07:25:25] anyway backport is not the right approach, and since they agreed to skip this time, I guess the TWN sync will catch up from `master`? [07:25:39] Thinking about this some more, I think showing the key is worse then showing the English fallback. Lets go ahead, the related keys will be renamed and catch up tomorrow. [07:25:46] catch up next week** [07:25:50] :-D [07:26:12] !log aqu@deploy1003 Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s) [07:26:55] !log aqu@deploy1003 Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] [07:27:16] looking at en.json I sort of wonder if you can rename all the keys to new name in translation json? [07:27:33] revi, we generally leave that to translatewiki.net [07:27:38] true [07:28:15] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1306710|Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711|Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s) [07:28:18] T415904: [Epic] Experiment Reader Footnote Click Intent - https://phabricator.wikimedia.org/T415904 [07:28:41] Msz2001: Thanks! [07:28:47] The WMDE-Fisch's patches are deployed. What is the decision here? Or do we let cmede in the queue before deciding on the abijeet's patch? [07:28:55] !log aqu@deploy1003 Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s) [07:28:56] yw [07:29:05] still not sure if this is the right approach but I am really die-hard against it so /mehg [07:29:08] s/mehg/meh [07:29:16] am not really die-hard [07:29:21] ^^' [07:29:37] alt+tabs at maximum speed, mangling words [07:29:44] (I wasn't analysing this patch that much, was focused on other things, so I don't have much opinion rn) [07:29:50] !log aqu@deploy1003 Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] [07:30:12] !log aqu@deploy1003 Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s) [07:30:21] (03PS1) 10Jelto: gitlab: also enable restricted robots.txt on replicas [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) [07:30:31] Can we apply this questioned patch to PatchDemo and check its effect there? [07:30:51] iirc there's already screenshot somewhere above [07:31:59] (03CR) 10Ayounsi: [C:03+2] rack depool: use build in reason fuction [cookbooks] - 10https://gerrit.wikimedia.org/r/1306637 (owner: 10Ayounsi) [07:32:16] I think pretty clear in what would happen. So I entrust abijeet with the decisions what's worse. The old message or the new message but only in English everywhere. [07:32:25] we know the outcome, the other languages will show the English string as a fallback. That's the same as adding a new feature that hasn't had any translations. [07:32:37] +1 [07:32:41] ^ [07:33:04] There are translations which should get deployed with the train next week: https://translatewiki.net/wiki/Special:MessageGroupStats?group=&messages=MediaWiki%3AExt-uls-empty-state-entrypoint-description&x=D#sortable:3=desc [07:33:25] Showing a message key imo is worse. So I vote that we go ahead with the change. [07:33:34] Okay, so I have nothing against deploying. Should I do it or abijeet would you prefer to do it yourself? [07:33:36] I ain't veto the decision :) [07:33:48] (ofc I don't have the right to veto lol) [07:34:12] revi, thanks for your inputs though. I agree it could have been handled better. [07:34:28] I am just yet another professional "yell at everyone" person /self-rant [07:34:35] I'm not releng, so technically I also can't veto ;) [07:35:05] Msz2001, I don't have deployer rights. Please help deploy if you can. [07:35:07] Msz2001: at least deployers can just refuse to do da patch (for those who can't deploy on their own), which is de facto veto :P [07:35:09] (03Merged) 10jenkins-bot: rack depool: use build in reason fuction [cookbooks] - 10https://gerrit.wikimedia.org/r/1306637 (owner: 10Ayounsi) [07:35:24] abijeet: I'll deploy then [07:35:27] (03CR) 10Jelto: [V:03+1] "PCC SUCCESS (NOOP 1 CORE_DIFF 3): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [07:35:57] revi: well, it's only a veto if you have a single deployer available :D [07:35:57] anyway since my patch is done and I still have an hour and a half for my workplace, should probably walk out now lol [07:36:26] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mszwarc@deploy1003 using scap backport" [extensions/UniversalLanguageSelector] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306850 (https://phabricator.wikimedia.org/T429882) (owner: 10Abijeet Patro) [07:36:45] (03PS2) 10Jelto: gitlab: also enable restricted robots.txt on replicas [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) [07:37:29] (03CR) 10CI reject: [V:04-1] gitlab: also enable restricted robots.txt on replicas [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [07:37:34] (03PS1) 10Arnaudb: backup: restrict gerrit-repo-data fileset to git and LFS [puppet] - 10https://gerrit.wikimedia.org/r/1306863 (https://phabricator.wikimedia.org/T411583) [07:42:37] (03PS3) 10Jelto: gitlab: also enable restricted robots.txt on replicas [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) [07:43:26] (03CR) 10Gkyziridis: [C:03+2] ml-services: Deploy qwen36-27b in eager mode to fix startup timeout [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306699 (https://phabricator.wikimedia.org/T425680) (owner: 10Gkyziridis) [07:44:43] (03CR) 10Jelto: [V:03+1] "PCC SUCCESS (NOOP 1 CORE_DIFF 2 DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compile" [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [07:45:06] (03Merged) 10jenkins-bot: ULS rewrite: change description key in EmptySearchEntrypoint [extensions/UniversalLanguageSelector] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306850 (https://phabricator.wikimedia.org/T429882) (owner: 10Abijeet Patro) [07:45:34] (03Merged) 10jenkins-bot: ml-services: Deploy qwen36-27b in eager mode to fix startup timeout [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306699 (https://phabricator.wikimedia.org/T425680) (owner: 10Gkyziridis) [07:45:34] !log mszwarc@deploy1003 Started scap sync-world: Backport for [[gerrit:1306850|ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] [07:45:37] T429882: CX: missing ULSv2 EMPTY_SEARCH entrypoint in languagesearcher - https://phabricator.wikimedia.org/T429882 [07:47:16] (03CR) 10Filippo Giunchedi: [C:03+1] "Thank you! LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1306690 (https://phabricator.wikimedia.org/T430651) (owner: 10Cathal Mooney) [07:49:02] (03CR) 10Jcrespo: "I won't vote but I am ok from the bacula point of view, but I won't have enough gerrit context to know it is the desired way to move forwa" [puppet] - 10https://gerrit.wikimedia.org/r/1306863 (https://phabricator.wikimedia.org/T411583) (owner: 10Arnaudb) [07:50:42] (03PS2) 10Muehlenhoff: Update redis-misc-canary alias [puppet] - 10https://gerrit.wikimedia.org/r/1305344 [07:51:14] (03CR) 10Muehlenhoff: [C:03+1] "Looks good (also raised within SRE IF and no objections were raised)" [puppet] - 10https://gerrit.wikimedia.org/r/1306161 (https://phabricator.wikimedia.org/T430479) (owner: 10Hashar) [07:51:52] !log gkyziridis@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . [07:52:29] 06SRE, 10SRE-tools, 06Infrastructure-Foundations, 13Patch-For-Review: Fix Pypi twine setup for pywmflib - https://phabricator.wikimedia.org/T430620#12074632 (10elukey) 05Open→03Resolved a:03elukey Ok this was a pebcak due to me not knowing how `bdist_wheel` works. I thought I needed to add the `w... [07:52:41] (03CR) 10Muehlenhoff: [C:03+2] Enable profile::docker::builder::prune_images on build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1306670 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [07:53:50] 06SRE, 10SRE-tools, 06Infrastructure-Foundations: Deploy wmflib 3.0.0 to production - https://phabricator.wikimedia.org/T430552#12074644 (10elukey) 3.0.0 deployed fleetwide, and 3.1.0 is ready to go. @MoritzMuehlenhoff we can do the 3.1.0 rollout next week, hopefully this will give enough time to people to f... [07:55:24] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [07:55:49] cmede: I think we might be unable to squeeze your patch, the current deployment is still building images (as it's i18n-related change) [07:56:05] (03CR) 10Elukey: [C:03+1] "Yes I'd say that config.yaml should be the "canonical" version about what's deployed on a cluster, but we'll surely forget to rename in th" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306327 (https://phabricator.wikimedia.org/T427401) (owner: 10JMeybohm) [07:56:18] Msz2001 yeah, i was thinking that... [07:56:20] it's all good [08:00:05] andre and brennen: Time to do the MediaWiki train - Utc-0+Utc-7 Version deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T0800). [08:00:15] * andre waiting for the backport to finish [08:02:36] (03CR) 10Ayounsi: LVS: add public vlan IPs/subnets for LVS still connected to L2 vlans (036 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1306690 (https://phabricator.wikimedia.org/T430651) (owner: 10Cathal Mooney) [08:03:36] !log mszwarc@deploy1003 mszwarc, abi: Backport for [[gerrit:1306850|ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [08:03:39] T429882: CX: missing ULSv2 EMPTY_SEARCH entrypoint in languagesearcher - https://phabricator.wikimedia.org/T429882 [08:03:51] abijeet: Is there anything to verify for your patch? [08:04:30] Msz2001, doing a quick check [08:06:22] (03PS3) 10Cathal Mooney: LVS: add public vlan IPs/subnets for LVS still connected to L2 vlans [puppet] - 10https://gerrit.wikimedia.org/r/1306690 (https://phabricator.wikimedia.org/T430651) [08:07:08] (03PS4) 10Jelto: gitlab: also enable restricted robots.txt on replicas [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) [08:07:19] (03CR) 10Cathal Mooney: LVS: add public vlan IPs/subnets for LVS still connected to L2 vlans (036 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1306690 (https://phabricator.wikimedia.org/T430651) (owner: 10Cathal Mooney) [08:08:13] (03PS4) 10Cathal Mooney: LVS: add public vlan IPs/subnets for LVS still connected to L2 vlans [puppet] - 10https://gerrit.wikimedia.org/r/1306690 (https://phabricator.wikimedia.org/T430651) [08:09:05] Msz2001, looks good. [08:09:11] !log mszwarc@deploy1003 mszwarc, abi: Continuing with deployment [08:13:43] Sorry for going into this train window, it's at "sync-prod-k8s" stage now, so should be done in a few minutes [08:14:47] np [08:15:49] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [08:16:35] (03CR) 10Arnaudb: [C:03+1] "lgtm thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [08:18:17] 06SRE, 10SRE-tools, 06Infrastructure-Foundations: Deploy wmflib 3.0.0 to production - https://phabricator.wikimedia.org/T430552#12074705 (10MoritzMuehlenhoff) >>! In T430552#12074635, @elukey wrote: > 3.0.0 deployed fleetwide, and 3.1.0 is ready to go. @MoritzMuehlenhoff we can do the 3.1.0 rollout next week... [08:21:45] !log mszwarc@deploy1003 Finished scap sync-world: Backport for [[gerrit:1306850|ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s) [08:21:46] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [08:21:48] andre: it's done, over to you [08:21:48] T429882: CX: missing ULSv2 EMPTY_SEARCH entrypoint in languagesearcher - https://phabricator.wikimedia.org/T429882 [08:22:04] Wow, this scap tool 45 mins [08:22:07] took* [08:22:33] ack [08:24:22] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.9 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306870 (https://phabricator.wikimedia.org/T423918) [08:24:25] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by aklapper@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306870 (https://phabricator.wikimedia.org/T423918) (owner: 10TrainBranchBot) [08:24:53] (03CR) 10Ayounsi: [C:03+1] LVS: add public vlan IPs/subnets for LVS still connected to L2 vlans [puppet] - 10https://gerrit.wikimedia.org/r/1306690 (https://phabricator.wikimedia.org/T430651) (owner: 10Cathal Mooney) [08:25:02] (03CR) 10STran: "Do we also need to remove the `ipoid_rw` and `ipoid_ro` users here too? Or will that be done as part of the database drop (I assume we're " [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [08:26:15] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.9 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306870 (https://phabricator.wikimedia.org/T423918) (owner: 10TrainBranchBot) [08:36:05] (03CR) 10Tiziano Fogli: docker_registry: migrate nrpe checks to alertmanager (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306351 (https://phabricator.wikimedia.org/T384321) (owner: 10Hnowlan) [08:36:32] !log aklapper@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs T423918 [08:36:35] T423918: 1.47.0-wmf.9 deployment blockers - https://phabricator.wikimedia.org/T423918 [08:38:51] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [08:38:58] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [08:39:15] FIRING: MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?panelId=18&fullscreen&orgId=1&var-datasource=codfw%20prometheus/ops - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [08:44:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [08:46:16] (03PS1) 10Marostegui: clouddb1027: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1306871 (https://phabricator.wikimedia.org/T409557) [08:46:17] Going to roll back the train [08:46:40] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.9 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306872 (https://phabricator.wikimedia.org/T423918) [08:46:43] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by aklapper@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306872 (https://phabricator.wikimedia.org/T423918) (owner: 10TrainBranchBot) [08:46:55] (03PS1) 10Jgiannelos: Parsoid read views: Bump enwiki traffic to 75% [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306873 [08:47:45] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.9 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306872 (https://phabricator.wikimedia.org/T423918) (owner: 10TrainBranchBot) [08:48:44] (03CR) 10Marostegui: [C:03+2] clouddb1027: Enable notifications [puppet] - 10https://gerrit.wikimedia.org/r/1306871 (https://phabricator.wikimedia.org/T409557) (owner: 10Marostegui) [08:49:06] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306873 (owner: 10Jgiannelos) [08:49:10] (03PS1) 10Elukey: spicerack: add management/config.yaml structure [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) [08:53:12] (03PS2) 10Elukey: spicerack: add management/config.yaml structure [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) [08:54:07] !log aklapper@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs T423918 [08:54:10] T423918: 1.47.0-wmf.9 deployment blockers - https://phabricator.wikimedia.org/T423918 [08:55:15] (03PS3) 10Elukey: spicerack: add management/config.yaml structure [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) [08:56:43] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) (owner: 10Elukey) [08:57:42] (03CR) 10Dreamy Jazz: "I was intending to leave the database until the last step (just in case something is still trying to access it)" [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [08:58:38] (03CR) 10Fabfur: [C:03+1] "as I was saying, I'd prefer to have a single map file, eventually with the cluster indicated as value in the map, but this is good to me a" [puppet] - 10https://gerrit.wikimedia.org/r/1306545 (https://phabricator.wikimedia.org/T402512) (owner: 10Elukey) [08:59:15] RESOLVED: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [09:02:04] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [09:02:12] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [09:05:07] (03PS3) 10Hnowlan: docker_registry: migrate nrpe checks to alertmanager [puppet] - 10https://gerrit.wikimedia.org/r/1306351 (https://phabricator.wikimedia.org/T384321) [09:10:50] (03PS1) 10Elukey: role::cluster::management: add fake mgmt password [labs/private] - 10https://gerrit.wikimedia.org/r/1306875 (https://phabricator.wikimedia.org/T429699) [09:11:12] (03CR) 10Hnowlan: docker_registry: migrate nrpe checks to alertmanager (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306351 (https://phabricator.wikimedia.org/T384321) (owner: 10Hnowlan) [09:11:22] (03CR) 10Elukey: [V:03+2 C:03+2] role::cluster::management: add fake mgmt password [labs/private] - 10https://gerrit.wikimedia.org/r/1306875 (https://phabricator.wikimedia.org/T429699) (owner: 10Elukey) [09:11:53] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) (owner: 10Elukey) [09:14:12] (03CR) 10Elukey: "Added Filippo for the Cloud deployment :)" [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) (owner: 10Elukey) [09:14:36] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [09:14:46] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [09:21:02] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [09:21:10] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [09:27:37] I'll have a private code change to deploy. Would anyone mind if I do it now? [09:28:14] (03CR) 10Elukey: Add sre.hosts.bmc-user-mgmt.py (037 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [09:29:23] Msz2001: yes, I think you could use the current train window - train is rolled back and currently blocked, and I don't see a patch to fix the blocker at T430778 yet [09:29:24] T430778: Wikimedia\Assert\PreconditionException: Precondition failed: This Title instance does not represent a proper page, but merely a link target. - https://phabricator.wikimedia.org/T430778 [09:29:46] Thanks, I'll proceed with deployment in a minute, then [09:33:19] Deploying [09:35:30] (03CR) 10Giuseppe Lavagetto: [C:03+2] cache::varnish: add rate-limit file generated from hiddenparma [puppet] - 10https://gerrit.wikimedia.org/r/1306502 (https://phabricator.wikimedia.org/T422249) (owner: 10Giuseppe Lavagetto) [09:39:09] !log mszwarc@deploy1003 Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s) [09:39:24] Done [09:39:36] (03PS1) 10Blake: services: Add a new mw-pretrain k8s service. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306878 [09:39:56] (03CR) 10Klausman: [V:03+1 C:03+2] "PCC SUCCESS (NOOP 3): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8846/console" [puppet] - 10https://gerrit.wikimedia.org/r/1294223 (https://phabricator.wikimedia.org/T420438) (owner: 10Elukey) [09:41:25] (03CR) 10Bartosz Wójtowicz: [C:03+2] ml-services: Add qwen3-14b deployment to llm namespace. (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306286 (https://phabricator.wikimedia.org/T426749) (owner: 10Bartosz Wójtowicz) [09:42:13] (03PS2) 10Blake: services: Add a new mw-pretrain k8s service. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306878 [09:43:29] (03Merged) 10jenkins-bot: ml-services: Add qwen3-14b deployment to llm namespace. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306286 (https://phabricator.wikimedia.org/T426749) (owner: 10Bartosz Wójtowicz) [09:44:25] (03CR) 10Klausman: [V:03+1 C:03+2] "PCC SUCCESS (NOOP 3): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8847/console" [puppet] - 10https://gerrit.wikimedia.org/r/1294223 (https://phabricator.wikimedia.org/T420438) (owner: 10Elukey) [09:46:25] (03PS1) 10Giuseppe Lavagetto: known-client-rate-limits: fix confd template [puppet] - 10https://gerrit.wikimedia.org/r/1306881 [09:46:37] (03CR) 10Giuseppe Lavagetto: [V:03+2 C:03+2] known-client-rate-limits: fix confd template [puppet] - 10https://gerrit.wikimedia.org/r/1306881 (owner: 10Giuseppe Lavagetto) [09:47:17] (03CR) 10Klausman: [V:03+1 C:03+2] "PCC SUCCESS (NOOP 3): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8848/console" [puppet] - 10https://gerrit.wikimedia.org/r/1294223 (https://phabricator.wikimedia.org/T420438) (owner: 10Elukey) [09:47:56] (03CR) 10Marostegui: "Can we resume this work?" [cookbooks] - 10https://gerrit.wikimedia.org/r/1277076 (https://phabricator.wikimedia.org/T419874) (owner: 10Federico Ceratto) [09:49:42] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-eqsin:et-0/0/0 (Transport: Hurricane Electric (dc4841.sin1) {#dc4841.dal5.sin}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:49:56] (03CR) 10Klausman: [V:03+1 C:03+2] "PCC SUCCESS (NOOP 3): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8849/console" [puppet] - 10https://gerrit.wikimedia.org/r/1294223 (https://phabricator.wikimedia.org/T420438) (owner: 10Elukey) [09:50:29] (03PS1) 10Aqu: Give commonswiki-monthly sqoop its own log [puppet] - 10https://gerrit.wikimedia.org/r/1306883 (https://phabricator.wikimedia.org/T427532) [09:51:23] !log bwojtowicz@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [09:52:25] (03CR) 10Klausman: [V:03+1 C:03+2] "PCC SUCCESS (NOOP 3): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/8850/console" [puppet] - 10https://gerrit.wikimedia.org/r/1294223 (https://phabricator.wikimedia.org/T420438) (owner: 10Elukey) [09:54:28] (03CR) 10Mvolz: [C:03+2] citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306355 (owner: 10PipelineBot) [09:55:19] (03PS1) 10Giuseppe Lavagetto: Template fix [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1306885 [09:55:29] (03CR) 10Giuseppe Lavagetto: [V:03+2 C:03+2] Template fix [software/hiddenparma/deploy] - 10https://gerrit.wikimedia.org/r/1306885 (owner: 10Giuseppe Lavagetto) [09:55:49] !log oblivian@cumin1003 START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" [09:55:51] !log oblivian@cumin1003 START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 [09:56:41] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003 [09:56:42] !log oblivian@cumin1003 END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003" [09:56:44] (03Merged) 10jenkins-bot: citoid: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306355 (owner: 10PipelineBot) [09:56:48] (03PS1) 10Dr0ptp4kt: Separate Commons log file for globalimagelinks sqoop [puppet] - 10https://gerrit.wikimedia.org/r/1306886 (https://phabricator.wikimedia.org/T427532) [09:58:43] (03Abandoned) 10Dr0ptp4kt: Separate Commons log file for globalimagelinks sqoop [puppet] - 10https://gerrit.wikimedia.org/r/1306886 (https://phabricator.wikimedia.org/T427532) (owner: 10Dr0ptp4kt) [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1000) [10:01:29] (03CR) 10Joal: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1306883 (https://phabricator.wikimedia.org/T427532) (owner: 10Aqu) [10:01:35] (03CR) 10Mvolz: [C:03+2] Update translators for zotero [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306305 (https://phabricator.wikimedia.org/T428915) (owner: 10Mvolz) [10:03:43] (03Merged) 10jenkins-bot: Update translators for zotero [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306305 (https://phabricator.wikimedia.org/T428915) (owner: 10Mvolz) [10:07:57] (03PS2) 10Gerrit maintenance bot: mariadb: Promote db2179 to s4 master [puppet] - 10https://gerrit.wikimedia.org/r/1305617 (https://phabricator.wikimedia.org/T430127) [10:15:12] !log cwilliams@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 T430127 [10:15:15] T430127: Switchover s4 master (db2240 -> db2179) - https://phabricator.wikimedia.org/T430127 [10:15:32] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Set db2179 with weight 0 T430127', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json [10:16:22] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) (owner: 10Elukey) [10:20:48] !log atsuko@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie [10:21:17] (03CR) 10CWilliams: [C:03+2] mariadb: Promote db2179 to s4 master [puppet] - 10https://gerrit.wikimedia.org/r/1305617 (https://phabricator.wikimedia.org/T430127) (owner: 10Gerrit maintenance bot) [10:23:09] !log atsuko@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie [10:23:13] !log Starting s4 codfw failover from db2240 to db2179 - T430127 [10:23:16] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:23:16] T430127: Switchover s4 master (db2240 -> db2179) - https://phabricator.wikimedia.org/T430127 [10:23:57] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Promote db2179 to s4 primary T430127', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json [10:26:02] !log atsuko@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie [10:26:36] !log installing nginx security updates [10:26:37] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:26:59] !log cwilliams@cumin1003 dbctl commit (dc=all): 'Depool db2240 T430127', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json [10:27:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:39:48] !log atsuko@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage [10:40:59] !log cwilliams@cumin1003 START - Cookbook sre.mysql.major-upgrade [10:40:59] !log cwilliams@cumin1003 dbmaint on s4@codfw T429893 [10:41:02] T429893: Migrate s4 section to Debian Trixie - https://phabricator.wikimedia.org/T429893 [10:41:08] !log cwilliams@cumin1003 START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet [10:41:20] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet [10:42:44] !log atsuko@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage [10:44:12] !log atsuko@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage [10:44:31] !log cwilliams@cumin1003 START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie [10:44:58] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage [10:45:48] RECOVERY - OSPF status on cr2-eqiad is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [10:45:50] RECOVERY - OSPF status on cr1-drmrs is OK: OSPFv2: 4/4 UP : OSPFv3: 4/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [10:48:05] (03PS1) 10Bartosz Wójtowicz: admin_ng: Fix KServe view RBAC on install_kserve_resources [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306892 [10:49:10] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage [10:53:01] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage [10:57:09] RESOLVED: [4x] CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.138) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [10:57:40] RESOLVED: [4x] BFDdown: BFD session down between cr1-drmrs and 185.15.58.138 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [10:58:10] (03PS1) 10Marostegui: mariadb: Add x4 [puppet] - 10https://gerrit.wikimedia.org/r/1306894 (https://phabricator.wikimedia.org/T404715) [10:58:28] (03CR) 10Marostegui: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306894 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [10:59:20] (03CR) 10Muehlenhoff: [C:03+2] Enable the weekly base build on build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1306695 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [11:00:05] mvolz: That opportune time for a Services – Citoid / Zotero deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1100). [11:00:24] !log cwilliams@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage [11:00:54] (03CR) 10Marostegui: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306894 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [11:04:31] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage [11:04:41] (03PS2) 10Marostegui: mariadb: Add x4 [puppet] - 10https://gerrit.wikimedia.org/r/1306894 (https://phabricator.wikimedia.org/T404715) [11:04:56] (03CR) 10Marostegui: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306894 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [11:06:08] (03CR) 10STran: [C:03+1] deployment_server: absent ipoid kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1306779 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [11:06:51] (03CR) 10STran: [C:03+1] deployment_server: remove ipoid users [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [11:07:17] (03PS4) 10Krinkle: varnish: Add edge fixup for corrupt upload.wm.o urls from mobileapps [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) [11:07:19] (03CR) 10Marostegui: "https://puppet-compiler.wmflabs.org/output/1306894/7113/" [puppet] - 10https://gerrit.wikimedia.org/r/1306894 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [11:08:21] (03CR) 10STran: [C:03+1] Remove ipoid chart and service definitions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306784 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [11:09:32] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie [11:12:03] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [11:13:13] (03PS1) 10Slyngshede: P:cache::varnish::frontend configure varnish for thumbnail [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [11:13:56] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/citoid: apply [11:14:20] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/citoid: apply [11:14:28] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie [11:14:56] (03PS3) 10Clément Goubert: deployment_server: absent ipoid kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1306779 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [11:14:56] (03PS4) 10Clément Goubert: deployment_server: remove ipoid users [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [11:14:56] (03PS1) 10Clément Goubert: service: ipoid to service_setup [puppet] - 10https://gerrit.wikimedia.org/r/1306899 (https://phabricator.wikimedia.org/T416623) [11:14:59] (03PS1) 10Clément Goubert: service: Remove ipoid service [puppet] - 10https://gerrit.wikimedia.org/r/1306900 (https://phabricator.wikimedia.org/T416623) [11:15:30] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie [11:16:21] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/citoid: apply [11:16:54] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/citoid: apply [11:17:32] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" [11:17:44] (03PS1) 10Cathal Mooney: Add INCLUDE statements for new PTR records for magru HE linknets [dns] - 10https://gerrit.wikimedia.org/r/1306901 (https://phabricator.wikimedia.org/T424839) [11:17:59] (03PS1) 10Clément Goubert: wmnet: Remove ipoid CNAME [dns] - 10https://gerrit.wikimedia.org/r/1306902 (https://phabricator.wikimedia.org/T416623) [11:18:55] (03CR) 10CI reject: [V:04-1] wmnet: Remove ipoid CNAME [dns] - 10https://gerrit.wikimedia.org/r/1306902 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [11:20:37] cmooney@cumin1003 netbox (PID 290836) is awaiting input [11:20:54] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" [11:20:54] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:21:34] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie [11:23:52] !log cmooney@cumin1003 START - Cookbook sre.dns.netbox [11:25:17] (03PS2) 10Slyngshede: P:cache::varnish::frontend configure varnish for thumbnail [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [11:27:08] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/citoid: apply [11:27:24] !log cmooney@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" [11:27:29] !log cmooney@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003" [11:27:29] !log cmooney@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:27:44] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/citoid: apply [11:28:44] (03PS2) 10Clément Goubert: service: ipoid to service_setup [puppet] - 10https://gerrit.wikimedia.org/r/1306899 (https://phabricator.wikimedia.org/T416623) [11:28:44] (03PS4) 10Clément Goubert: deployment_server: absent ipoid kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1306779 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [11:28:44] (03PS5) 10Clément Goubert: deployment_server: remove ipoid users [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [11:28:45] (03PS2) 10Clément Goubert: service: Remove ipoid service [puppet] - 10https://gerrit.wikimedia.org/r/1306900 (https://phabricator.wikimedia.org/T416623) [11:28:46] (03PS1) 10Clément Goubert: services_proxy: Remove ipoid listener [puppet] - 10https://gerrit.wikimedia.org/r/1306903 (https://phabricator.wikimedia.org/T416623) [11:28:55] !log mvolz@deploy1003 helmfile [staging] START helmfile.d/services/zotero: apply [11:29:42] FIRING: CoreRouterInterfaceDown: Core router interface down - cr1-magru:et-0/0/0 (Transport: Hurricane Electric (dc4841.sao4) {#changeme_magru_he_cct}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [11:30:15] !log mvolz@deploy1003 helmfile [staging] DONE helmfile.d/services/zotero: apply [11:31:40] !log cwilliams@cumin1003 START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed [11:33:35] (03PS5) 10Krinkle: varnish: Add edge fixup for corrupt upload.wm.o urls from mobileapps [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) [11:35:36] (03CR) 10Aqu: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306883 (https://phabricator.wikimedia.org/T427532) (owner: 10Aqu) [11:36:00] !log mvolz@deploy1003 helmfile [eqiad] START helmfile.d/services/zotero: apply [11:36:24] (03CR) 10Krinkle: "@Tgr If you want to run this locally:" [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) (owner: 10Krinkle) [11:36:29] !log mvolz@deploy1003 helmfile [eqiad] DONE helmfile.d/services/zotero: apply [11:39:34] (03PS3) 10CWilliams: Allow a single replica for sre.mysql.major-upgrade [cookbooks] - 10https://gerrit.wikimedia.org/r/1305682 (https://phabricator.wikimedia.org/T429758) [11:39:37] (03PS6) 10Krinkle: varnish: Add edge fixup for corrupt upload.wm.o urls from mobileapps [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) [11:40:09] !log mvolz@deploy1003 helmfile [codfw] START helmfile.d/services/zotero: apply [11:40:40] !log mvolz@deploy1003 helmfile [codfw] DONE helmfile.d/services/zotero: apply [11:40:48] (03PS7) 10Krinkle: varnish: Add edge fixup for corrupt upload.wm.o urls from mobileapps [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) [11:40:50] (03CR) 10Krinkle: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) (owner: 10Krinkle) [11:47:21] (03CR) 10Dreamy Jazz: [C:03+1] services_proxy: Remove ipoid listener [puppet] - 10https://gerrit.wikimedia.org/r/1306903 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [11:47:59] (03CR) 10Cathal Mooney: [C:03+2] Add INCLUDE statements for new PTR records for magru HE linknets [dns] - 10https://gerrit.wikimedia.org/r/1306901 (https://phabricator.wikimedia.org/T424839) (owner: 10Cathal Mooney) [11:48:04] (03PS1) 10Clément Goubert: trafficserver::backend: Remove X-W-D for /w/rest.php [puppet] - 10https://gerrit.wikimedia.org/r/1306905 (https://phabricator.wikimedia.org/T428909) [11:48:44] !log atsuko@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie [11:49:01] !log cmooney@dns2005 START - running authdns-update [11:49:20] (03CR) 10Zabe: [C:03+2] BETA: Update interwiki cache [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306000 (owner: 10Zabe) [11:50:42] (03Merged) 10jenkins-bot: BETA: Update interwiki cache [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306000 (owner: 10Zabe) [11:50:44] !log cmooney@dns2005 END - running authdns-update [11:52:16] !log atsuko@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie [11:56:02] (03CR) 10Filippo Giunchedi: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) (owner: 10Elukey) [11:56:50] (03CR) 10CWilliams: [C:03+2] Allow a single replica for sre.mysql.major-upgrade [cookbooks] - 10https://gerrit.wikimedia.org/r/1305682 (https://phabricator.wikimedia.org/T429758) (owner: 10CWilliams) [11:58:10] (03PS3) 10Slyngshede: P:cache::varnish::frontend configure varnish for thumbnail [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) [11:59:50] (03Merged) 10jenkins-bot: Allow a single replica for sre.mysql.major-upgrade [cookbooks] - 10https://gerrit.wikimedia.org/r/1305682 (https://phabricator.wikimedia.org/T429758) (owner: 10CWilliams) [12:00:09] !log drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade T426343 [12:00:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:00:12] T426343: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343 [12:01:00] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [12:03:10] (03CR) 10Ladsgroup: "This is great. We can add them later but there are a couple more stuff like https://gerrit.wikimedia.org/r/c/operations/puppet/+/1306395 t" [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [12:03:50] (03CR) 10Filippo Giunchedi: [V:03+2 C:03+2] admin: add monathierse to analytics-privatedata-users [puppet] - 10https://gerrit.wikimedia.org/r/1306495 (https://phabricator.wikimedia.org/T430304) (owner: 10Filippo Giunchedi) [12:05:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:06:25] (03PS1) 10Neriah: PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) [12:09:01] !log atsuko@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage [12:09:32] !log atsuko@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage [12:10:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.17% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:11:38] (03CR) 10CDobbins: "My first-pass read of this is that it looks good, but given all the regexes, I really think this should have testing." [puppet] - 10https://gerrit.wikimedia.org/r/1306896 (https://phabricator.wikimedia.org/T427465) (owner: 10Slyngshede) [12:14:42] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage [12:17:12] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed [12:17:13] !log cwilliams@cumin1003 END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0) [12:19:22] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage [12:28:23] (03PS1) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [12:29:02] (03CR) 10CI reject: [V:04-1] varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [12:30:18] (03PS7) 10Elukey: Add sre.hosts.bmc-user-mgmt.py [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) [12:32:05] (03CR) 10Elukey: Add sre.hosts.bmc-user-mgmt.py (033 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [12:33:22] (03PS2) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [12:38:43] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie [12:40:59] PROBLEM - Postfix SMTP on crm2001 is CRITICAL: CRITICAL - Certificate crm2001.codfw.wmnet expires in 15 day(s) (Fri 17 Jul 2026 12:40:00 PM GMT +0000). https://wikitech.wikimedia.org/wiki/Mail%23Troubleshooting [12:42:07] (03CR) 10Ladsgroup: [C:03+1] mariadb: Add x4 [puppet] - 10https://gerrit.wikimedia.org/r/1306894 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [12:42:34] (03CR) 10Dreamy Jazz: [C:03+1] service: ipoid to service_setup [puppet] - 10https://gerrit.wikimedia.org/r/1306899 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [12:42:36] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie [12:42:51] (03CR) 10Dreamy Jazz: [C:03+1] service: Remove ipoid service [puppet] - 10https://gerrit.wikimedia.org/r/1306900 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [12:43:34] (03CR) 10Marostegui: [C:03+2] mariadb: Add x4 [puppet] - 10https://gerrit.wikimedia.org/r/1306894 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [12:45:24] (03CR) 10Jgiannelos: "This will fail again. The problem in production is that for a bad title we end up in a situation where the title is in NS_SPECIAL (because" [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) (owner: 10Neriah) [12:45:29] (03CR) 10Jgiannelos: [C:04-1] PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) (owner: 10Neriah) [12:48:27] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db2229 to s6 master [puppet] - 10https://gerrit.wikimedia.org/r/1306912 (https://phabricator.wikimedia.org/T430814) [12:50:09] !log fceratto@cumin1003 START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet [12:50:10] !log fceratto@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet [12:51:41] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 T430814 [12:51:42] (03CR) 10Dzahn: [C:03+2] proxy: Allow outbount HTTPS connections to port 25000 [puppet] - 10https://gerrit.wikimedia.org/r/1306161 (https://phabricator.wikimedia.org/T430479) (owner: 10Hashar) [12:51:45] T430814: Switchover s6 master (db2214 -> db2229) - https://phabricator.wikimedia.org/T430814 [12:51:50] !log fceratto@cumin1003 dbctl commit (dc=all): 'Set db2229 with weight 0 T430814', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json [12:51:52] 06SRE, 06Infrastructure-Foundations: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12075814 (10MoritzMuehlenhoff) >>! In T427282#12072623, @ssingh wrote: > I haven't caught up with the ticket yet as I was out but note that `urldownloader1005` failed again (just eqiad), so the above... [12:52:15] (03PS2) 10Anzx: eswikisource: add wikibooks as importsource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306456 (https://phabricator.wikimedia.org/T430537) [12:52:29] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306456 (https://phabricator.wikimedia.org/T430537) (owner: 10Anzx) [12:54:25] 10SRE-SLO: Sloth: enable alerting - https://phabricator.wikimedia.org/T428617#12075827 (10tappof) 05Open→03Resolved a:03tappof [12:54:46] (03PS1) 10Marostegui: db2247: Add note about its future [puppet] - 10https://gerrit.wikimedia.org/r/1306914 (https://phabricator.wikimedia.org/T404715) [12:55:13] 10SRE-SLO, 10Observability-Alerting, 06SRE Observability (FY2025/2026-Q3): sloth deployment - https://phabricator.wikimedia.org/T414579#12075833 (10tappof) [12:55:25] (03CR) 10Dzahn: gitlab: also enable restricted robots.txt on replicas (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [12:56:26] (03CR) 10Marostegui: "this is a noop" [puppet] - 10https://gerrit.wikimedia.org/r/1306914 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [12:56:28] (03CR) 10Marostegui: [C:03+2] db2247: Add note about its future [puppet] - 10https://gerrit.wikimedia.org/r/1306914 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [12:56:44] (03CR) 10Federico Ceratto: [C:03+2] mariadb: Promote db2229 to s6 master [puppet] - 10https://gerrit.wikimedia.org/r/1306912 (https://phabricator.wikimedia.org/T430814) (owner: 10Gerrit maintenance bot) [12:56:47] (03CR) 10Dzahn: "@ssingh@wikimedia.org what would you say about minimum TTL?" [dns] - 10https://gerrit.wikimedia.org/r/1306459 (https://phabricator.wikimedia.org/T425441) (owner: 10Arnaudb) [12:57:04] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install [12:57:12] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12075843 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=0e1e62e2-283a-46e8-8264-f9dc495aa361) set by... [12:58:14] (03PS1) 10Ozge: ml-services: editing-suggestions model bump to 20260701125423 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306915 (https://phabricator.wikimedia.org/T430812) [12:58:42] (03PS1) 10Dreamy Jazz: Move non temporary accounts settings out TA section [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306916 [12:58:49] (03CR) 10Ozge: [V:03+2 C:03+2] ml-services: editing-suggestions model bump to 20260701125423 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306915 (https://phabricator.wikimedia.org/T430812) (owner: 10Ozge) [12:59:24] !log Starting s6 codfw failover from db2214 to db2229 - T430814 [12:59:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:59:27] T430814: Switchover s6 master (db2214 -> db2229) - https://phabricator.wikimedia.org/T430814 [13:00:00] !log fceratto@cumin1003 dbctl commit (dc=all): 'Promote db2229 to s6 primary T430814', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: gettimeofday() says it's time for UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1300) [13:00:05] cmede, nemo-yiannis, and anzx: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:09] FYI if there is time left at the end of this backport window, I may want to backport one of the T430778 train blocker patches. If not, I'll do it later today (as the train would still be blocked). [13:00:10] T430778: Wikimedia\Assert\PreconditionException: Precondition failed: This Title instance does not represent a proper page, but merely a link target. - https://phabricator.wikimedia.org/T430778 [13:00:22] o/ [13:00:29] o/ [13:00:42] i can deploy mine [13:00:53] andre: My initial patch doesn't fix the issue but the follow up should. I don't think we can review this in this window, lets unblock the train after [13:00:58] (03Merged) 10jenkins-bot: ml-services: editing-suggestions model bump to 20260701125423 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306915 (https://phabricator.wikimedia.org/T430812) (owner: 10Ozge) [13:00:58] I can also deploy mine [13:01:06] need someone to deploy mine [13:01:13] (03CR) 10DLynch: "> This version replaces reference chars ("[", "]") with a special char" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306915 (https://phabricator.wikimedia.org/T430812) (owner: 10Ozge) [13:01:30] * TheresNoTime is around to deploy if needed [13:01:35] nemo-yiannis: ah, only now I'm matching IRC names and Phab names, heh :D Yeah let's unblock later today [13:01:45] (03PS3) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [13:01:59] !log installing python3.13 security updates [13:02:00] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:02:00] nemo-yiannis: but feel free to get the first one deployed if you feel like [13:02:06] ok [13:02:09] on it [13:02:25] RESOLVED: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:02:28] o/ sorry, I was distracted. also around for a bit [13:03:02] who's first? [13:03:25] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jgiannelos@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306873 (owner: 10Jgiannelos) [13:03:44] that answers that :D [13:03:50] :D [13:04:14] !log fceratto@cumin1003 dbctl commit (dc=all): 'Depool db2214 T430814', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json [13:04:22] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover [13:04:37] (03CR) 10Tiziano Fogli: [C:03+1] "+1 from me. Up to you whether to wait for feedback from @ltoscano@wikimedia.org as well or just proceed." [puppet] - 10https://gerrit.wikimedia.org/r/1306351 (https://phabricator.wikimedia.org/T384321) (owner: 10Hnowlan) [13:04:49] (03Merged) 10jenkins-bot: Parsoid read views: Bump enwiki traffic to 75% [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306873 (owner: 10Jgiannelos) [13:05:16] !log jgiannelos@deploy1003 Started scap sync-world: Backport for [[gerrit:1306873|Parsoid read views: Bump enwiki traffic to 75%]] [13:05:53] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover [13:06:48] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance [13:06:59] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' . [13:07:30] !log jgiannelos@deploy1003 jgiannelos: Backport for [[gerrit:1306873|Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:08:20] !log ozge@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [13:08:49] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:09:22] !log jgiannelos@deploy1003 jgiannelos: Continuing with deployment [13:11:23] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [13:11:28] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [13:12:04] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.5 point update - https://phabricator.wikimedia.org/T427072#12075913 (10MoritzMuehlenhoff) [13:13:04] !log installing qemu security updates [13:13:04] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:13:45] !log jgiannelos@deploy1003 Finished scap sync-world: Backport for [[gerrit:1306873|Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s) [13:14:05] i am done with my patch [13:14:52] shall i go? [13:15:06] (03PS5) 10Jelto: gitlab: also enable restricted robots.txt on replicas [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) [13:15:18] !log rebooting routing-engine 1 on cr1-eqiad T417873 [13:15:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:15:20] T417873: eqiad: upgrade routers (2026) - https://phabricator.wikimedia.org/T417873 [13:15:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by caro@deploy1003 using scap backport" [extensions/VisualEditor] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306842 (https://phabricator.wikimedia.org/T430741) (owner: 10Medelius) [13:16:15] (03PS1) 10Dreamy Jazz: Remove TA patrol rights from users on fishbowl + private [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306925 (https://phabricator.wikimedia.org/T425048) [13:16:22] (03CR) 10Jelto: gitlab: also enable restricted robots.txt on replicas (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [13:17:34] jouncebot: nowandnext [13:17:34] For the next 0 hour(s) and 42 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1300) [13:17:34] In 0 hour(s) and 42 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1400) [13:18:27] (03CR) 10Jelto: [V:03+1] "PCC SUCCESS (NOOP 1 DIFF 1 CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compile" [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [13:18:33] (03PS4) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [13:21:01] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306916 (owner: 10Dreamy Jazz) [13:21:07] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306925 (https://phabricator.wikimedia.org/T425048) (owner: 10Dreamy Jazz) [13:21:49] (03CR) 10Kamila Součková: [C:03+1] services: Add a new mw-pretrain k8s service. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306878 (owner: 10Blake) [13:24:14] (03CR) 10Dzahn: gitlab: also enable restricted robots.txt on replicas (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [13:26:17] (03Merged) 10jenkins-bot: EditCheck: fix pre-save focusedAction error [extensions/VisualEditor] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306842 (https://phabricator.wikimedia.org/T430741) (owner: 10Medelius) [13:26:45] !log caro@deploy1003 Started scap sync-world: Backport for [[gerrit:1306842|EditCheck: fix pre-save focusedAction error (T430741)]] [13:26:48] T430741: VisualEditor freezes in pre-save if all checks are collapsed - https://phabricator.wikimedia.org/T430741 [13:27:03] !log route-engine failover cr1-eqiad [13:27:03] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:28:50] !log caro@deploy1003 caro: Backport for [[gerrit:1306842|EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:29:40] (03CR) 10Tiziano Fogli: [C:03+1] "Although a probe is already defined in hieradata/common/service.yaml:" [puppet] - 10https://gerrit.wikimedia.org/r/1306351 (https://phabricator.wikimedia.org/T384321) (owner: 10Hnowlan) [13:29:51] PROBLEM - OSPF status on cr2-eqdfw is CRITICAL: OSPFv2: 6/7 UP : OSPFv3: 6/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:30:24] !log caro@deploy1003 caro: Continuing with deployment [13:30:29] (03PS1) 10Muehlenhoff: Revert urldownloader in codfw to 2003 [dns] - 10https://gerrit.wikimedia.org/r/1306926 (https://phabricator.wikimedia.org/T427282) [13:30:55] FIRING: [3x] PyBalBGPUnstable: PyBal BGP sessions on instance lvs1018 with peer 208.80.154.196 are failing #page - https://wikitech.wikimedia.org/wiki/PyBal#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPyBalBGPUnstable [13:31:12] (03CR) 10Elukey: [C:03+1] "I left some comments here and there but overall it looks good to me. The only mildly concerning thing is that a cluster role for the contr" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305973 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:31:14] FIRING: PfwCoreBGPDown: ... [13:31:20] Fundraising Firewall core BGP session down between pfw1-eqiad and cr1-eqiad (208.80.154.200) - group Production - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqiad&var-device=pfw1-eqiad:9804&var-bgp_group=Production&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DPfwCoreBGPDown [13:31:27] !ack [13:31:28] 8123 (ACKED) [3x] PyBalBGPUnstable lvs sre (pybal 64600 208.80.154.196 eqiad) [13:31:32] FIRING: [4x] CloudCoreBGPDown: Cloud (WMCS) BGP session down between cloudsw1-c8-eqiad and cr1-eqiad (10.64.147.16) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCloudCoreBGPDown [13:31:36] (03PS5) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [13:31:46] <_joe_> topranks: I assume this is you? [13:32:13] _joe_: yep, damn I forogt about the cloudsw [13:32:34] as is the pybal one [13:32:35] <_joe_> no I mean the lvs pag.e we just got [13:32:40] <_joe_> ah ok [13:32:42] (03CR) 10Elukey: [C:03+1] topolvm: customise the imported chart for WMF [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305974 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:32:51] RECOVERY - OSPF status on cr2-eqdfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:32:54] _joe_: we we're wrong to say they'd not page [13:33:09] lvs is set to page, I'll downtime the lvs specifically when I do cr2 later [13:33:44] topranks: yeah, we put in that page so that we know if the lvs BGP session is failing essentially [13:33:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:34:22] sukhe: yeah it makes sense, we incorrectly assessed that wouldn't happen for hosts earlier (wikikube etc don't page), should have considered LVS [13:34:29] (03CR) 10Elukey: [C:03+1] topolvm: tighten controller RBAC [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305975 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:34:42] RESOLVED: CoreRouterInterfaceDown: Core router interface down - pfw1-eqiad:xe-0/2/0 (Core: cr1-eqiad:xe-3/1/7 {#delete_me_and_replace_with_two_cables1}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=pfw1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:34:45] !log caro@deploy1003 Finished scap sync-world: Backport for [[gerrit:1306842|EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s) [13:34:48] T430741: VisualEditor freezes in pre-save if all checks are collapsed - https://phabricator.wikimedia.org/T430741 [13:35:07] alright i'm done, thanks! [13:35:25] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad [13:35:33] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12075997 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=c1065c24-e5fc-46f8-8e8c-c08a491f7877) set by... [13:36:06] (03CR) 10Ssingh: [C:03+1] "Thank you!" [dns] - 10https://gerrit.wikimedia.org/r/1306926 (https://phabricator.wikimedia.org/T427282) (owner: 10Muehlenhoff) [13:36:11] RESOLVED: PfwCoreBGPDown: ... [13:36:11] Fundraising Firewall core BGP session down between pfw1-eqiad and cr1-eqiad (208.80.154.200) - group Production - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=eqiad&var-device=pfw1-eqiad:9804&var-bgp_group=Production&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DPfwCoreBGPDown [13:36:32] RESOLVED: [4x] CloudCoreBGPDown: Cloud (WMCS) BGP session down between cloudsw1-c8-eqiad and cr1-eqiad (10.64.147.16) - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCloudCoreBGPDown [13:36:48] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db1160 to s4 master [puppet] - 10https://gerrit.wikimedia.org/r/1306927 (https://phabricator.wikimedia.org/T430817) [13:36:54] (03PS1) 10Gerrit maintenance bot: wmnet: Update s4-master alias [dns] - 10https://gerrit.wikimedia.org/r/1306928 (https://phabricator.wikimedia.org/T430817) [13:37:01] (03CR) 10Elukey: [C:03+1] topolvm: scrape controller/node metrics via prometheus.io annotations [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306222 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:37:01] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad [13:37:08] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12076012 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=c9fa9a87-4c81-4974-9c0a-6a8c8205b869) set by... [13:37:17] !log jmm@cumin2003 START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org [13:37:25] FIRING: SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:37:46] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12076017 (10ops-monitoring-bot) VM urldownloader1005.wikimedia.org rebooted by jmm@cumin2003 with reason: bump resources [13:38:38] (03CR) 10Jelto: [V:03+1 C:03+2] gitlab: also enable restricted robots.txt on replicas [puppet] - 10https://gerrit.wikimedia.org/r/1306861 (https://phabricator.wikimedia.org/T430563) (owner: 10Jelto) [13:38:46] (03CR) 10Elukey: [C:03+1] topolvm-crds: add the TopoLVM CRD for version 0.38.1 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305976 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:39:28] (03CR) 10Elukey: [C:03+1] admin_ng: define the topolvm CSI releases [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306223 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:39:47] (03CR) 10Elukey: [C:03+1] admin_ng: enable the topolvm CSI driver on dse-k8s [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305978 (https://phabricator.wikimedia.org/T429331) (owner: 10Btullis) [13:39:51] (03PS6) 10Giuseppe Lavagetto: cache::varnish: switch known client rate limits to hp-generated data [puppet] - 10https://gerrit.wikimedia.org/r/1306503 (https://phabricator.wikimedia.org/T422249) [13:41:10] 06SRE, 06Infrastructure-Foundations, 10netops: GSHUT (and other?) community matching/actions not working on SR-Linux - https://phabricator.wikimedia.org/T430810#12076032 (10cmooney) [13:41:46] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org [13:41:50] !log atsuko@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie [13:43:35] !log jmm@cumin2003 START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org [13:44:04] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Move URL downloaders to trixie - https://phabricator.wikimedia.org/T427282#12076043 (10ops-monitoring-bot) VM urldownloader1006.wikimedia.org rebooted by jmm@cumin2003 with reason: bump resources [13:44:32] !log atsuko@cumin2003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie [13:44:51] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2086 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:44:57] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2061 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:44:57] PROBLEM - ElasticSearch health check for shards on 9443 on search.svc.codfw.wmnet is CRITICAL: CRITICAL - elasticsearch https://search.svc.codfw.wmnet:9443/_cluster/health error while fetching: HTTPSConnectionPool(host=search.svc.codfw.wmnet, port=9443): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:44:57] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2063 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:44:59] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2081 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:44:59] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2084 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:44:59] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2074 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:44:59] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2070 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:01] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2090 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:01] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2093 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:01] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2073 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:01] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2077 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:01] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2088 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:03] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2097 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:03] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2104 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:03] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2099 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:03] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2098 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:04] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2111 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:04] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2105 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:09] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2114 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:09] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2112 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:11] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2065 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:11] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2087 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:19] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2102 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:19] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2067 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:45:47] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2106 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [13:46:56] (03PS1) 10Ozge: ml-services: Bump revscoring production images to 2026-06-23-094330-publish [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306930 (https://phabricator.wikimedia.org/T429675) [13:47:24] (03CR) 10Andrew Bogott: [C:03+2] magnum policy.yaml: replace rule:admin_or_user with rule:admin_or_member [puppet] - 10https://gerrit.wikimedia.org/r/1306775 (https://phabricator.wikimedia.org/T430680) (owner: 10Andrew Bogott) [13:48:06] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org [13:48:31] (03PS2) 10Ozge: ml-services: Bump revscoring production images to 2026-06-23-094330-publish [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306930 (https://phabricator.wikimedia.org/T429675) [13:48:54] (03PS3) 10Ozge: ml-services: Bump revscoring production images to 2026-06-23-094330-publish [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306930 (https://phabricator.wikimedia.org/T429675) [13:51:44] (03CR) 10Giuseppe Lavagetto: [C:03+2] cache::varnish: switch known client rate limits to hp-generated data [puppet] - 10https://gerrit.wikimedia.org/r/1306503 (https://phabricator.wikimedia.org/T422249) (owner: 10Giuseppe Lavagetto) [13:52:00] jouncebot: nowandnext [13:52:00] For the next 0 hour(s) and 7 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1300) [13:52:00] In 0 hour(s) and 7 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1400) [13:52:29] anzx: Did you find a deployer? [13:52:52] yes [13:53:09] need a deployer [13:53:26] I can bundle yours with mine [13:53:29] Let me review it [13:53:31] ok [13:55:47] (03PS1) 10Gerrit maintenance bot: mariadb: Promote db2159 to s7 master [puppet] - 10https://gerrit.wikimedia.org/r/1306931 (https://phabricator.wikimedia.org/T430826) [13:56:05] (03CR) 10Dreamy Jazz: [C:03+2] eswikisource: add wikibooks as importsource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306456 (https://phabricator.wikimedia.org/T430537) (owner: 10Anzx) [13:56:11] (03PS6) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [13:56:53] !log reboot routing-enginer RE0 on cr1-eqiad T417873 [13:56:55] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:56:56] T417873: eqiad: upgrade routers (2026) - https://phabricator.wikimedia.org/T417873 [13:57:04] (03Merged) 10jenkins-bot: eswikisource: add wikibooks as importsource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306456 (https://phabricator.wikimedia.org/T430537) (owner: 10Anzx) [13:57:12] (03CR) 10TrainBranchBot: [C:03+2] "Copied votes on follow-up patch sets have been updated:" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306456 (https://phabricator.wikimedia.org/T430537) (owner: 10Anzx) [13:57:13] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306916 (owner: 10Dreamy Jazz) [13:57:13] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306925 (https://phabricator.wikimedia.org/T425048) (owner: 10Dreamy Jazz) [13:57:22] !log fceratto@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 T430826 [13:57:25] RESOLVED: SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:57:25] T430826: Switchover s7 master (db2220 -> db2159) - https://phabricator.wikimedia.org/T430826 [13:57:55] (03PS1) 10Jforrester: wikifunctions: Upgrade evaluators from 2026-06-25-145651 to 2026-06-30-213833 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306932 (https://phabricator.wikimedia.org/T411110) [13:58:08] (03PS1) 10Jforrester: wikifunctions: Upgrade orchestrator from 2026-06-23-115555 to 2026-07-01-115510 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306933 (https://phabricator.wikimedia.org/T411110) [13:58:11] (03Merged) 10jenkins-bot: Move non temporary accounts settings out TA section [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306916 (owner: 10Dreamy Jazz) [13:58:14] (03Merged) 10jenkins-bot: Remove TA patrol rights from users on fishbowl + private [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306925 (https://phabricator.wikimedia.org/T425048) (owner: 10Dreamy Jazz) [13:58:42] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1306456|eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916|Move non temporary accounts settings out TA section]], [[gerrit:1306925|Remove TA patrol rights from users on fishbowl + private (T425048)]] [13:58:48] T430537: Add wikibooks to wgImportSources in eswikisource - https://phabricator.wikimedia.org/T430537 [13:58:48] T425048: Turn off temporary account viewer stuffs on private and fishbowl wikis - https://phabricator.wikimedia.org/T425048 [13:59:07] !log fceratto@cumin1003 dbctl commit (dc=all): 'Set db2159 with weight 0 T430826', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json [14:00:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1400) [14:00:26] Okie-dokie. [14:00:33] (03CR) 10Slyngshede: varnish: add tests for text thumbnail config (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [14:00:41] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade evaluators from 2026-06-25-145651 to 2026-06-30-213833 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306932 (https://phabricator.wikimedia.org/T411110) (owner: 10Jforrester) [14:00:52] (03PS7) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [14:00:53] !log dreamyjazz@deploy1003 anzx, dreamyjazz: Backport for [[gerrit:1306456|eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916|Move non temporary accounts settings out TA section]], [[gerrit:1306925|Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:01:02] looking [14:01:11] Thanks [14:01:32] FIRING: [2x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:01:36] (03CR) 10Muehlenhoff: [C:03+2] Revert urldownloader in codfw to 2003 [dns] - 10https://gerrit.wikimedia.org/r/1306926 (https://phabricator.wikimedia.org/T427282) (owner: 10Muehlenhoff) [14:01:43] !log jmm@dns1004 START - running authdns-update [14:02:25] FIRING: SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:02:51] (03Merged) 10jenkins-bot: wikifunctions: Upgrade evaluators from 2026-06-25-145651 to 2026-06-30-213833 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306932 (https://phabricator.wikimedia.org/T411110) (owner: 10Jforrester) [14:02:51] (03CR) 10Federico Ceratto: [C:03+2] mariadb: Promote db2159 to s7 master [puppet] - 10https://gerrit.wikimedia.org/r/1306931 (https://phabricator.wikimedia.org/T430826) (owner: 10Gerrit maintenance bot) [14:03:10] !log cmooney@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad [14:03:23] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12076313 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=aab5730e-3995-411e-aa0c-e48fc42c7ad5) set by... [14:03:27] !log flipping cr1-eqiad active routing-enginer back to RE0 T417873 [14:03:28] Dreamy_Jazz: looks good ok to sync, [14:03:29] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:03:30] T417873: eqiad: upgrade routers (2026) - https://phabricator.wikimedia.org/T417873 [14:03:35] Thanks [14:03:39] Still testing mine [14:03:39] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:03:47] !log jmm@dns1004 END - running authdns-update [14:04:11] !log Starting s7 codfw failover from db2220 to db2159 - T430826 [14:04:14] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:04:15] T430826: Switchover s7 master (db2220 -> db2159) - https://phabricator.wikimedia.org/T430826 [14:04:23] !log dreamyjazz@deploy1003 anzx, dreamyjazz: Continuing with deployment [14:04:33] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:04:35] I think I have the right downtimes now - but I cr1-eqiad will go offline again momentarily [14:04:37] !log filippo@cumin1003 conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org [14:04:44] !log filippo@cumin1003 conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org [14:04:52] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:05:04] !log fceratto@cumin1003 dbctl commit (dc=all): 'Promote db2159 to s7 primary T430826', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json [14:06:17] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:06:23] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:06:32] FIRING: [6x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:06:58] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:07:20] (03CR) 10Jforrester: [C:03+2] wikifunctions: Upgrade orchestrator from 2026-06-23-115555 to 2026-07-01-115510 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306933 (https://phabricator.wikimedia.org/T411110) (owner: 10Jforrester) [14:07:29] !log fceratto@cumin1003 dbctl commit (dc=all): 'Depool db2220 T430826', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json [14:07:36] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover [14:07:51] PROBLEM - OSPF status on cr2-eqdfw is CRITICAL: OSPFv2: 6/7 UP : OSPFv3: 6/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:08:40] (03PS1) 10Muehlenhoff: docker-registry: Allow image builds to be pushed from build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1306936 (https://phabricator.wikimedia.org/T417389) [14:08:43] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1306456|eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916|Move non temporary accounts settings out TA section]], [[gerrit:1306925|Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s) [14:08:48] T430537: Add wikibooks to wgImportSources in eswikisource - https://phabricator.wikimedia.org/T430537 [14:08:48] T425048: Turn off temporary account viewer stuffs on private and fishbowl wikis - https://phabricator.wikimedia.org/T425048 [14:09:13] Dreamy_Jazz: thanks for deploying [14:09:17] Np [14:09:33] (03Merged) 10jenkins-bot: wikifunctions: Upgrade orchestrator from 2026-06-23-115555 to 2026-07-01-115510 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306933 (https://phabricator.wikimedia.org/T411110) (owner: 10Jforrester) [14:09:45] FIRING: CirrusStreamingUpdaterFlinkJobUnstable: cirrus_streaming_updater_consumer_search_codfw in codfw (k8s) is unstable - https://wikitech.wikimedia.org/wiki/Search#Streaming_Updater - https://grafana.wikimedia.org/d/K9x0c4aVk/flink-app?var-datasource=codfw+prometheus%2Fk8s&var-namespace=cirrus-streaming-updater&var-helm_release=consumer-search - https://alerts.wikimedia.org/?q=alertname%3DCirrusStreamingUpdaterFlinkJobUnstable [14:10:22] !log jforrester@deploy1003 helmfile [staging] START helmfile.d/services/wikifunctions: apply [14:10:40] (03CR) 10Bartosz Wójtowicz: [C:03+1] "Thank you!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306930 (https://phabricator.wikimedia.org/T429675) (owner: 10Ozge) [14:10:47] !log jforrester@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifunctions: apply [14:10:51] RECOVERY - OSPF status on cr2-eqdfw is OK: OSPFv2: 7/7 UP : OSPFv3: 7/7 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:11:19] !log jforrester@deploy1003 helmfile [codfw] START helmfile.d/services/wikifunctions: apply [14:11:32] FIRING: [9x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:12:04] !log jforrester@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply [14:12:12] !log jforrester@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifunctions: apply [14:12:40] !log jforrester@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply [14:13:08] !log fceratto@cumin1003 END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover [14:13:47] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover [14:14:06] !log re-enable routing-engine graceful-failover on cr1-eqiad T417873 [14:14:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:14:09] T417873: eqiad: upgrade routers (2026) - https://phabricator.wikimedia.org/T417873 [14:15:01] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch2100 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [14:15:21] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover [14:16:32] FIRING: [15x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:16:49] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover [14:16:49] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [14:17:12] (03CR) 10Ozge: [C:03+2] ml-services: Bump revscoring production images to 2026-06-23-094330-publish [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306930 (https://phabricator.wikimedia.org/T429675) (owner: 10Ozge) [14:17:53] jouncebot: nowandnext [14:17:53] For the next 0 hour(s) and 42 minute(s): Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1400) [14:17:53] In 0 hour(s) and 12 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1430) [14:17:54] (03PS1) 10Dreamy Jazz: Remove group permissions definitions later in the request [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306938 (https://phabricator.wikimedia.org/T425048) [14:18:01] Need to make a follow-up to my deploy [14:18:03] (03CR) 10CI reject: [V:04-1] Remove group permissions definitions later in the request [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306938 (https://phabricator.wikimedia.org/T425048) (owner: 10Dreamy Jazz) [14:18:14] (03PS2) 10Dreamy Jazz: Remove group permissions definitions later in the request [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306938 (https://phabricator.wikimedia.org/T425048) [14:20:18] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306938 (https://phabricator.wikimedia.org/T425048) (owner: 10Dreamy Jazz) [14:20:32] (03Merged) 10jenkins-bot: ml-services: Bump revscoring production images to 2026-06-23-094330-publish [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306930 (https://phabricator.wikimedia.org/T429675) (owner: 10Ozge) [14:21:32] FIRING: [18x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:21:42] (03Merged) 10jenkins-bot: Remove group permissions definitions later in the request [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306938 (https://phabricator.wikimedia.org/T425048) (owner: 10Dreamy Jazz) [14:22:01] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover [14:22:09] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1306938|Remove group permissions definitions later in the request (T425048)]] [14:22:12] T425048: Turn off temporary account viewer stuffs on private and fishbowl wikis - https://phabricator.wikimedia.org/T425048 [14:24:19] !log dreamyjazz@deploy1003 dreamyjazz: Backport for [[gerrit:1306938|Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:26:23] (03CR) 10Muehlenhoff: [C:03+2] Drop use of MW_APPSERVER_NETWORKS for ircstream now that mw* servers are gone [puppet] - 10https://gerrit.wikimedia.org/r/1214094 (https://phabricator.wikimedia.org/T411508) (owner: 10Muehlenhoff) [14:26:32] FIRING: [26x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:26:43] !log dreamyjazz@deploy1003 dreamyjazz: Continuing with deployment [14:29:36] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover [14:30:05] Deploy window Wikifunctions Services UTC Afternoon (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1400) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1430) [14:31:06] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1306938|Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s) [14:31:08] !log POWERING DOWN CR1-EQIAD for line card installation T426343 [14:31:09] T425048: Turn off temporary account viewer stuffs on private and fishbowl wikis - https://phabricator.wikimedia.org/T425048 [14:31:11] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:31:12] T426343: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343 [14:31:32] FIRING: [30x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [14:31:58] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2104 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: green, timed_out: False, number_of_nodes: 27, number_of_data_nodes: 27, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1723, active_shards: 5169, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [14:31:58] assigned_shards: 0, number_of_pending_tasks: 21690, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 2957539, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [14:31:58] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2111 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: green, timed_out: False, number_of_nodes: 27, number_of_data_nodes: 27, discovered_master: True, active_primary_shards: 1723, active_shards: 5169, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, delayed_unassigned_shards: 0, number [14:31:58] ing_tasks: 21690, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 2957585, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [14:31:58] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2070 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: green, timed_out: False, number_of_nodes: 27, number_of_data_nodes: 27, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1723, active_shards: 5169, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 0, de [14:32:44] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2086 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: yellow, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1723, active_shards: 4786, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 383, [14:32:44] _unassigned_shards: 383, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 92.59044302573032 https://wikitech.wikimedia.org/wiki/Search%23Administration [14:32:50] RECOVERY - ElasticSearch health check for shards on 9443 on search.svc.codfw.wmnet is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: yellow, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1723, active_shards: 4786, relocating_shards: 0, initializing_shards: 0, unassigned_sha [14:32:50] , delayed_unassigned_shards: 383, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 92.59044302573032 https://wikitech.wikimedia.org/wiki/Search%23Administration [14:32:50] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2061 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: yellow, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, active_primary_shards: 1723, active_shards: 4786, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 383, delayed_unassigned_shards: 383, n [14:32:50] _pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 92.59044302573032 https://wikitech.wikimedia.org/wiki/Search%23Administration [14:32:50] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch2063 is OK: OK - elasticsearch status production-search-omega-codfw: cluster_name: production-search-omega-codfw, status: yellow, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1723, active_shards: 4786, relocating_shards: 0, initializing_shards: 0, unassigned_shards: 383, [14:32:51] _unassigned_shards: 383, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 92.59044302573032 https://wikitech.wikimedia.org/wiki/Search%23Administration [14:55:49] (03CR) 10Gergő Tisza: [C:03+1] "Ack. I don't think there's much point in testing the tests, those are run during deployment anyway." [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) (owner: 10Krinkle) [14:56:14] RECOVERY - Host re0.cr2-eqiad.mgmt is UP: PING OK - Packet loss = 0%, RTA = 1.10 ms [14:57:19] sorry topranks that wasn't aimed at you, it was meant as "what Andrew's days would look like if he worked in the datacenter" [14:57:44] RECOVERY - OSPF status on cr1-magru is OK: OSPFv2: 3/3 UP : OSPFv3: 3/3 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [14:58:44] RECOVERY - OSPF status on cr1-drmrs is OK: OSPFv2: 4/4 UP : OSPFv3: 4/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [15:00:42] 10ops-ulsfo, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: ULSFO: OOB IPV6 down - https://phabricator.wikimedia.org/T430599#12076680 (10Papaul) Information sent to DR [15:00:43] (03CR) 10Gergő Tisza: [C:03+1] varnish: Add edge fixup for corrupt upload.wm.o urls from mobileapps (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) (owner: 10Krinkle) [15:02:14] RECOVERY - Host cr2-eqiad IPv6 is UP: PING OK - Packet loss = 0%, RTA = 1.15 ms [15:02:48] !log ongoing maintenance on lsw1-a8-codfw [15:02:49] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:04:02] RECOVERY - Check unit status of sync-puppet-volatile on puppetserver1002 is OK: OK: Status of the systemd unit sync-puppet-volatile https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:04:02] RECOVERY - Check unit status of sync-puppet-volatile on puppetserver2001 is OK: OK: Status of the systemd unit sync-puppet-volatile https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:04:28] We lost some PostgreSQL databases behind Airflow. I think I'll be reverting to backup. [15:04:30] <_joe_> what now [15:04:36] <_joe_> !incidents [15:04:37] 8120 (ACKED) es1039 (paged)/MariaDB Replica SQL: es7 (paged) [15:04:37] 8121 (ACKED) es1039 (paged)/MariaDB Replica IO: es7 (paged) [15:04:37] 8122 (ACKED) es1039 (paged)/MariaDB Replica Lag: es7 (paged) [15:04:37] 8123 (ACKED) [3x] PyBalBGPUnstable lvs sre (pybal 64600 208.80.154.196 eqiad) [15:04:38] 8134 (ACKED) Host 10.3.0.1 [15:04:38] 8135 (ACKED) Host cloudelastic.wikimedia.org [15:04:38] 8143 (UNACKED) NELHigh sre (thanos-rule@main tcp.timed_out) [15:04:39] 8127 (RESOLVED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [15:04:39] 8126 (RESOLVED) VarnishUnavailable global sre (varnish-text thanos-rule@main) [15:04:40] 8124 (RESOLVED) [50x] ProbeDown sre () [15:04:40] 8141 (RESOLVED) NELHigh sre (thanos-rule@main tcp.timed_out) [15:04:40] RECOVERY - Check unit status of push_cross_cluster_settings_9400 on cirrussearch2106 is OK: OK: Status of the systemd unit push_cross_cluster_settings_9400 https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:04:41] 8119 (RESOLVED) es1039 (paged)/mysqld processes (paged) [15:04:41] 8118 (RESOLVED) Host es1039 (paged) [15:04:42] 8112 (RESOLVED) es1039 (paged)/MariaDB read only es7 (paged) [15:04:42] 8116 (RESOLVED) es1048 (paged)/MariaDB Replica Lag: es7 (paged) [15:04:43] 8114 (RESOLVED) es1040 (paged)/MariaDB Replica Lag: es7 (paged) [15:04:43] 8115 (RESOLVED) es1035 (paged)/MariaDB Replica Lag: es7 (paged) [15:04:44] 8117 (RESOLVED) es2039 (paged)/MariaDB Replica Lag: es7 (paged) [15:04:44] 8111 (RESOLVED) es2039 (paged)/MariaDB Replica IO: es7 (paged) [15:04:45] 8108 (RESOLVED) es1035 (paged)/MariaDB Replica IO: es7 (paged) [15:04:45] 8109 (RESOLVED) es1048 (paged)/MariaDB Replica IO: es7 (paged) [15:04:46] 8110 (RESOLVED) es1040 (paged)/MariaDB Replica IO: es7 (paged) [15:04:46] 8113 (RESOLVED) es1039 (paged)/mysqld processes (paged) [15:04:47] 8107 (RESOLVED) Host es1039 (paged) [15:05:05] jinxer-wm is missing [15:06:18] !incidents [15:06:19] 8120 (ACKED) es1039 (paged)/MariaDB Replica SQL: es7 (paged) [15:06:19] 8121 (ACKED) es1039 (paged)/MariaDB Replica IO: es7 (paged) [15:06:19] 8122 (ACKED) es1039 (paged)/MariaDB Replica Lag: es7 (paged) [15:06:19] 8123 (ACKED) [3x] PyBalBGPUnstable lvs sre (pybal 64600 208.80.154.196 eqiad) [15:06:19] 8134 (ACKED) Host 10.3.0.1 [15:06:20] 8135 (ACKED) Host cloudelastic.wikimedia.org [15:06:20] 8143 (UNACKED) NELHigh sre (thanos-rule@main tcp.timed_out) [15:06:21] 8127 (RESOLVED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [15:06:21] 8126 (RESOLVED) VarnishUnavailable global sre (varnish-text thanos-rule@main) [15:06:22] 8124 (RESOLVED) [50x] ProbeDown sre () [15:06:22] 8141 (RESOLVED) NELHigh sre (thanos-rule@main tcp.timed_out) [15:06:23] 8119 (RESOLVED) es1039 (paged)/mysqld processes (paged) [15:06:23] 8118 (RESOLVED) Host es1039 (paged) [15:06:24] 8112 (RESOLVED) es1039 (paged)/MariaDB read only es7 (paged) [15:06:24] 8116 (RESOLVED) es1048 (paged)/MariaDB Replica Lag: es7 (paged) [15:06:25] 8114 (RESOLVED) es1040 (paged)/MariaDB Replica Lag: es7 (paged) [15:06:25] 8115 (RESOLVED) es1035 (paged)/MariaDB Replica Lag: es7 (paged) [15:06:26] 8117 (RESOLVED) es2039 (paged)/MariaDB Replica Lag: es7 (paged) [15:06:26] 8111 (RESOLVED) es2039 (paged)/MariaDB Replica IO: es7 (paged) [15:06:27] 8108 (RESOLVED) es1035 (paged)/MariaDB Replica IO: es7 (paged) [15:06:27] 8109 (RESOLVED) es1048 (paged)/MariaDB Replica IO: es7 (paged) [15:06:28] 8110 (RESOLVED) es1040 (paged)/MariaDB Replica IO: es7 (paged) [15:06:28] 8113 (RESOLVED) es1039 (paged)/mysqld processes (paged) [15:06:29] 8107 (RESOLVED) Host es1039 (paged) [15:06:31] !ack [15:06:32] 8143 (ACKED) NELHigh sre (thanos-rule@main tcp.timed_out) [15:06:34] !log pt1979@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet [15:06:50] (03PS9) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [15:07:19] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment for rscout - https://phabricator.wikimedia.org/T430594#12076734 (10fgiunchedi) @Rsilvola FYI this is pending your approval [15:07:43] !log pt1979@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet [15:09:35] !log fceratto@cumin1003 END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover [15:10:54] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:11:32] FIRING: [24x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [15:11:35] (03PS10) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [15:12:06] !log pt1979@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde [15:12:11] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply [15:12:13] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: codfw: rack A8 maintenance 2026-07-01 10:00 am CT - https://phabricator.wikimedia.org/T429856#12076785 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=c0d4701e-c375-434f-b42c-5115f14a1c21) set by pt1979@cumin1003 fo... [15:12:15] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply [15:12:33] !log atsuko@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage [15:12:34] <_joe_> !log restarted manually alertmanager-irc-relay [15:12:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:12:41] FIRING: [32x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:15:45] RESOLVED: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:16:32] RESOLVED: [13x] SLOBudgetBurn: Search update lag is below 95% target in codfw - https://alerts.wikimedia.org/?q=alertname%3DSLOBudgetBurn [15:17:11] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.14 point update - https://phabricator.wikimedia.org/T426759#12076844 (10MoritzMuehlenhoff) [15:17:32] (03PS11) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [15:19:16] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply [15:19:59] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply [15:20:02] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage [15:21:06] !log Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - T429861 [15:21:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:21:09] T429861: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861 [15:22:33] !log brett@cumin2002 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - T429861 [15:23:14] (03CR) 10Jgiannelos: PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) (owner: 10Neriah) [15:23:36] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply [15:23:40] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply [15:25:56] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply [15:26:39] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply [15:29:23] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply [15:30:21] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply [15:31:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [15:32:16] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply [15:32:25] (03PS2) 10Jgiannelos: PageBundleParserOutputConverter: Avoid revision lookup for bogus title [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) [15:32:51] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply [15:35:00] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply [15:35:31] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply [15:36:45] PROBLEM - BFD status on ssw1-a8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:36:45] PROBLEM - BFD status on ssw1-a1-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:36:57] (03CR) 10Jgiannelos: "recheck" [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) (owner: 10Jgiannelos) [15:37:23] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply [15:37:25] FIRING: [25x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:37:25] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to deployment for rscout - https://phabricator.wikimedia.org/T430594#12076920 (10Rsilvola) Approving as Randall's manager. [15:37:39] FIRING: CoreBGPDown: Core BGP session down between ssw1-a8-codfw and lsw1-a8-codfw (10.192.252.10) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=ssw1-a8-codfw:9804&var-bgp_group=EVPN_IBGP&var-bgp_neighbor=lsw1-a8-codfw - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:38:24] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply [15:38:51] FIRING: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/7 (Core: lsw1-a8-codfw:et-0/0/55 {#230403800025}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:39:37] (03CR) 10CI reject: [V:04-1] PageBundleParserOutputConverter: Avoid revision lookup for bogus title [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) (owner: 10Jgiannelos) [15:40:06] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply [15:41:05] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply [15:41:14] (03Abandoned) 10Jgiannelos: PageBundleParserOutputConverter: Avoid revision lookup for bogus title [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) (owner: 10Jgiannelos) [15:41:33] FIRING: [2x] KubernetesCalicoDown: wikikube-worker2042.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [15:41:33] (03PS2) 10Neriah: PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) [15:42:25] FIRING: [25x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:42:39] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-a8-codfw (10.192.252.10) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:47:39] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-a8-codfw (10.192.252.10) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:47:43] RECOVERY - BFD status on ssw1-a8-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:47:43] RECOVERY - BFD status on ssw1-a1-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:48:24] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie [15:48:27] (03Restored) 10Jgiannelos: PageBundleParserOutputConverter: Avoid revision lookup for bogus title [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) (owner: 10Jgiannelos) [15:48:37] (03PS3) 10Jgiannelos: PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) (owner: 10Neriah) [15:48:37] (03PS3) 10Jgiannelos: PageBundleParserOutputConverter: Avoid revision lookup for bogus title [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) [15:48:51] RESOLVED: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/7 (Core: lsw1-a8-codfw:et-0/0/55 {#230403800025}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:49:49] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [15:50:58] !log fceratto@cumin1003 START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover [15:50:59] !log ozge@deploy1003 helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . [15:51:09] !log fceratto@cumin1003 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover [15:51:12] !log ozge@deploy1003 helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' . [15:51:33] RESOLVED: [2x] KubernetesCalicoDown: wikikube-worker2042.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [15:52:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-a8-codfw (10.192.252.10) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:53:34] (03PS12) 10CDobbins: varnish: add tests for text thumbnail config [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) [15:55:43] (03CR) 10CDobbins: varnish: add tests for text thumbnail config (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [15:55:44] !log pt1979@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet [15:55:46] !log pt1979@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet [15:57:42] !log atsuko@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie [15:59:21] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) (owner: 10Jgiannelos) [15:59:30] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) (owner: 10Neriah) [15:59:41] !log pt1979@cumin1003 START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt [15:59:43] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461) {#3909}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [15:59:43] !log pt1979@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt [16:00:09] !log atsuko@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie [16:00:29] !log ongoing maintenance on lsw1-b2-codfw [16:00:29] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:02:38] (03CR) 10CDobbins: "text: 0 tests failed, 0 tests skipped, 41 tests passed" [puppet] - 10https://gerrit.wikimedia.org/r/1306911 (https://phabricator.wikimedia.org/T427465) (owner: 10CDobbins) [16:04:05] inflatador: hey are you the right person to ping to depool wdqs2026? [16:04:18] papaul sure, do you need me to do that? [16:04:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [16:05:42] inflatador: yes please https://phabricator.wikimedia.org/T429861 [16:05:44] pt1979@cumin1003 downtime (PID 561825) is awaiting input [16:06:43] !log pt1979@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde [16:06:55] 10ops-codfw, 06SRE, 06Data-Persistence, 06Data-Platform-SRE, and 5 others: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861#12077048 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=e7fbbad5-cec3-439b-bf2e-a0a0bbae2b6b) set by pt1979@cumin1003... [16:07:29] !log pt1979@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance [16:09:43] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:11:33] (03CR) 10Jasmine: [C:03+2] Add Kubernetes POD IP reverse range delegations for wikikube-ctrl1005 [dns] - 10https://gerrit.wikimedia.org/r/1302996 (https://phabricator.wikimedia.org/T418920) (owner: 10Jasmine) [16:12:44] !log jasmine@dns1004 START - running authdns-update [16:13:43] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, July 01 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-it" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1305773 (https://phabricator.wikimedia.org/T430227) (owner: 10Jdrewniak) [16:14:43] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:14:50] !log jasmine@dns1004 END - running authdns-update [16:15:41] (03CR) 10Jasmine: [C:03+2] Add new control plane wikikube-ctrl1005 to etcd-server SRV record [dns] - 10https://gerrit.wikimedia.org/r/1300942 (https://phabricator.wikimedia.org/T418920) (owner: 10Jasmine) [16:15:54] !log jasmine@dns1004 START - running authdns-update [16:16:01] !log atsuko@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage [16:16:18] !log jhancock@cumin2002 START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART [16:16:31] !log jhancock@cumin2002 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART [16:18:06] !log jasmine@dns1004 END - running authdns-update [16:18:14] !log atsuko@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage [16:19:37] (03CR) 10Krinkle: varnish: Add edge fixup for corrupt upload.wm.o urls from mobileapps (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) (owner: 10Krinkle) [16:19:57] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage [16:23:51] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage [16:28:33] !log dancy@deploy1003 Installing scap version "4.271.0" for 2 host(s) [16:30:28] !log dancy@deploy1003 Installation of scap version "4.271.0" completed for 2 hosts [16:33:08] (03PS8) 10Krinkle: varnish: Add edge fixup for corrupt upload.wm.o urls from mobileapps [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) [16:33:43] PROBLEM - BFD status on ssw1-a1-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [16:33:43] PROBLEM - BFD status on ssw1-a8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [16:34:39] FIRING: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b2-codfw (10.192.252.11) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [16:34:51] FIRING: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/8 (Core: lsw1-b2-codfw:et-0/0/55 {#230403800009}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [16:35:30] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 02 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployca" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1305773 (https://phabricator.wikimedia.org/T430227) (owner: 10Jdrewniak) [16:37:33] FIRING: [5x] KubernetesCalicoDown: wikikube-ctrl2004.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [16:37:41] FIRING: [2x] ProbeDown: Service wikikube-ctrl2004:6443 has failed probes (http_codfw_kube_apiserver_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#wikikube-ctrl2004:6443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:38:02] !ack [16:38:02] 8144 (ACKED) [2x] ProbeDown sre (wikikube-ctrl2004:6443 probes/custom codfw) [16:38:26] There's codfw Rack B2 work going on right now, the switch is restarting AFAICT [16:38:28] hi folks, that may be me [16:38:40] (looking) [16:39:03] <_joe_> I assumed it was the codfw maintenance [16:39:43] FIRING: JobUnavailable: Reduced availability for job mysql-test in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:40:08] <_joe_> jasmine_: what were you doing that might have caused this? [16:40:17] it went down at 16:32:00 [16:40:23] <_joe_> the server is unreachable as far as I can tell [16:40:33] federico3: Yes, it's been down for around that long [16:40:42] (the switch) [16:41:05] <_joe_> ok, we still have redundancy for the k8s apiserver afaict [16:41:06] d'you have a task or logs from the switch? [16:41:22] https://phabricator.wikimedia.org/T429861 [16:41:40] FIRING: [4x] KubernetesRsyslogDown: rsyslog on wikikube-worker2129:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [16:42:03] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-codfw in codfw - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=codfw%20prometheus/ops&var-kafka_cluster=logging-codfw - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [16:42:19] !ack [16:42:20] All incidents are already acked. [16:42:23] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie [16:43:43] RECOVERY - BFD status on ssw1-a1-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [16:43:43] RECOVERY - BFD status on ssw1-a8-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [16:44:00] brett: are we confident that's the reason? Can I put it in the task? [16:44:27] federico3: yes [16:44:27] <_joe_> servers should be back up now [16:44:39] RESOLVED: [2x] CoreBGPDown: Core BGP session down between ssw1-a1-codfw and lsw1-b2-codfw (10.192.252.11) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [16:44:43] RESOLVED: JobUnavailable: Reduced availability for job mysql-test in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:44:51] RESOLVED: [2x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-a1-codfw:et-0/0/8 (Core: lsw1-b2-codfw:et-0/0/55 {#230403800009}) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [16:45:40] <_joe_> yeah the servers are back, kafka and the kube apiserver should be too [16:45:49] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [16:47:03] RESOLVED: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-codfw in codfw - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=codfw%20prometheus/ops&var-kafka_cluster=logging-codfw - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [16:47:21] 10ops-codfw, 06SRE, 06Data-Persistence, 06Data-Platform-SRE, and 5 others: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861#12077348 (10FCeratto-WMF) This triggered the alert `ProbeDown: Service wikikube-ctrl2004...` - the probe indicated a downtime from 16:32:0... [16:47:32] RESOLVED: [5x] KubernetesCalicoDown: wikikube-ctrl2004.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [16:48:33] !log pt1979@cumin1003 START - Cookbook sre.hosts.remove-downtime for 59 hosts [16:49:02] !log pt1979@cumin1003 END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts [16:49:06] !log Start pybal on lvs2012 - T429861 [16:49:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:49:09] T429861: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861 [16:49:18] !log atsuko@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie [16:50:17] joe: i think you're right (I'm adding a new control plane to eqiad and needed to backtrack for other reasons and thought it might be related) [16:51:36] !log brett@cumin2002 START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet [16:51:38] !log brett@cumin2002 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet [16:51:40] RESOLVED: [4x] KubernetesRsyslogDown: rsyslog on wikikube-worker2129:9105 is missing kubernetes logs - https://wikitech.wikimedia.org/wiki/Kubernetes/Logging#Common_issues - https://alerts.wikimedia.org/?q=alertname%3DKubernetesRsyslogDown [16:51:45] 10ops-codfw, 06SRE, 06Data-Persistence, 06Data-Platform-SRE, and 5 others: codfw: rack B2 maintenance 2026-07-01 11:00 am CT - https://phabricator.wikimedia.org/T429861#12077406 (10FCeratto-WMF) It impacted db2226 https://grafana.wikimedia.org/goto/bfqtb4uge6i9sf?orgId=default and db2203 https://grafana.wi... [16:51:46] !log pt1979@cumin1003 START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt [16:51:47] !log pt1979@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt [16:52:39] !log pt1979@cumin1003 START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet [16:52:40] !log pt1979@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet [16:54:53] FIRING: KubernetesAPILatency: High Kubernetes API latency (LIST pods) on k8s-mlserve@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s-mlserve&var-latency_percentile=0.95&var-verb=LIST - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [16:55:10] (03CR) 10JHathaway: Add sre.hosts.bmc-user-mgmt.py (037 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1302859 (https://phabricator.wikimedia.org/T426180) (owner: 10Elukey) [16:56:11] RESOLVED: [2x] ProbeDown: Service wikikube-ctrl2004:6443 has failed probes (http_codfw_kube_apiserver_ip4) #page - https://wikitech.wikimedia.org/wiki/Runbook#wikikube-ctrl2004:6443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:57:37] (03PS3) 10Hashar: zuul: remove wikimedia.cloud.org from no_proxy [puppet] - 10https://gerrit.wikimedia.org/r/1306952 (https://phabricator.wikimedia.org/T430479) [16:57:37] (03CR) 10Hashar: "https://puppet-compiler.wmflabs.org/output/1306952/7120/zuul2002.codfw.wmnet/index.html" [puppet] - 10https://gerrit.wikimedia.org/r/1306952 (https://phabricator.wikimedia.org/T430479) (owner: 10Hashar) [16:57:54] !log pt1979@cumin1003 START - Cookbook sre.hosts.remove-downtime for 30 hosts [16:58:11] !log pt1979@cumin1003 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1700) [17:06:50] (03PS1) 10Btullis: datahub: move datahub secrets to datahub_config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306961 (https://phabricator.wikimedia.org/T402408) [17:17:57] (03CR) 10Btullis: [C:03+2] datahub: move datahub secrets to datahub_config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306961 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [17:20:15] (03Merged) 10jenkins-bot: datahub: move datahub secrets to datahub_config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306961 (https://phabricator.wikimedia.org/T402408) (owner: 10Btullis) [17:20:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [17:24:24] (03PS40) 10CDobbins: varnish: Add CSP report-only header value [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) [17:27:23] (03PS41) 10CDobbins: varnish: Add CSP report-only header value [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) [17:34:39] FIRING: CirrusSearchThreadPoolRejectionsTooHigh: cirrussearch1086-production-search-eqiad is rejecting excessive amounts of queries due to a full thread pool - https://w.wiki/DTaY - https://grafana.wikimedia.org/goto/aoZBw8pNR?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchThreadPoolRejectionsTooHigh [17:35:16] (03CR) 10CDobbins: "upload: 0 tests failed, 0 tests skipped, 20 tests passed" [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) (owner: 10CDobbins) [17:39:39] RESOLVED: CirrusSearchThreadPoolRejectionsTooHigh: cirrussearch1086-production-search-eqiad is rejecting excessive amounts of queries due to a full thread pool - https://w.wiki/DTaY - https://grafana.wikimedia.org/goto/aoZBw8pNR?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchThreadPoolRejectionsTooHigh [17:40:24] !log jasmine@cumin2002 START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie [17:40:36] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware, 13Patch-For-Review: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12077763 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-ctrl1005.eqiad.wmnet... [17:41:38] RECOVERY - Check unit status of statograph_post on alert1002 is OK: OK: Status of the systemd unit statograph_post https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:46:25] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps2011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:48:25] (03PS1) 10Cwhite: opensearch: add config required to authenticate curator requests [puppet] - 10https://gerrit.wikimedia.org/r/1306969 (https://phabricator.wikimedia.org/T350516) [17:49:03] (03CR) 10CI reject: [V:04-1] opensearch: add config required to authenticate curator requests [puppet] - 10https://gerrit.wikimedia.org/r/1306969 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [17:50:39] (03PS2) 10Cwhite: opensearch: add config required to authenticate curator requests [puppet] - 10https://gerrit.wikimedia.org/r/1306969 (https://phabricator.wikimedia.org/T350516) [17:51:21] (03CR) 10CI reject: [V:04-1] opensearch: add config required to authenticate curator requests [puppet] - 10https://gerrit.wikimedia.org/r/1306969 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [17:58:04] (03PS1) 10WMDE-Fisch: Fix how to check the treatment group [extensions/Cite] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306970 (https://phabricator.wikimedia.org/T415904) [17:58:16] (03CR) 10Daniel Kinzler: smokepy: Add interactive pod (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1301423 (https://phabricator.wikimedia.org/T424825) (owner: 10Daniel Kinzler) [17:58:39] (03PS1) 10WMDE-Fisch: Fix how to check the treatment group [extensions/Cite] (wmf/1.47.0-wmf.8) - 10https://gerrit.wikimedia.org/r/1306971 (https://phabricator.wikimedia.org/T415904) [18:00:05] andre and brennen: Time to do the MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot) deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T1800). [18:00:10] o/ Let's try again [18:00:27] nemo-yiannis: If you want to backport the two Parsoid patches, then the stage is yours. I'd try to move the train forward again once you've finished. TIA :) [18:00:38] on it [18:01:03] is there a way to do both, or should i do one by one ? [18:01:30] IIRC scap backport allows to pass several IDs [18:01:44] I have no clue though if it understands that one depends on each other, but I'd guess so [18:02:22] (and I don't know for Spiderpig, I admit. Too many ways to deploy things.) [18:02:38] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [extensions/Cite] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306970 (https://phabricator.wikimedia.org/T415904) (owner: 10WMDE-Fisch) [18:02:53] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [extensions/Cite] (wmf/1.47.0-wmf.8) - 10https://gerrit.wikimedia.org/r/1306971 (https://phabricator.wikimedia.org/T415904) (owner: 10WMDE-Fisch) [18:03:10] i put both on spiderpig, lets see how it goes [18:03:19] thanks! [18:03:23] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jgiannelos@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) (owner: 10Jgiannelos) [18:03:23] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jgiannelos@deploy1003 using scap backport" [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) (owner: 10Neriah) [18:08:28] (03Merged) 10jenkins-bot: PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) (owner: 10Neriah) [18:08:36] (03Merged) 10jenkins-bot: PageBundleParserOutputConverter: Avoid revision lookup for bogus title [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306950 (https://phabricator.wikimedia.org/T430778) (owner: 10Jgiannelos) [18:09:02] !log jgiannelos@deploy1003 Started scap sync-world: Backport for [[gerrit:1306950|PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910|PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] [18:09:05] T430778: Wikimedia\Assert\PreconditionException: Precondition failed: This Title instance does not represent a proper page, but merely a link target. - https://phabricator.wikimedia.org/T430778 [18:11:02] andre with the train not rolled out, i don't think i can test the fix [18:11:04] !log jgiannelos@deploy1003 jgiannelos, neriah: Backport for [[gerrit:1306950|PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910|PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [18:11:27] nemo-yiannis: yeah we'll find out in about 20min :) [18:11:33] hmmm [18:13:49] FIRING: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [18:13:58] !log jgiannelos@deploy1003 jgiannelos, neriah: Continuing with deployment [18:14:03] (03PS3) 10Cwhite: opensearch: add config required to authenticate curator requests [puppet] - 10https://gerrit.wikimedia.org/r/1306969 (https://phabricator.wikimedia.org/T350516) [18:17:11] (03PS42) 10CDobbins: varnish: Add CSP report-only header value [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) [18:18:17] !log jgiannelos@deploy1003 Finished scap sync-world: Backport for [[gerrit:1306950|PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910|PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s) [18:18:20] T430778: Wikimedia\Assert\PreconditionException: Precondition failed: This Title instance does not represent a proper page, but merely a link target. - https://phabricator.wikimedia.org/T430778 [18:19:01] ok done [18:19:24] nemo-yiannis, thanks so much! Alright, I'm going to move the train again [18:19:47] (03PS1) 10TrainBranchBot: group1 to 1.47.0-wmf.9 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306975 (https://phabricator.wikimedia.org/T423918) [18:19:50] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by aklapper@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306975 (https://phabricator.wikimedia.org/T423918) (owner: 10TrainBranchBot) [18:20:44] (03Merged) 10jenkins-bot: group1 to 1.47.0-wmf.9 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306975 (https://phabricator.wikimedia.org/T423918) (owner: 10TrainBranchBot) [18:21:08] (03PS1) 10Bking: wdqs: Add depool metadata to hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1306976 (https://phabricator.wikimedia.org/T327300) [18:22:11] (03CR) 10ArielGlenn: smokepy: use live mount for test files (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1302104 (https://phabricator.wikimedia.org/T424825) (owner: 10Daniel Kinzler) [18:22:39] FIRING: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch2061-production-search-codfw is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [18:27:03] !log aklapper@deploy1003 rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs T423918 [18:27:06] T423918: 1.47.0-wmf.9 deployment blockers - https://phabricator.wikimedia.org/T423918 [18:27:39] FIRING: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch2076-production-search-codfw is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [18:27:48] (03PS7) 10JHathaway: redfish: add find_accounts [software/spicerack] - 10https://gerrit.wikimedia.org/r/1303559 (https://phabricator.wikimedia.org/T426180) [18:27:54] (03CR) 10Gergő Tisza: [C:03+1] varnish: Add edge fixup for corrupt upload.wm.o urls from mobileapps [puppet] - 10https://gerrit.wikimedia.org/r/1306230 (https://phabricator.wikimedia.org/T427623) (owner: 10Krinkle) [18:28:33] nemo-yiannis: finished deployment but the testcase at https://www.wikidata.org/w/rest.php/v1/revision/2512536582/html seems to fail again [18:29:02] (or is there some caching involved?) [18:29:16] (03PS8) 10JHathaway: redfish: add find_accounts [software/spicerack] - 10https://gerrit.wikimedia.org/r/1303559 (https://phabricator.wikimedia.org/T426180) [18:33:31] andre looking [18:33:49] (03CR) 10ArielGlenn: rest-gateway: Dockerize system tests (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1297666 (https://phabricator.wikimedia.org/T424825) (owner: 10Daniel Kinzler) [18:33:49] is this on debug node? [18:34:07] nemo-yiannis, I dumped the new stacktrace into https://phabricator.wikimedia.org/T430778#12077937 [18:34:56] uggh, different error now [18:35:20] what request are you using to reproduce this ? [18:36:16] nemo-yiannis, I open https://www.wikidata.org/w/rest.php/v1/revision/2512536582/html in a browser [18:36:25] that's the testcase URL in the ticket [18:37:46] is the error rate the same as before ? [18:38:08] nemo-yiannis: No, it's much much lower [18:38:14] yeah [18:38:42] (03CR) 10JHathaway: [C:03+2] redfish: add find_accounts (031 comment) [software/spicerack] - 10https://gerrit.wikimedia.org/r/1303559 (https://phabricator.wikimedia.org/T426180) (owner: 10JHathaway) [18:39:28] nemo-yiannis, you can click the "Find reqId in Logstash" link in my last comment in https://phabricator.wikimedia.org/T430778#12077937 to see yourself [18:40:10] nemo-yiannis: I guess the question is if this justifies another rollback or not to the previous version, given the lower error rate [18:40:25] yeah, also its definitely not wikipedia view [18:40:36] its a rest api call [18:42:51] I only see one revision failing: https://logstash.wikimedia.org/goto/ead731aad0ded0d6a65bf4abe8a65c17 [18:42:57] (03Merged) 10jenkins-bot: redfish: add find_accounts [software/spicerack] - 10https://gerrit.wikimedia.org/r/1303559 (https://phabricator.wikimedia.org/T426180) (owner: 10JHathaway) [18:43:27] nemo-yiannis, ah, I realize that the automatic URL being passed sucks. Garr. basically: go to https://logstash.wikimedia.org/app/dashboards#/view/mediawiki-errors, enter "Invariant failed: Should be Parsoid content" into the search box [18:43:51] yeah, looks like the rate is low enough not to roll back for now. That's my take at least. I'm going to comment on the task to summarize. [18:43:56] ok [18:45:21] nemo-yiannis: Cool, let's keep it as-is for now. Thanks again for the fixes, and sorry for all the hassle! [18:45:24] we may have to revert a different patch of cscott's to stop this one .. or fix it. I'll ping scott on slack. [18:45:31] ah, alright [18:45:46] It is very suspicious that we only see one revision in the logs though [18:47:52] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware, 13Patch-For-Review: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12077958 (10VRiley-WMF) Physically rebooted the machine and it seemed like it came back up with no issues. Please let us know if we can do an... [18:49:15] jasmine@cumin2002 reimage (PID 3737901) is awaiting input [18:50:56] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie [18:52:36] (03CR) 10Ayounsi: [C:03+1] "thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1306976 (https://phabricator.wikimedia.org/T327300) (owner: 10Bking) [18:52:57] oh is it only one revision for all the log noise from before as well? [18:54:08] (03CR) 10BCornwall: [C:03+1] "as a nit, I think removing the enforcement on testwiki should have been another CR to preserve atomicity." [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) (owner: 10CDobbins) [18:54:50] nemo-yiannis, also see "/w/rest.php/v1/page/Q2479618/html?printable=yes&safemode=1&uselang=en" [18:55:16] and the earlier noise was for lots of different urls. [18:56:29] (03CR) 10BCornwall: varnish: Add CSP report-only header value [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) (owner: 10CDobbins) [18:57:07] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie [18:58:01] (03CR) 10BCornwall: varnish: Add CSP report-only header value (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) (owner: 10CDobbins) [18:58:29] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: Install new MPC10E-10C line cards on cr1-eqiad and cr2-eqiad slot 0. - https://phabricator.wikimedia.org/T426343#12078004 (10cmooney) Going to re-schedule this for next Wednesday July 8th at 14:00 UTC. Fairly confident things went ok... [19:05:21] (03CR) 10Bking: [C:03+2] wdqs: Add depool metadata to hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1306976 (https://phabricator.wikimedia.org/T327300) (owner: 10Bking) [19:07:16] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage [19:07:39] RESOLVED: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch2076-production-search-codfw is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [19:08:15] (03PS43) 10CDobbins: varnish: Add CSP report-only header value [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) [19:09:48] (03CR) 10JHathaway: [C:03+1] Puppet 8: Replace legacy facts [puppet] - 10https://gerrit.wikimedia.org/r/1306797 (https://phabricator.wikimedia.org/T372666) (owner: 10Ladsgroup) [19:13:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs/main-internal on k8s-dse@eqiad in state pending-install - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=eqiad&var-cluster=k8s-dse&var-namespace=wdqs - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [19:15:11] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage [19:17:20] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage [19:18:54] !log jasmine@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage [19:19:31] (03CR) 10JHathaway: spicerack: add management/config.yaml structure (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1306874 (https://phabricator.wikimedia.org/T429699) (owner: 10Elukey) [19:24:29] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage [19:25:44] (03PS44) 10CDobbins: varnish: Add CSP report-only header value [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) [19:28:02] 06SRE, 10Infrastructure Security, 06Infrastructure-Foundations, 10netops: Lumen transport eqiad codfw down July 2026 - https://phabricator.wikimedia.org/T430874 (10cmooney) 03NEW p:05Triage→03Medium [19:28:28] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage [19:30:43] (03CR) 10RLazarus: "recheck" [dns] - 10https://gerrit.wikimedia.org/r/1306902 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [19:31:38] (03CR) 10BCornwall: [C:03+1] wmnet: Update s4-master alias [dns] - 10https://gerrit.wikimedia.org/r/1306928 (https://phabricator.wikimedia.org/T430817) (owner: 10Gerrit maintenance bot) [19:34:01] (03CR) 10BCornwall: [C:03+1] "I'm guessing that opensearch-ipoid is wanted to be kept, but making a note of it here just in case." [dns] - 10https://gerrit.wikimedia.org/r/1306902 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [19:34:04] (03PS45) 10CDobbins: varnish: Add CSP report-only header value [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) [19:34:42] (03CR) 10RLazarus: [C:03+1] services_proxy: Remove ipoid listener [puppet] - 10https://gerrit.wikimedia.org/r/1306903 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [19:34:46] (03CR) 10RLazarus: [C:03+1] "Yes, that's the replacement for the one we're removing. Thanks for checking!" [dns] - 10https://gerrit.wikimedia.org/r/1306902 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [19:38:17] (03CR) 10RLazarus: "This needs to go to lvs_setup first, and then service_setup, right? (https://wikitech.wikimedia.org/wiki/LVS#Remove_network_probes_/_monit" [puppet] - 10https://gerrit.wikimedia.org/r/1306899 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [19:38:59] (03CR) 10RLazarus: [C:03+1] service: Remove ipoid service [puppet] - 10https://gerrit.wikimedia.org/r/1306900 (https://phabricator.wikimedia.org/T416623) (owner: 10Clément Goubert) [19:39:13] (03PS46) 10CDobbins: varnish: Add CSP report-only header value [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) [19:42:11] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie [19:42:40] FIRING: SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:43:42] !log jasmine@cumin2002 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" [19:44:13] !log jasmine@cumin2002 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002" [19:44:15] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie [19:44:18] (03CR) 10CDobbins: "text: 0 tests failed, 0 tests skipped, 40 tests passed" [puppet] - 10https://gerrit.wikimedia.org/r/1297217 (https://phabricator.wikimedia.org/T117618) (owner: 10CDobbins) [19:44:27] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware, 13Patch-For-Review: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12078109 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-ctrl1005.eqiad.wmnet with... [19:46:44] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie [19:59:14] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor I � Unicode. All rise for UTC late backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T2000). [20:00:05] nemo-yiannis: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:03:30] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware, 13Patch-For-Review: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12078121 (10jasmine_) >>! In T418920#12040455, @MLechvien-WMF wrote: > @jasmine_ can you create the decommissioning task for wikikube-ctrl100... [20:04:08] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware, 13Patch-For-Review: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12078122 (10jasmine_) >>! In T418920#12077958, @VRiley-WMF wrote: > Physically rebooted the machine and it seemed like it came back up with n... [20:06:24] (03CR) 10JHathaway: "@taavi@wikimedia.org let me know how you would like to break this commit up. Happy to do per module patch sets or you are welcome to break" [puppet] - 10https://gerrit.wikimedia.org/r/1305982 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:06:33] (03PS1) 10Kamila Součková: mediawiki/php: fix apt component for bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1306979 (https://phabricator.wikimedia.org/T423714) [20:11:20] (03PS2) 10Kamila Součková: mediawiki/php: fix apt component for bookworm [puppet] - 10https://gerrit.wikimedia.org/r/1306979 (https://phabricator.wikimedia.org/T423714) [20:11:39] (03CR) 10RLazarus: [C:03+1] Remove ipoid chart and service definitions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1306784 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [20:11:47] (03CR) 10RLazarus: [C:03+1] deployment_server: absent ipoid kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1306779 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [20:11:55] (03CR) 10Kamila Součková: "This should be a no-op on bullseye, and a fix on bookworm. I don't think we have any bookworm hosts to pcc against though (I am testing th" [puppet] - 10https://gerrit.wikimedia.org/r/1306979 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [20:11:58] (03CR) 10RLazarus: [C:03+1] deployment_server: remove ipoid users [puppet] - 10https://gerrit.wikimedia.org/r/1306782 (https://phabricator.wikimedia.org/T416623) (owner: 10Dreamy Jazz) [20:12:02] (03CR) 10Kamila Součková: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1306979 (https://phabricator.wikimedia.org/T423714) (owner: 10Kamila Součková) [20:19:40] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage [20:24:40] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage [20:28:34] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie [20:30:41] PROBLEM - ganeti-noded running on ganeti1029 is CRITICAL: PROCS CRITICAL: 3 processes with UID = 0 (root), command name ganeti-noded https://wikitech.wikimedia.org/wiki/Ganeti [20:31:41] RECOVERY - ganeti-noded running on ganeti1029 is OK: PROCS OK: 2 processes with UID = 0 (root), command name ganeti-noded https://wikitech.wikimedia.org/wiki/Ganeti [20:39:38] (03CR) 10JHathaway: "@brouberol@wikimedia.org let me know how you would like me to dice this commit up, to make it more digestible, happy to do per module patc" [puppet] - 10https://gerrit.wikimedia.org/r/1305987 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:41:19] (03CR) 10JHathaway: "@ssingh@wikimedia.org let me know how you want me to break up this commit, happy to do a patch per module, or you can break it up how you " [puppet] - 10https://gerrit.wikimedia.org/r/1305984 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:41:58] 2 !incidents [20:45:54] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage [20:46:07] (03CR) 10JHathaway: "@matthieulec@wikimedia.org happy to break this up in any way that is more digestible for the team, e.g. per module, of course you folks ar" [puppet] - 10https://gerrit.wikimedia.org/r/1305986 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [20:50:18] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage [20:50:46] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie [20:55:08] FIRING: KubernetesAPILatency: High Kubernetes API latency (LIST pods) on k8s-mlserve@eqiad - https://wikitech.wikimedia.org/wiki/Kubernetes - https://grafana.wikimedia.org/d/ddNd-sLnk/kubernetes-api-details?var-site=eqiad&var-cluster=k8s-mlserve&var-latency_percentile=0.95&var-verb=LIST - https://alerts.wikimedia.org/?q=alertname%3DKubernetesAPILatency [20:55:13] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - ml-ctrl_6443: Servers ml-serve-ctrl1001.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [20:56:13] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [21:00:05] Deploy window Wikifunctions Services UTC Late (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T2100) [21:01:08] (03CR) 10JHathaway: "@hnowlan@wikimedia.org let me know how you want to break this up for the team, happy to redo as per module patches, or let you folks break" [puppet] - 10https://gerrit.wikimedia.org/r/1305989 (https://phabricator.wikimedia.org/T372666) (owner: 10JHathaway) [21:12:37] (03PS1) 10Jdlrobson: Remove unused user skin preference config [skins/Vector] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306983 (https://phabricator.wikimedia.org/T358273) [21:15:23] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie [21:19:48] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie [21:34:21] !log bking@cumin2003 START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm [21:34:39] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host wcqs2003 [21:35:06] !log bking@cumin2003 START - Cookbook sre.dns.netbox [21:36:57] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage [21:41:07] bking@cumin2003 reimage (PID 2183740) is awaiting input [21:42:31] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" [21:42:35] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003" [21:42:36] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [21:42:36] !log bking@cumin2003 START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:42:40] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:42:41] !log bking@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003 [21:42:53] !log bking@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003 [21:42:53] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003 [21:43:11] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage [21:43:54] (03PS1) 10Subramanya Sastry: Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306985 (https://phabricator.wikimedia.org/T430194) [21:46:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps2011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:50:35] !log jasmine@cumin2002 START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie [21:50:51] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware, 13Patch-For-Review: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12078322 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jasmine@cumin2002 for host wikikube-ctrl1005.eqiad.wmnet... [22:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260701T2200) [22:01:57] !log bking@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage [22:03:36] !log jasmine@cumin2002 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage [22:09:13] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage [22:10:19] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie [22:13:00] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage [22:17:25] FIRING: [2x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:22:25] FIRING: [2x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:29:06] !log jasmine@cumin2002 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie [22:29:21] 06SRE, 06ServiceOps new, 10ServiceOps-Upgrades-Hardware, 13Patch-For-Review: wikikube-ctrl100[56] implementation tracking - https://phabricator.wikimedia.org/T418920#12078381 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jasmine@cumin2002 for host wikikube-ctrl1005.eqiad.wmnet with... [22:36:23] (03PS1) 10BPirkle: REST: use RestExternalModules config variable [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306988 (https://phabricator.wikimedia.org/T428375) [22:37:07] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm [23:06:25] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:12:45] (03CR) 10C. Scott Ananian: [C:03+1] "Thanks!" [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306910 (https://phabricator.wikimedia.org/T430778) (owner: 10Neriah) [23:30:03] (03Abandoned) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1306795 (owner: 10TrainBranchBot) [23:30:55] (03PS3) 10Cwhite: prometheus: refactor prometheus-es-exporter to use config file [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) [23:31:35] (03CR) 10CI reject: [V:04-1] prometheus: refactor prometheus-es-exporter to use config file [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:32:38] (03PS4) 10Cwhite: prometheus: refactor prometheus-es-exporter to use config file [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) [23:36:09] (03CR) 10Cwhite: "+1!" [puppet] - 10https://gerrit.wikimedia.org/r/1305718 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:42:29] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1306995 [23:42:29] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1306995 (owner: 10TrainBranchBot) [23:44:39] FIRING: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch2084-production-search-codfw is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [23:45:18] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1306985 (https://phabricator.wikimedia.org/T430194) (owner: 10Subramanya Sastry) [23:47:55] (03PS1) 10C. Scott Ananian: [parser] When expanding an extension tag with a title, use a new frame [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306996 (https://phabricator.wikimedia.org/T430344) [23:48:13] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Thursday, July 02 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [core] (wmf/1.47.0-wmf.9) - 10https://gerrit.wikimedia.org/r/1306996 (https://phabricator.wikimedia.org/T430344) (owner: 10C. Scott Ananian) [23:49:39] RESOLVED: CirrusSearchNodeIndexingNotIncreasing: Elasticsearch instance cirrussearch2084-production-search-codfw is not indexing - https://wikitech.wikimedia.org/wiki/Search/Elasticsearch_Administration#Indexing_hung_and_not_making_progress - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [23:50:31] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1306995 (owner: 10TrainBranchBot) [23:50:50] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2106 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:50] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2093 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:50] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2086 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:50] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2068 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:54] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2090 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:54] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2089 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:56] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2111 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:56] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2113 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:56] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2095 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:56] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2098 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:56] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2105 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:57] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2065 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:57] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch2070 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [23:50:59] !log cscott@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [23:51:28] !log cscott@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [23:51:30] !log cscott@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [23:51:51] ^ looking into this cirrussearch noise [23:52:00] !log cscott@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [23:58:59] !log T429844 [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election [23:59:02] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [23:59:03] T429844: Migrate production OpenSearch clusters from 1.x-2.x - CODFW - https://phabricator.wikimedia.org/T429844