[00:35:07] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [01:10:36] (03PS1) 10TrainBranchBot: Branch commit for wmf/1.47.0-wmf.19 [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337665 (https://phabricator.wikimedia.org/T430838) [01:10:39] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/1.47.0-wmf.19 [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337665 (https://phabricator.wikimedia.org/T430838) (owner: 10TrainBranchBot) [01:11:11] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1337666 [01:11:11] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1337666 (owner: 10TrainBranchBot) [01:19:26] (03Merged) 10jenkins-bot: Branch commit for wmf/1.47.0-wmf.19 [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337665 (https://phabricator.wikimedia.org/T430838) (owner: 10TrainBranchBot) [01:20:31] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1337666 (owner: 10TrainBranchBot) [01:26:23] PROBLEM - jenkins_service_running on releases1003 is CRITICAL: PROCS CRITICAL: 3 processes with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [01:27:23] RECOVERY - jenkins_service_running on releases1003 is OK: PROCS OK: 1 process with regex args .*/bin/java .*-jar /usr/share/java/jenkins.war https://wikitech.wikimedia.org/wiki/Jenkins [01:32:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:00:04] Deploy window Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T0200) [02:00:04] Deploy window Automatic deployment of MediaWiki to pretrain wikis - see mw:Pretrain (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T0200) [02:00:38] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [02:01:46] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:09:27] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 41s) [02:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:49:53] (03PS4) 10Tim Starling: Enable Produnto on pilot wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324966 (https://phabricator.wikimedia.org/T421436) [03:00:05] Deploy window Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T0300) [03:01:54] (03PS1) 10TrainBranchBot: testwikis to 1.47.0-wmf.19 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337668 (https://phabricator.wikimedia.org/T430838) [03:01:57] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by mwpresync@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337668 (https://phabricator.wikimedia.org/T430838) (owner: 10TrainBranchBot) [03:03:00] (03Merged) 10jenkins-bot: testwikis to 1.47.0-wmf.19 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337668 (https://phabricator.wikimedia.org/T430838) (owner: 10TrainBranchBot) [03:03:21] !log mwpresync@deploy1003 Started scap sync-world: testwikis to 1.47.0-wmf.19 refs T430838 [03:03:24] T430838: 1.47.0-wmf.19 deployment blockers - https://phabricator.wikimedia.org/T430838 [03:20:21] PROBLEM - Improperly owned -0:0- files in /srv/mediawiki-staging on deploy2002 is CRITICAL: Improperly owned (0:0) files in /srv/mediawiki-staging https://wikitech.wikimedia.org/wiki/Monitoring/bad_directory_owner [03:30:15] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker2005 has a BGP session which is not in the 'established' state. [03:30:21] RECOVERY - Improperly owned -0:0- files in /srv/mediawiki-staging on deploy2002 is OK: Files ownership is ok. https://wikitech.wikimedia.org/wiki/Monitoring/bad_directory_owner [03:39:51] !log mwpresync@deploy1003 Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs T430838 (duration: 36m 30s) [03:39:54] T430838: 1.47.0-wmf.19 deployment blockers - https://phabricator.wikimedia.org/T430838 [04:00:05] Deploy window Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T0400) [04:02:30] !log mwpresync@deploy1003 Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s) [04:35:07] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [04:55:03] (03CR) 10Ayounsi: [C:03+2] Add support for short hostname in query [software/homer] - 10https://gerrit.wikimedia.org/r/1337571 (owner: 10Ayounsi) [04:56:37] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 08 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1335706 (https://phabricator.wikimedia.org/T436426) (owner: 10Hamish) [04:59:45] PROBLEM - zuul_service_running on contint1002 is CRITICAL: PROCS CRITICAL: 4 processes with regex args bin/zuul-server https://www.mediawiki.org/wiki/Continuous_integration/Zuul [05:00:45] RECOVERY - zuul_service_running on contint1002 is OK: PROCS OK: 2 processes with regex args bin/zuul-server https://www.mediawiki.org/wiki/Continuous_integration/Zuul [05:01:08] (03Merged) 10jenkins-bot: Add support for short hostname in query [software/homer] - 10https://gerrit.wikimedia.org/r/1337571 (owner: 10Ayounsi) [05:07:28] !log Add grafana-plugins 0.10 to bookworm-wikimedia - T436056 [05:07:31] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:07:32] T436056: Move grafana-plugins from Gerrit to Gitlab to ease building the package - https://phabricator.wikimedia.org/T436056 [05:21:16] (03CR) 10Ryan Kemper: [C:04-1] "found an issue with my original gitlab patch and therefore this image (a delayed edit can outrank a subsequent delete during dedupe, meani" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334719 (https://phabricator.wikimedia.org/T436736) (owner: 10DCausse) [05:32:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:51:36] (03PS1) 10Ayounsi: Capirca: optimize fetching the latest completed run [software/homer] - 10https://gerrit.wikimedia.org/r/1337811 [05:55:20] (03PS2) 10Ayounsi: Capirca: optimize fetching the latest completed run [software/homer] - 10https://gerrit.wikimedia.org/r/1337811 [05:55:20] (03PS3) 10Ayounsi: capirca: python 3.12 deprecates datetime.utcnow() [software/homer] - 10https://gerrit.wikimedia.org/r/1212243 (owner: 10E75ti) [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T0600) [06:00:05] marostegui, cezmunsta, and federico3: #bothumor My software never has bugs. It just develops random features. Rise for Primary database switchover. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T0600). [06:00:39] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [06:13:45] (03CR) 10Ayounsi: [C:03+2] capirca: python 3.12 deprecates datetime.utcnow() [software/homer] - 10https://gerrit.wikimedia.org/r/1212243 (owner: 10E75ti) [06:26:17] (03CR) 10Mszwarc: [C:03+1] Split out edit and block-based filters from activity filters [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337610 (https://phabricator.wikimedia.org/T436508) (owner: 10STran) [06:38:47] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - linkrecommendation-external_4006: Servers wikikube-worker1291.eqiad.wmnet, wikikube-worker1322.eqiad.wmnet, wikikube-worker1042.eqiad.wmnet, wikikube-worker1268.eqiad.wmnet, wikikube-worker1067.eqiad.wmnet, wikikube-worker1148.eqiad.wmnet, wikikube-worker1259.eqiad.wmnet, wikikube-worker1155.eqiad.wmnet, wikikube-worker1345.eqiad.wmnet, wikikube-work [06:38:47] qiad.wmnet, wikikube-worker1036.eqiad.wmnet, wikikube-worker1380.eqiad.wmnet, wikikube-worker1310.eqiad.wmnet, wikikube-worker1049.eqiad.wmnet, wikikube-worker1371.eqiad.wmnet, wikikube-worker1260.eqiad.wmnet, wikikube-worker1094.eqiad.wmnet, wikikube-worker1132.eqiad.wmnet, wikikube-worker1016.eqiad.wmnet, wikikube-worker1071.eqiad.wmnet, wikikube-worker1279.eqiad.wmnet, wikikube-worker1157.eqiad.wmnet, wikikube-worker1358.eqiad.wmnet, w [06:38:47] worker1072.eqiad.wmnet, wikikube-worker1307.eqiad.wmnet, wikikube-worker1287.eqiad.wmnet, wikikube-worker1244.eqiad.wmnet, wikikube-worker1278.eqiad.wmnet, wikikube-worker1377.eqiad.wmn https://wikitech.wikimedia.org/wiki/PyBal [06:38:47] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - linkrecommendation-external_4006: Servers wikikube-worker1144.eqiad.wmnet, wikikube-worker1042.eqiad.wmnet, wikikube-worker1268.eqiad.wmnet, wikikube-worker1383.eqiad.wmnet, wikikube-worker1067.eqiad.wmnet, wikikube-worker1160.eqiad.wmnet, wikikube-worker1116.eqiad.wmnet, wikikube-worker1050.eqiad.wmnet, wikikube-worker1036.eqiad.wmnet, wikikube-work [06:38:47] qiad.wmnet, wikikube-worker1079.eqiad.wmnet, wikikube-worker1132.eqiad.wmnet, wikikube-worker1247.eqiad.wmnet, wikikube-worker1273.eqiad.wmnet, wikikube-worker1260.eqiad.wmnet, wikikube-worker1279.eqiad.wmnet, wikikube-worker1157.eqiad.wmnet, wikikube-worker1338.eqiad.wmnet, wikikube-worker1313.eqiad.wmnet, wikikube-worker1287.eqiad.wmnet, wikikube-worker1270.eqiad.wmnet, wikikube-worker1244.eqiad.wmnet, wikikube-worker1037.eqiad.wmnet, w [06:38:48] worker1278.eqiad.wmnet, wikikube-worker1340.eqiad.wmnet, wikikube-worker1377.eqiad.wmnet, wikikube-worker1119.eqiad.wmnet, wikikube-worker1066.eqiad.wmnet, wikikube-worker1336.eqiad.wmn https://wikitech.wikimedia.org/wiki/PyBal [06:49:47] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [06:49:47] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [06:50:20] (03PS2) 10Muehlenhoff: Rename puppetmaster::puppetdb to puppetdb [puppet] - 10https://gerrit.wikimedia.org/r/1331470 [06:56:11] (03PS2) 10Muehlenhoff: Remove puppet/config-master records pointint to puppetmaster2001 [dns] - 10https://gerrit.wikimedia.org/r/1237463 (https://phabricator.wikimedia.org/T416606) [07:00:04] hi :) [07:00:05] Amir1, urbanecm, and awight: I, the Bot under the Fountain, call upon thee, The Deployer, to do UTC morning backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T0700). [07:00:05] hamishcz: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:06:17] (03CR) 10Filippo Giunchedi: [C:03+1] Remove puppet/config-master records pointint to puppetmaster2001 [dns] - 10https://gerrit.wikimedia.org/r/1237463 (https://phabricator.wikimedia.org/T416606) (owner: 10Muehlenhoff) [07:10:17] Is anybody around pls? [07:11:55] (03CR) 10Muehlenhoff: [C:03+2] Remove puppet/config-master records pointint to puppetmaster2001 [dns] - 10https://gerrit.wikimedia.org/r/1237463 (https://phabricator.wikimedia.org/T416606) (owner: 10Muehlenhoff) [07:11:59] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 T404715', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json [07:12:03] T404715: Setup x4 section - https://phabricator.wikimedia.org/T404715 [07:12:17] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 T404715', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json [07:12:34] !log jmm@dns1004 START - running authdns-update [07:13:09] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 T404715', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json [07:14:43] !log jmm@dns1004 END - running authdns-update [07:18:14] !log mvernon@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet [07:22:29] !log jmm@cumin1004 START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org [07:22:30] !log jmm@cumin1004 START - Cookbook sre.dns.netbox [07:22:44] (03CR) 10Brouberol: [C:03+1] wdqs: extend envoy rq_time histogram buckets [deployment-charts] - 10https://gerrit.wikimedia.org/r/1336094 (https://phabricator.wikimedia.org/T428631) (owner: 10Gmodena) [07:26:18] !log jmm@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004" [07:26:21] !log jmm@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004" [07:26:21] !log jmm@cumin1004 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [07:26:21] !log jmm@cumin1004 START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors [07:26:24] !log jmm@cumin1004 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors [07:27:04] !log jmm@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004" [07:27:07] !log jmm@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004" [07:27:12] !log jmm@cumin1004 START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie [07:27:25] 06SRE, 06Infrastructure-Foundations, 07LDAP, 13Patch-For-Review: Migrate the r/w LDAP servers to Trixie and MDB storage - https://phabricator.wikimedia.org/T331699#12296305 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin1004 for host ldap-replica1006.wikimedia.org with... [07:28:40] !log mvernon@cumin2003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet [07:28:53] !log mvernon@cumin2003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet [07:29:21] !log mvernon@cumin2003 START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet [07:30:05] !log Add grafana-plugins 0.15 to bookworm-wikimedia - T436056 [07:30:07] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:30:08] T436056: Move grafana-plugins from Gerrit to Gitlab to ease building the package - https://phabricator.wikimedia.org/T436056 [07:30:15] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker2005 has a BGP session which is not in the 'established' state. [07:32:40] FIRING: [3x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:35:03] (03PS1) 10Muehlenhoff: Remove puppet/config-master records pointing to puppetmaster1001 [dns] - 10https://gerrit.wikimedia.org/r/1337819 (https://phabricator.wikimedia.org/T416606) [07:36:36] Looks like no one back to me lol [07:37:40] FIRING: [4x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:37:42] !log jmm@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage [07:38:30] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet [07:38:32] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet [07:38:59] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [07:40:45] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [07:41:17] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm [07:42:40] RESOLVED: [4x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:42:53] (03CR) 10Volans: [C:03+1] "LGTM, I've tested the options via netbox APIs in my browser and they seems to be working as expected." [software/homer] - 10https://gerrit.wikimedia.org/r/1337811 (owner: 10Ayounsi) [07:43:16] !log jmm@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage [07:44:45] (03CR) 10Ayounsi: [C:03+2] Capirca: optimize fetching the latest completed run [software/homer] - 10https://gerrit.wikimedia.org/r/1337811 (owner: 10Ayounsi) [07:44:50] (03PS1) 10Ayounsi: CHANGELOG: add changelog for release v0.11.3 [software/homer] - 10https://gerrit.wikimedia.org/r/1337821 [07:45:25] FIRING: [3x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:46:30] (03CR) 10Majavah: [C:03+1] Rename puppetmaster::puppetdb to puppetdb [puppet] - 10https://gerrit.wikimedia.org/r/1331470 (owner: 10Muehlenhoff) [07:46:46] (03PS1) 10Muehlenhoff: base::kernel: blacklist RDS modules [puppet] - 10https://gerrit.wikimedia.org/r/1337822 [07:50:05] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337822 (owner: 10Muehlenhoff) [07:50:25] RESOLVED: [4x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:50:29] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet [07:53:34] mvernon@cumin1003 upgrade-firmware (PID 3692674) is awaiting input [07:53:58] (03CR) 10Volans: [C:03+1] "LGTM" [software/homer] - 10https://gerrit.wikimedia.org/r/1337821 (owner: 10Ayounsi) [07:54:24] (03Merged) 10jenkins-bot: Capirca: optimize fetching the latest completed run [software/homer] - 10https://gerrit.wikimedia.org/r/1337811 (owner: 10Ayounsi) [07:54:25] (03Merged) 10jenkins-bot: capirca: python 3.12 deprecates datetime.utcnow() [software/homer] - 10https://gerrit.wikimedia.org/r/1212243 (owner: 10E75ti) [07:54:52] (03PS2) 10Muehlenhoff: base::kernel: blacklist RDS modules [puppet] - 10https://gerrit.wikimedia.org/r/1337822 [07:55:03] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499#12296343 (10elukey) Nothing to report after the switch, the metrics look good on the Registry's dashboard and nobody came up with... [07:55:11] (03CR) 10Ayounsi: [C:03+2] CHANGELOG: add changelog for release v0.11.3 [software/homer] - 10https://gerrit.wikimedia.org/r/1337821 (owner: 10Ayounsi) [07:57:23] (03CR) 10Ayounsi: [C:03+1] base::kernel: blacklist RDS modules [puppet] - 10https://gerrit.wikimedia.org/r/1337822 (owner: 10Muehlenhoff) [07:57:30] (03CR) 10Volans: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1337822 (owner: 10Muehlenhoff) [07:58:25] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage [07:58:54] !log jmm@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie [07:58:54] !log jmm@cumin1004 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org [07:59:06] 06SRE, 06Infrastructure-Foundations, 07LDAP, 13Patch-For-Review: Migrate the r/w LDAP servers to Trixie and MDB storage - https://phabricator.wikimedia.org/T331699#12296350 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin1004 for host ldap-replica1006.wikimedia.org with OS t... [08:01:32] (03Merged) 10jenkins-bot: CHANGELOG: add changelog for release v0.11.3 [software/homer] - 10https://gerrit.wikimedia.org/r/1337821 (owner: 10Ayounsi) [08:01:42] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12296356 (10Geertivp) Problem is that other admins seem to have lost their admin right as well... how is that possible? Strange enough we are on the e-mail list of administrators wikimediabe-l-... [08:03:00] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 08 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#depl" [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337610 (https://phabricator.wikimedia.org/T436508) (owner: 10STran) [08:03:43] (03CR) 10Muehlenhoff: [C:03+2] base::kernel: blacklist RDS modules [puppet] - 10https://gerrit.wikimedia.org/r/1337822 (owner: 10Muehlenhoff) [08:03:53] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage [08:03:56] (03CR) 10Elukey: [C:03+1] base::kernel: blacklist RDS modules (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1337822 (owner: 10Muehlenhoff) [08:05:39] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Create new S3 backends for the Docker Registry service - https://phabricator.wikimedia.org/T427175#12296365 (10elukey) The Registry is now completely on S3, swift is not used anymore. I am going to wait more days before calling it over but so far it... [08:08:46] mvernon@cumin1003 upgrade-firmware (PID 3692674) is awaiting input [08:11:28] (03PS1) 10Ayounsi: Release 0.11.3 [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1337848 [08:15:13] (03PS2) 10Ayounsi: Release 0.11.3 [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1337848 [08:18:25] FIRING: [3x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:19:05] (03PS3) 10Ayounsi: Release 0.11.3 [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1337848 [08:19:47] (03CR) 10Volans: [C:03+1] "LGTM, thx" [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1337848 (owner: 10Ayounsi) [08:20:44] (03CR) 10Ayounsi: [C:03+2] Release 0.11.3 [software/homer/deploy] - 10https://gerrit.wikimedia.org/r/1337848 (owner: 10Ayounsi) [08:23:25] FIRING: [4x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:25:26] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm [08:26:11] !log mvernon@cumin1003 START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet [08:27:08] (03CR) 10Cathal Mooney: [C:03+1] Capirca: optimize fetching the latest completed run [software/homer] - 10https://gerrit.wikimedia.org/r/1337811 (owner: 10Ayounsi) [08:28:25] RESOLVED: [4x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:29:13] !log ayounsi@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003 [08:30:02] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003 [08:30:45] (03PS1) 10Jelto: service::catalog: Set ipip for echostore, eventstreams, kartotherian [puppet] - 10https://gerrit.wikimedia.org/r/1337851 (https://phabricator.wikimedia.org/T420436) [08:30:48] (03PS1) 10Jelto: service::catalog: Set ipip_encapsulation for echostore, eventstreams, kartotherian [puppet] - 10https://gerrit.wikimedia.org/r/1337852 (https://phabricator.wikimedia.org/T420436) [08:31:26] (03PS1) 10Dpogorzelski: lw-studio: bump pgsql version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337853 [08:31:38] (03PS2) 10Jelto: service::catalog: Set ipip for echostore, eventstreams, kartotherian [puppet] - 10https://gerrit.wikimedia.org/r/1337852 (https://phabricator.wikimedia.org/T420436) [08:31:45] (03CR) 10CI reject: [V:04-1] service::catalog: Set ipip for echostore, eventstreams, kartotherian [puppet] - 10https://gerrit.wikimedia.org/r/1337852 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [08:32:31] (03PS2) 10Jelto: service::catalog: Set ipip for echostore, eventstreams, kartotherian codfw [puppet] - 10https://gerrit.wikimedia.org/r/1337851 (https://phabricator.wikimedia.org/T420436) [08:32:45] (03PS3) 10Jelto: service::catalog: Set ipip for echostore, eventstreams, kartotherian eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1337852 (https://phabricator.wikimedia.org/T420436) [08:33:25] FIRING: [8x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:34:42] !log ayounsi@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003 [08:35:07] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [08:35:32] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003 [08:36:46] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet [08:36:48] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet [08:37:49] !log ayounsi@cumin1003 START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003 [08:37:53] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [08:38:25] RESOLVED: [8x] ProbeDown: Service aqs2006-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:38:38] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003 [08:39:39] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [08:41:28] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm [08:42:48] (03PS1) 10Marostegui: mariadb: Prepare x4 masters [puppet] - 10https://gerrit.wikimedia.org/r/1337854 (https://phabricator.wikimedia.org/T404715) [08:43:45] (03CR) 10Marostegui: [C:03+2] mariadb: Prepare x4 masters [puppet] - 10https://gerrit.wikimedia.org/r/1337854 (https://phabricator.wikimedia.org/T404715) (owner: 10Marostegui) [08:43:57] !log marostegui@cumin1003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split [08:45:40] FIRING: [3x] ProbeDown: Service aqs2007-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:48:55] (03CR) 10Brouberol: "Did you have time to take a look during the last meeting?" [puppet] - 10https://gerrit.wikimedia.org/r/1319853 (https://phabricator.wikimedia.org/T421952) (owner: 10Brouberol) [08:50:40] RESOLVED: [4x] ProbeDown: Service aqs2007-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:52:27] 06SRE, 10Wikimedia-Mailing-lists: Mailing list logging in/ownership issue - https://phabricator.wikimedia.org/T436830#12296482 (10Aklapper) > how is that possible? You may receive mail to a different address via forwards than the one which you have in mind, for example. Impossible to tell without specific ema... [08:52:54] (03CR) 10Brouberol: [C:03+1] lw-studio: bump pgsql version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337853 (owner: 10Dpogorzelski) [08:53:48] (03CR) 10Dpogorzelski: [C:03+2] lw-studio: bump pgsql version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337853 (owner: 10Dpogorzelski) [08:54:46] (03CR) 10Dpogorzelski: [V:03+2 C:03+2] lw-studio: bump pgsql version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337853 (owner: 10Dpogorzelski) [08:55:09] (03CR) 10Muehlenhoff: "Yes, we discussed it yesterday, but I hadn't gotten around yet to follow up: The change per se was deemed perfectly fine, but instead of t" [puppet] - 10https://gerrit.wikimedia.org/r/1319853 (https://phabricator.wikimedia.org/T421952) (owner: 10Brouberol) [08:55:09] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync [08:55:12] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync [08:57:03] (03PS1) 10Zabe: Use virtual domain in NameTableStore for collation [core] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337856 (https://phabricator.wikimedia.org/T405812) [08:57:17] (03PS1) 10Zabe: Use virtual domain in NameTableStore for collation [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337857 (https://phabricator.wikimedia.org/T405812) [08:58:33] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet [09:00:12] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage [09:01:37] I'm around too [09:01:41] o/ [09:01:59] !log Starting x4 split from s4, RO time on commons needed T404715 [09:02:01] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:02:02] T404715: Setup x4 section - https://phabricator.wikimedia.org/T404715 [09:02:10] o/ [09:02:10] 06SRE, 10SRE-Access-Requests: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271 (10dkertesz) 03NEW [09:02:29] !log marostegui@cumin1003 dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance T404715', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json [09:02:35] starting the split [09:03:19] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage [09:03:57] for the record: https://phabricator.wikimedia.org/P96390 [09:04:16] replication disconnected from x4 [09:04:43] starting to change dbctl with new masters [09:04:54] 06SRE, 10SRE-Access-Requests: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12296549 (10dkertesz) [09:05:18] !log marostegui@cumin1003 dbctl commit (dc=all): 'Set x4 masters T404715', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json [09:05:30] 06SRE, 10SRE-Access-Requests: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12296552 (10Fabfur) a:03Fabfur [09:06:15] 06SRE, 10SRE-Access-Requests: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12296556 (10Fabfur) [09:06:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?panelId=18&fullscreen&orgId=1&var-datasource=codfw%20prometheus/ops - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [09:06:37] there's something going with dbctl [09:06:42] when removing old hosts from x4 db1160 [09:07:18] fixed [09:07:50] !log marostegui@cumin1003 dbctl commit (dc=all): 'Remove old s4 masters from x4 T404715', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json [09:07:53] T404715: Setup x4 section - https://phabricator.wikimedia.org/T404715 [09:07:54] Ok I am ready to go RW [09:07:57] zabe Amir1 ^ [09:08:30] checking [09:08:37] can i go RW? [09:08:37] In orchestrator, it still shows the codfw hosts as part of s4 instead of x4. Is that just a display bug? [09:08:42] yep [09:08:53] (03PS1) 10Muehlenhoff: thumbor-plugins: Rebuild against latest package versions in Trixie [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337861 [09:08:54] https://orchestrator.wikimedia.org/web/cluster/x4 [09:08:58] this is only eqiad [09:09:13] okay then [09:09:23] mmm [09:09:23] ready from my pov [09:09:30] if that is not an issue [09:09:33] looks good to me too [09:09:34] ok let me fix x4 codfw orchestrator to make fully sure [09:10:40] thanks [09:10:42] https://orchestrator.wikimedia.org/web/cluster/x4 [09:10:45] looks happy [09:10:47] going to go RW [09:10:56] yay [09:11:12] mvernon@cumin1003 upgrade-firmware (PID 3731154) is awaiting input [09:11:15] FIRING: [8x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [09:11:22] !log marostegui@cumin1003 dbctl commit (dc=all): 'Make x4 and s4 RW again T404715', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json [09:11:27] Done in dbctl, changing masters in mariadb [09:11:55] all done [09:11:57] we should be live [09:12:03] can you guys check please? [09:12:10] should Commons fail like this when on RO? https://usercontent.irccloud-cdn.com/file/LZ71kBcn/image.png [09:12:24] i'd expect to be able to still...read stuff [09:12:31] yeah [09:12:36] not nice [09:12:39] no, but not related to the split at least [09:13:08] I see recentchanges in commons moving [09:13:33] zabe: can you check if page changes is being reflected on x4 too? [09:13:41] (03CR) 10Muehlenhoff: [C:03+2] Additional necessary settings for the openldap::rw_mdb role [puppet] - 10https://gerrit.wikimedia.org/r/1335860 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [09:13:42] yep, will do [09:13:46] thanks [09:14:04] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploy v1.5.4 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337862 (https://phabricator.wikimedia.org/T421813) [09:14:48] So far logstash looks fine [09:15:24] (03PS1) 10Daniel Kertesz: admin: move dkertesz from ldap_only_users to users [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) [09:15:25] CentralAuth is at fault... I'll fill a ticket for the broken read only, sounds like something that'd affect other wikis too [09:15:50] urbanecm: thanks [09:16:15] RESOLVED: [8x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [09:18:06] page writes seem to be duplicated as intended [09:18:07] FIRING: [2x] ProbeDown: Service aqs2007-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:20:40] (03PS1) 10Marostegui: mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) [09:21:25] filled T437273, not sure if i should cross link it from somewhere [09:21:25] T437273: MediaWiki is not capable to operate in read only mode - https://phabricator.wikimedia.org/T437273 [09:21:27] (03CR) 10Muehlenhoff: [C:03+1] "LGTM (haven't validated the SSH keys, it's sufficient if the onboarding buddy does that)" [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) (owner: 10Daniel Kertesz) [09:22:27] !log mvernon@cumin1003 START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet [09:22:28] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm [09:22:56] 10SRE-SLO, 10Citoid, 06Editing-team, 07Sustainability (Incident Followup): Improve monitoring in citoid so that url-downloader failures are detected - https://phabricator.wikimedia.org/T381372#12296619 (10MLechvien-WMF) In the meantime the observability of citoid has been improved in T430629 . > @CDanis a... [09:23:07] FIRING: [4x] ProbeDown: Service aqs2007-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:23:08] (03PS1) 10Ayounsi: Reclaim DNS v6 PTR for decom GRE tunnels [dns] - 10https://gerrit.wikimedia.org/r/1337865 [09:23:34] !log installing rsync security updates [09:23:35] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:23:57] (03PS2) 10Brouberol: admin: allow analytics-admin members to impersonate any analytics user [puppet] - 10https://gerrit.wikimedia.org/r/1319853 (https://phabricator.wikimedia.org/T421952) [09:23:59] zabe: everything looks good from your side? [09:24:08] (03CR) 10CI reject: [V:04-1] Reclaim DNS v6 PTR for decom GRE tunnels [dns] - 10https://gerrit.wikimedia.org/r/1337865 (owner: 10Ayounsi) [09:24:20] Yes. I am a bit surprised tbh. [09:24:33] Thanks for your help. :) [09:24:34] (03CR) 10Elukey: [C:03+1] mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [09:24:46] zabe: hahah [09:24:59] (03CR) 10Marostegui: "recheck" [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [09:25:29] (03CR) 10Zabe: [C:03+1] mysql.py: Add x4 (031 comment) [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [09:25:46] jouncebot: nowandnext [09:25:46] No deployments scheduled for the next 0 hour(s) and 34 minute(s) [09:25:46] In 0 hour(s) and 34 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1000) [09:25:56] (03CR) 10Zabe: [C:03+2] Use virtual domain in NameTableStore for collation [core] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337856 (https://phabricator.wikimedia.org/T405812) (owner: 10Zabe) [09:25:56] (03CR) 10Zabe: [C:03+2] Use virtual domain in NameTableStore for collation [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337857 (https://phabricator.wikimedia.org/T405812) (owner: 10Zabe) [09:26:06] (03CR) 10CI reject: [V:04-1] mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [09:26:23] (03CR) 10Volans: mysql.py: Add x4 (031 comment) [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [09:26:59] (03CR) 10Zabe: [C:03+1] mysql.py: Add x4 (031 comment) [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [09:28:07] FIRING: [8x] ProbeDown: Service aqs2007-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:29:48] !log remove GRE tunnels eqiad-drmrs eqdfw-ulsfo [09:29:49] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:30:13] (03Merged) 10jenkins-bot: Use virtual domain in NameTableStore for collation [core] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337856 (https://phabricator.wikimedia.org/T405812) (owner: 10Zabe) [09:31:00] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [09:31:20] (03Merged) 10jenkins-bot: Use virtual domain in NameTableStore for collation [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337857 (https://phabricator.wikimedia.org/T405812) (owner: 10Zabe) [09:31:23] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet [09:31:24] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet [09:32:09] !log ayounsi@cumin1003 START - Cookbook sre.dns.netbox [09:32:15] \o/ [09:32:40] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:32:40] (03PS1) 10Dpogorzelski: cnpg: fix major-upgrade access to the kube API [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337867 [09:32:41] !log zabe@deploy1003 Started scap sync-world: Backport for [[gerrit:1337856|Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857|Use virtual domain in NameTableStore for collation (T405812)]] [09:32:44] T405812: Migrate categorylinks to virtual domain - https://phabricator.wikimedia.org/T405812 [09:32:47] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [09:33:07] RESOLVED: [8x] ProbeDown: Service aqs2007-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:33:21] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm [09:33:35] later today I drop a couple of tables from x4 and s4 which would trigger fatal if it tries to read from out-dated tables [09:33:59] in one replica each cluster first [09:34:06] (03PS2) 10Dpogorzelski: cnpg: fix major-upgrade access to the kube API [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337867 [09:34:39] FIRING: CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=drmrs&var-device=cr1-drmrs:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [09:34:59] !log ayounsi@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003" [09:35:10] FIRING: BFDdown: BFD session down between cr2-eqiad and 185.15.58.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:35:12] (03CR) 10Ayounsi: "recheck" [dns] - 10https://gerrit.wikimedia.org/r/1337865 (owner: 10Ayounsi) [09:35:41] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003" [09:35:42] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [09:37:08] !log zabe@deploy1003 zabe: Backport for [[gerrit:1337856|Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857|Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [09:38:07] FIRING: [8x] ProbeDown: Service aqs2007-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:38:27] (03PS2) 10Ayounsi: Reclaim DNS v6 PTR for decom GRE tunnels [dns] - 10https://gerrit.wikimedia.org/r/1337865 [09:39:13] !log zabe@deploy1003 zabe: Continuing with deployment [09:39:22] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12296675 (10MoritzMuehlenhoff) [09:40:10] RESOLVED: BFDdown: BFD session down between cr2-eqiad and 185.15.58.151 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [09:42:06] (03CR) 10Volans: [C:03+1] "LGTM" [dns] - 10https://gerrit.wikimedia.org/r/1337865 (owner: 10Ayounsi) [09:42:46] (03CR) 10Ayounsi: [C:03+2] Reclaim DNS v6 PTR for decom GRE tunnels [dns] - 10https://gerrit.wikimedia.org/r/1337865 (owner: 10Ayounsi) [09:43:09] !log ayounsi@dns1004 START - running authdns-update [09:44:19] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet [09:45:23] !log ayounsi@dns1004 END - running authdns-update [09:45:26] (03PS1) 10Kevin Bazira: admin_ng: Add tool-server namespace to LiftWing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337869 (https://phabricator.wikimedia.org/T436892) [09:45:51] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12296705 (10BTullis) A few weeks later, after our upgrade to bookworm on all of these `an-worker` hosts, I run the command again to see how many hosts have few... [09:45:56] !log zabe@deploy1003 Finished scap sync-world: Backport for [[gerrit:1337856|Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857|Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s) [09:45:59] T405812: Migrate categorylinks to virtual domain - https://phabricator.wikimedia.org/T405812 [09:48:07] FIRING: [4x] ProbeDown: Service aqs2008-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:51:34] (03CR) 10Btullis: [C:03+1] cnpg: fix major-upgrade access to the kube API [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337867 (owner: 10Dpogorzelski) [09:51:51] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage [09:53:07] RESOLVED: [4x] ProbeDown: Service aqs2008-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:53:29] (03PS2) 10Marostegui: mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) [09:54:42] (03CR) 10Dpogorzelski: [C:03+2] cnpg: fix major-upgrade access to the kube API [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337867 (owner: 10Dpogorzelski) [09:54:59] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage [09:57:13] (03CR) 10CI reject: [V:04-1] mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [09:58:20] (03Merged) 10jenkins-bot: cnpg: fix major-upgrade access to the kube API [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337867 (owner: 10Dpogorzelski) [09:59:03] !log mvernon@cumin1003 END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet [09:59:08] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet [09:59:46] Amir1: For my own organizational purposes, when do you expect to clean up one x4 replica so I can move stuff to sanitarium, not in a rush, but just for me to organise [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1000) [10:00:38] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:01:13] marostegui: later today. Would tomorrow be okay with you? If you want faster I can do it but have to leave some meetings [10:01:37] !log mvernon@cumin1003 END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet [10:01:46] (03PS3) 10Marostegui: mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) [10:01:51] Amir1: sure, that works [10:02:19] Awesome [10:04:19] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet [10:04:51] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync [10:04:54] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync [10:05:06] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet [10:07:29] 06SRE-OnFire, 06MediaWiki-Engineering, 06ServiceOps, 07Sustainability (Incident Followup): Reduce the amount of messages sent through channel:Memcached during failures - https://phabricator.wikimedia.org/T390529#12296766 (10MLechvien-WMF) [10:07:50] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet [10:07:53] (03PS1) 10Btullis: ceph: Add a script to compare CephX keyrings with the cluster [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) [10:07:55] (03PS1) 10Btullis: ceph: Make the CephX import guard compare the key [puppet] - 10https://gerrit.wikimedia.org/r/1337876 (https://phabricator.wikimedia.org/T437233) [10:08:42] (03CR) 10CI reject: [V:04-1] ceph: Add a script to compare CephX keyrings with the cluster [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) (owner: 10Btullis) [10:08:45] (03CR) 10Zabe: [C:03+1] mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [10:08:53] 06SRE-OnFire, 06MediaWiki-Engineering, 06ServiceOps, 07Sustainability (Incident Followup): Reduce the amount of messages sent through channel:Memcached during failures - https://phabricator.wikimedia.org/T390529#12296775 (10MLechvien-WMF) Can we assess if this task is still relevant and scope it well? As t... [10:10:54] (03PS2) 10Btullis: ceph: Add a script to compare CephX keyrings with the cluster [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) [10:10:54] (03PS2) 10Btullis: ceph: Make the CephX import guard compare the key [puppet] - 10https://gerrit.wikimedia.org/r/1337876 (https://phabricator.wikimedia.org/T437233) [10:11:44] (03CR) 10CI reject: [V:04-1] ceph: Add a script to compare CephX keyrings with the cluster [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) (owner: 10Btullis) [10:12:30] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm [10:13:35] (03CR) 10Marostegui: [C:03+2] mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [10:13:51] (03CR) 10Marostegui: [V:03+2 C:03+2] mysql.py: Add x4 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1337864 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [10:18:41] !log mvernon@cumin1003 START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet [10:19:39] (03PS4) 10Samtar: IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307862 (https://phabricator.wikimedia.org/T431355) [10:22:07] FIRING: [2x] ProbeDown: Service aqs2009-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:22:21] (03PS3) 10Btullis: ceph: Add a script to compare CephX keyrings with the cluster [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) [10:22:21] (03PS3) 10Btullis: ceph: Make the CephX import guard compare the key [puppet] - 10https://gerrit.wikimedia.org/r/1337876 (https://phabricator.wikimedia.org/T437233) [10:23:05] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1319853 (https://phabricator.wikimedia.org/T421952) (owner: 10Brouberol) [10:23:29] (03CR) 10Brouberol: [C:03+2] "Thanks Moritz!" [puppet] - 10https://gerrit.wikimedia.org/r/1319853 (https://phabricator.wikimedia.org/T421952) (owner: 10Brouberol) [10:24:17] 10SRE-SLO, 10Citoid, 06Editing-team, 07Sustainability (Incident Followup): Create probe for citoid that tests whether requests can access the outside internet - https://phabricator.wikimedia.org/T381372#12296827 (10Mvolz) p:05High→03Medium [10:26:35] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [10:27:07] FIRING: [4x] ProbeDown: Service aqs2009-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:27:50] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet [10:27:51] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet [10:28:22] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [10:28:40] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm [10:29:02] (03PS5) 10Samtar: IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307862 (https://phabricator.wikimedia.org/T431355) [10:29:09] (03PS1) 10Kevin Bazira: ml-services: Add tool-server deployment configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337879 (https://phabricator.wikimedia.org/T436892) [10:30:10] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet [10:30:37] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307862 (https://phabricator.wikimedia.org/T431355) (owner: 10Samtar) [10:31:40] (03Merged) 10jenkins-bot: IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1307862 (https://phabricator.wikimedia.org/T431355) (owner: 10Samtar) [10:32:02] (03PS1) 10Brouberol: data: fix sudo permissions associated to analytics-admins [puppet] - 10https://gerrit.wikimedia.org/r/1337880 (https://phabricator.wikimedia.org/T421952) [10:32:03] !log samtar@deploy1003 Started scap sync-world: Backport for [[gerrit:1307862|IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] [10:32:06] T431355: Enable WatchstarPopover on beta - https://phabricator.wikimedia.org/T431355 [10:32:07] RESOLVED: [4x] ProbeDown: Service aqs2009-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:34:33] (03PS2) 10Brouberol: data: fix sudo permissions associated to analytics-admins [puppet] - 10https://gerrit.wikimedia.org/r/1337880 (https://phabricator.wikimedia.org/T421952) [10:34:49] !log mvernon@cumin1003 END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet [10:36:48] !log samtar@deploy1003 samtar: Backport for [[gerrit:1307862|IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [10:37:05] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet [10:37:07] FIRING: [4x] ProbeDown: Service aqs2009-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:37:26] !log mvernon@cumin1003 END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet [10:38:19] !log samtar@deploy1003 samtar: Continuing with deployment [10:39:18] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet [10:39:54] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1337880 (https://phabricator.wikimedia.org/T421952) (owner: 10Brouberol) [10:43:01] !log samtar@deploy1003 Finished scap sync-world: Backport for [[gerrit:1307862|IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s) [10:43:04] T431355: Enable WatchstarPopover on beta - https://phabricator.wikimedia.org/T431355 [10:43:25] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip for echostore, eventstreams, kartotherian codfw [puppet] - 10https://gerrit.wikimedia.org/r/1337851 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [10:43:30] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip for echostore, eventstreams, kartotherian eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1337852 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [10:44:06] marostegui: which host do you want me to clean up? any preference [10:44:41] Amir1: db1260 would work [10:45:22] (03CR) 10Dragoniez: "Thanks :) I actually meant that we could limit the change to the following line for viwiki and leave all other lines unchanged:" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1335401 (https://phabricator.wikimedia.org/T437006) (owner: 10Tryvix1509) [10:45:25] awesome [10:46:48] (03CR) 10Tryvix1509: "Done" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1335401 (https://phabricator.wikimedia.org/T437006) (owner: 10Tryvix1509) [10:47:24] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage [10:52:07] RESOLVED: [4x] ProbeDown: Service aqs2009-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:52:27] mvernon@cumin1003 upgrade-firmware (PID 3741900) is awaiting input [10:52:57] (03PS3) 10Muehlenhoff: Setup a syncrepl cluster on trixie/MDB [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) [10:53:04] (03PS4) 10Muehlenhoff: Setup a syncrepl cluster on trixie/MDB [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) [10:53:10] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage [10:54:45] Amir1: mind creating a task under https://phabricator.wikimedia.org/T404715#12296713 for tracking the db1260 clean up and to have it there for posterity? [10:54:52] even if it is already done, just to track it [10:55:54] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [10:56:44] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync [10:56:46] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync [10:57:04] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync [10:57:04] (03CR) 10JMeybohm: [C:03+1] "tbh I don't think you need to bump the chart version for this. There is no functional change that requires a rollout and if you don't bump" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333699 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [10:57:06] !log dpogorzelski@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync [11:00:21] marostegui: sure thing [11:01:29] (03PS4) 10Btullis: ceph: Add a script to compare CephX keyrings with the cluster [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) [11:01:29] (03PS4) 10Btullis: ceph: Make the CephX import guard compare the key [puppet] - 10https://gerrit.wikimedia.org/r/1337876 (https://phabricator.wikimedia.org/T437233) [11:01:56] (03PS1) 10Muehlenhoff: amd::gpu: Remove support for bullseye [puppet] - 10https://gerrit.wikimedia.org/r/1337885 [11:03:07] (03CR) 10Btullis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) (owner: 10Btullis) [11:03:53] (03CR) 10Clément Goubert: "Couple of last changes" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332764 (https://phabricator.wikimedia.org/T434109) (owner: 10DLynch) [11:03:57] (03PS1) 10Muehlenhoff: Stop building a Bullseye base image [puppet] - 10https://gerrit.wikimedia.org/r/1337886 [11:07:44] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337885 (owner: 10Muehlenhoff) [11:08:06] (03CR) 10Blake: [C:03+2] mesh: upgrade to 1.3.3 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334741 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:08:36] (03PS1) 10Kevin Bazira: profile::k8s::deployment_server: add config for tool-server [puppet] - 10https://gerrit.wikimedia.org/r/1337888 (https://phabricator.wikimedia.org/T436892) [11:09:06] (03PS1) 10Marostegui: mariadb: Clean up x4 references [puppet] - 10https://gerrit.wikimedia.org/r/1337890 (https://phabricator.wikimedia.org/T437229) [11:09:48] (03PS1) 10Marostegui: wmnet: Add x4 [dns] - 10https://gerrit.wikimedia.org/r/1337891 (https://phabricator.wikimedia.org/T437229) [11:09:56] (03CR) 10Marostegui: "This is a noop" [puppet] - 10https://gerrit.wikimedia.org/r/1337890 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:09:58] (03CR) 10Marostegui: [C:03+2] mariadb: Clean up x4 references [puppet] - 10https://gerrit.wikimedia.org/r/1337890 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:10:22] (03Merged) 10jenkins-bot: mesh: upgrade to 1.3.3 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334741 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:10:56] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm [11:10:58] (03CR) 10Blake: [C:03+2] mesh: Add switch to enable k8s native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334742 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:11:10] (03PS2) 10Blake: mesh: Add switch to enable k8s native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334742 (https://phabricator.wikimedia.org/T417800) [11:11:10] (03CR) 10CI reject: [V:04-1] mesh: Add switch to enable k8s native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334742 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:11:18] (03CR) 10Blake: [V:03+2 C:03+2] mesh: Add switch to enable k8s native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334742 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:12:17] !log mvernon@cumin1003 START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet [11:12:27] (03PS1) 10Marostegui: mariadb: Clean up x4 references in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1337893 (https://phabricator.wikimedia.org/T437229) [11:12:41] (03CR) 10Marostegui: "This is a noop" [puppet] - 10https://gerrit.wikimedia.org/r/1337893 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:13:13] (03CR) 10Marostegui: [C:03+2] mariadb: Clean up x4 references in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1337893 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:13:57] (03Merged) 10jenkins-bot: mesh: Add switch to enable k8s native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334742 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:14:18] (03CR) 10Blake: [C:03+2] statsd: Upgrade to 1.0.5 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334738 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:15:40] FIRING: [3x] ProbeDown: Service aqs2010-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:16:53] (03PS1) 10Marostegui: mariadb.yaml: Add x4 [puppet] - 10https://gerrit.wikimedia.org/r/1337894 (https://phabricator.wikimedia.org/T437229) [11:16:57] (03Merged) 10jenkins-bot: statsd: Upgrade to 1.0.5 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334738 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:17:00] (03PS2) 10Blake: statsd: Enable a native k8s sidecar option. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334739 (https://phabricator.wikimedia.org/T417800) [11:17:49] (03PS1) 10Muehlenhoff: Remove obsolete sub secrets [labs/private] - 10https://gerrit.wikimedia.org/r/1337896 [11:17:50] (03PS1) 10Kevin Bazira: tool-server: Add envoy services-proxy listener [puppet] - 10https://gerrit.wikimedia.org/r/1337895 (https://phabricator.wikimedia.org/T436892) [11:20:05] (03CR) 10Blake: [C:03+2] statsd: Enable a native k8s sidecar option. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334739 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:20:09] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [11:20:40] FIRING: [4x] ProbeDown: Service aqs2010-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:21:54] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet [11:21:55] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [11:21:56] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet [11:22:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:22:17] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm [11:22:45] (03Merged) 10jenkins-bot: statsd: Enable a native k8s sidecar option. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334739 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:23:15] (03CR) 10Blake: [C:03+2] httpd: Add version 1.0.3 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334012 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:23:23] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet [11:23:27] (03PS1) 10SomeRandomDeveloper: Skip DB read-only check for CentralAuthSessionManager [extensions/CentralAuth] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337898 (https://phabricator.wikimedia.org/T437273) [11:23:40] (03PS1) 10SomeRandomDeveloper: Skip DB read-only check for CentralAuthSessionManager [extensions/CentralAuth] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337899 (https://phabricator.wikimedia.org/T437273) [11:24:26] (03PS1) 10Kevin Bazira: service::catalog: Add tool-server [puppet] - 10https://gerrit.wikimedia.org/r/1337900 (https://phabricator.wikimedia.org/T436892) [11:25:40] FIRING: [4x] ProbeDown: Service aqs2010-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:26:01] (03Merged) 10jenkins-bot: httpd: Add version 1.0.3 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334012 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:26:20] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Remove obsolete sub secrets [labs/private] - 10https://gerrit.wikimedia.org/r/1337896 (owner: 10Muehlenhoff) [11:26:21] (03PS2) 10Blake: httpd: Add switch to enable Kubernetes native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334013 (https://phabricator.wikimedia.org/T417800) [11:27:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.73% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:27:51] (03PS1) 10Muehlenhoff: Stub secrets for new Trixie OpenLDAP roles [labs/private] - 10https://gerrit.wikimedia.org/r/1337901 (https://phabricator.wikimedia.org/T331699) [11:28:54] (03PS1) 10Marostegui: check_depooled.sh: Add x4 [software] - 10https://gerrit.wikimedia.org/r/1337902 (https://phabricator.wikimedia.org/T437229) [11:29:36] (03CR) 10Marostegui: "This is a noop" [software] - 10https://gerrit.wikimedia.org/r/1337902 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:29:56] (03CR) 10Marostegui: [C:03+2] check_depooled.sh: Add x4 [software] - 10https://gerrit.wikimedia.org/r/1337902 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:30:15] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker2005 has a BGP session which is not in the 'established' state. [11:30:27] (03CR) 10Blake: [C:03+2] httpd: Add switch to enable Kubernetes native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334013 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:30:33] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Stub secrets for new Trixie OpenLDAP roles [labs/private] - 10https://gerrit.wikimedia.org/r/1337901 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [11:30:40] FIRING: [4x] ProbeDown: Service aqs2010-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:31:03] (03Merged) 10jenkins-bot: check_depooled.sh: Add x4 [software] - 10https://gerrit.wikimedia.org/r/1337902 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:33:15] (03PS1) 10Muehlenhoff: Remove Hiera entries for hosts long decommissioned [labs/private] - 10https://gerrit.wikimedia.org/r/1337903 [11:33:28] (03Merged) 10jenkins-bot: httpd: Add switch to enable Kubernetes native sidecar behaviour. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334013 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:34:18] (03PS1) 10Kevin Bazira: Add tool-server CNAMEs to k8s-ingress-ml-serve [dns] - 10https://gerrit.wikimedia.org/r/1337904 (https://phabricator.wikimedia.org/T436892) [11:34:21] (03CR) 10Blake: [C:03+2] php-fpm: Add version 1.0.2 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334015 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:35:41] (03PS1) 10Muehlenhoff: Remove mcrouter stub secrets for long-decommissioned mw baremetal hosts [labs/private] - 10https://gerrit.wikimedia.org/r/1337905 [11:35:55] mvernon@cumin1003 upgrade-firmware (PID 3747795) is awaiting input [11:36:03] (03PS5) 10Muehlenhoff: Setup a syncrepl cluster on trixie/MDB [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) [11:36:31] (03CR) 10Majavah: [C:03+1] Remove Hiera entries for hosts long decommissioned [labs/private] - 10https://gerrit.wikimedia.org/r/1337903 (owner: 10Muehlenhoff) [11:36:43] (03CR) 10JMeybohm: [C:03+1] Stop building a Bullseye base image [puppet] - 10https://gerrit.wikimedia.org/r/1337886 (owner: 10Muehlenhoff) [11:39:14] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage [11:40:41] !log dropping unneeded tables from x4 - db1260 (T437278) [11:40:44] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:40:45] T437278: Drop unneeded tables from x4 and s4 - https://phabricator.wikimedia.org/T437278 [11:41:07] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [11:41:39] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Remove Hiera entries for hosts long decommissioned [labs/private] - 10https://gerrit.wikimedia.org/r/1337903 (owner: 10Muehlenhoff) [11:42:32] (03CR) 10Blake: [V:03+2 C:03+2] php-fpm: Add version 1.0.2 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334015 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:43:24] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage [11:43:49] (03CR) 10Zabe: [C:03+1] wmnet: Add x4 [dns] - 10https://gerrit.wikimedia.org/r/1337891 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:44:21] (03CR) 10Marostegui: [C:03+2] wmnet: Add x4 [dns] - 10https://gerrit.wikimedia.org/r/1337891 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:44:24] marostegui: I have dropped biggest stuff from db1260, to drop all tables that are not needed (around seventy more), I'll need half a day but does right now work? [11:44:25] (03PS2) 10Blake: php-fpm: Add version 1.0.2 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334015 (https://phabricator.wikimedia.org/T417800) [11:44:33] https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&refresh=5m&var-server=db1260&var-datasource=000000026&var-cluster=mysql&viewPanel=panel-28&from=now-1h&to=now&timezone=utc [11:44:36] !log marostegui@dns1004 START - running authdns-update [11:45:07] (03CR) 10Blake: [V:03+2 C:03+2] php-fpm: Add version 1.0.2 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334015 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:46:36] (03CR) 10Zabe: [C:03+1] mariadb.yaml: Add x4 [puppet] - 10https://gerrit.wikimedia.org/r/1337894 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [11:46:44] !log marostegui@dns1004 END - running authdns-update [11:47:49] (03Merged) 10jenkins-bot: php-fpm: Add version 1.0.2 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334015 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:49:04] (03PS2) 10Blake: php-fpm: Make php-fpm-exporter a sidecar. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334016 (https://phabricator.wikimedia.org/T417800) [11:49:08] I go back to the event, will do more later [11:49:20] (03PS6) 10Muehlenhoff: Setup a syncrepl cluster on trixie/MDB [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) [11:50:18] (03PS2) 10Muehlenhoff: Remove mcrouter stub secrets for long-decommissioned mw baremetal hosts [labs/private] - 10https://gerrit.wikimedia.org/r/1337905 [11:52:44] Amir1: not super urgent, but I think it would be nice to rename/drop the link tables in one of the s4 hosts, to make sure we are not still reading outdated data somewhere (basically querying wrongly the other way around) [11:53:49] (03CR) 10Zabe: [C:03+2] Skip DB read-only check for CentralAuthSessionManager [extensions/CentralAuth] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337898 (https://phabricator.wikimedia.org/T437273) (owner: 10SomeRandomDeveloper) [11:53:50] (03CR) 10Zabe: [C:03+2] Skip DB read-only check for CentralAuthSessionManager [extensions/CentralAuth] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337899 (https://phabricator.wikimedia.org/T437273) (owner: 10SomeRandomDeveloper) [11:54:02] (03CR) 10Blake: [C:03+2] php-fpm: Make php-fpm-exporter a sidecar. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334016 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:54:30] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [11:56:01] (03Merged) 10jenkins-bot: Skip DB read-only check for CentralAuthSessionManager [extensions/CentralAuth] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337898 (https://phabricator.wikimedia.org/T437273) (owner: 10SomeRandomDeveloper) [11:56:04] (03Merged) 10jenkins-bot: Skip DB read-only check for CentralAuthSessionManager [extensions/CentralAuth] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337899 (https://phabricator.wikimedia.org/T437273) (owner: 10SomeRandomDeveloper) [11:56:19] (03Merged) 10jenkins-bot: php-fpm: Make php-fpm-exporter a sidecar. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1334016 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:56:34] !log zabe@deploy1003 Started scap sync-world: Backport for [[gerrit:1337898|Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899|Skip DB read-only check for CentralAuthSessionManager (T437273)]] [11:56:37] T437273: MediaWiki is not capable to operate in read only mode anymore - https://phabricator.wikimedia.org/T437273 [11:57:26] (03CR) 10Blake: [C:03+2] mcrouter: update to 1.3.6 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333855 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [11:58:14] (03PS2) 10Dpogorzelski: ml: images for Lift Wing Studio [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1333876 [12:00:00] (03Merged) 10jenkins-bot: mcrouter: update to 1.3.6 (copy) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333855 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1200) [12:00:27] (03PS3) 10Blake: mcrouter: Upgrade to 1.3.6. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333856 (https://phabricator.wikimedia.org/T417800) [12:00:45] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm [12:01:11] !log zabe@deploy1003 zabe, somerandomdeveloper: Backport for [[gerrit:1337898|Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899|Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [12:01:40] !log zabe@deploy1003 zabe, somerandomdeveloper: Continuing with deployment [12:01:49] !log mvernon@cumin1003 START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet [12:05:40] FIRING: [8x] ProbeDown: Service aqs2010-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:06:18] (03PS2) 10Kevin Bazira: ml-services: Add tool-server deployment configs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337879 (https://phabricator.wikimedia.org/T436892) [12:06:28] !log zabe@deploy1003 Finished scap sync-world: Backport for [[gerrit:1337898|Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899|Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s) [12:06:31] T437273: MediaWiki is not capable to operate in read only mode anymore - https://phabricator.wikimedia.org/T437273 [12:07:47] Thanks zabe! [12:08:01] yw:) [12:08:10] (03PS2) 10Cwhite: WikimediaEvents: enable client-side error logging for plwikisource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1196484 (https://phabricator.wikimedia.org/T340187) [12:09:17] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [12:10:24] (03PS1) 10Muehlenhoff: Create a separate role openldap::replica_mdb [puppet] - 10https://gerrit.wikimedia.org/r/1337910 (https://phabricator.wikimedia.org/T331699) [12:10:49] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet [12:10:50] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet [12:11:04] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [12:11:08] (03CR) 10CI reject: [V:04-1] Create a separate role openldap::replica_mdb [puppet] - 10https://gerrit.wikimedia.org/r/1337910 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [12:11:47] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm [12:11:50] !log mvernon@cumin1003 START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet [12:13:02] (03CR) 10Brouberol: [C:03+1] ceph: Add a script to compare CephX keyrings with the cluster [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) (owner: 10Btullis) [12:13:33] (03CR) 10Brouberol: [C:03+1] ceph: Make the CephX import guard compare the key [puppet] - 10https://gerrit.wikimedia.org/r/1337876 (https://phabricator.wikimedia.org/T437233) (owner: 10Btullis) [12:13:33] (03CR) 10Clément Goubert: [C:03+1] Remove mcrouter stub secrets for long-decommissioned mw baremetal hosts [labs/private] - 10https://gerrit.wikimedia.org/r/1337905 (owner: 10Muehlenhoff) [12:15:40] FIRING: [8x] ProbeDown: Service aqs2010-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:16:06] (03CR) 10Brouberol: [C:03+2] data: fix sudo permissions associated to analytics-admins [puppet] - 10https://gerrit.wikimedia.org/r/1337880 (https://phabricator.wikimedia.org/T421952) (owner: 10Brouberol) [12:17:08] (03PS2) 10Muehlenhoff: Create a separate role openldap::replica_mdb [puppet] - 10https://gerrit.wikimedia.org/r/1337910 (https://phabricator.wikimedia.org/T331699) [12:17:34] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Remove mcrouter stub secrets for long-decommissioned mw baremetal hosts [labs/private] - 10https://gerrit.wikimedia.org/r/1337905 (owner: 10Muehlenhoff) [12:20:40] RESOLVED: [7x] ProbeDown: Service aqs2010-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:22:13] !log klausman@cumin1004 START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad [12:22:16] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet [12:23:57] (03CR) 10JMeybohm: [V:03+1] "PCC SUCCESS (NOOP 13): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9379/consol" [puppet] - 10https://gerrit.wikimedia.org/r/1330128 (owner: 10Giuseppe Lavagetto) [12:24:23] mvernon@cumin1003 upgrade-firmware (PID 3751655) is awaiting input [12:26:42] (03CR) 10Elukey: [C:03+1] Stop building a Bullseye base image [puppet] - 10https://gerrit.wikimedia.org/r/1337886 (owner: 10Muehlenhoff) [12:27:41] (03PS1) 10Slyngshede: C:mediawiki::tools::cache_warmup remove mobile urls [puppet] - 10https://gerrit.wikimedia.org/r/1337911 (https://phabricator.wikimedia.org/T436781) [12:28:15] (03CR) 10Slyngshede: "The code for handling the mobile URIs can be removed in a separate CR." [puppet] - 10https://gerrit.wikimedia.org/r/1337911 (https://phabricator.wikimedia.org/T436781) (owner: 10Slyngshede) [12:29:10] (03CR) 10Elukey: [C:03+1] "Very ignorant about the change: are the two new ldap-rw hosts sync with each other using the other node as master?" [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [12:31:49] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage [12:32:19] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet [12:32:35] (03PS3) 10Muehlenhoff: Create a separate role openldap::replica_mdb [puppet] - 10https://gerrit.wikimedia.org/r/1337910 (https://phabricator.wikimedia.org/T331699) [12:35:07] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [12:36:59] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage [12:37:41] (03CR) 10Muehlenhoff: [C:03+2] Stop building a Bullseye base image [puppet] - 10https://gerrit.wikimedia.org/r/1337886 (owner: 10Muehlenhoff) [12:38:50] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet [12:38:50] (03PS3) 10Federico Ceratto: data-persistence: Alert on depooled hosts without silence [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) [12:38:51] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet [12:38:56] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet [12:40:14] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad [12:40:16] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations, 07Epic: Migrate Docker images running in Production away from Bullseye - https://phabricator.wikimedia.org/T416452#12297211 (10MoritzMuehlenhoff) We also stopped building new Bullseye base images now: https://gerrit.wikimedia.org/r/c/operations/pupp... [12:40:28] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad [12:41:03] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations, 07Epic: Migrate Docker images running in Production away from Bullseye - https://phabricator.wikimedia.org/T416452#12297212 (10MoritzMuehlenhoff) [12:41:23] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device lsw1-f1-eqiad [12:41:25] 06SRE, 06Data-Platform-SRE, 06Infrastructure-Foundations, 07Epic: Migrate Docker images running in Production away from Bullseye - https://phabricator.wikimedia.org/T416452#12297213 (10MoritzMuehlenhoff) [12:41:29] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad [12:42:01] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr2-eqiad [12:42:09] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad [12:42:31] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr1-eqiad [12:42:44] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad [12:42:55] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw [12:43:03] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw [12:43:59] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet [12:44:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-drmrs and cr2-eqiad (185.15.58.150) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=drmrs&var-device=cr1-drmrs:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [12:44:53] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device lsw1-e1-eqiad [12:45:00] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad [12:45:20] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad [12:45:26] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad [12:45:51] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device lsw1-e2-eqiad [12:45:58] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad [12:48:07] FIRING: [2x] ProbeDown: Service aqs2011-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:49:16] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for echostore, eventstreams, kartotherian codfw [puppet] - 10https://gerrit.wikimedia.org/r/1337851 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [12:49:20] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad [12:49:33] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad [12:49:36] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device lsw1-e1-eqiad [12:49:36] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad [12:49:39] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad [12:49:39] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad [12:49:42] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device lsw1-e2-eqiad [12:49:42] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad [12:49:45] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad [12:49:51] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad [12:49:54] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device lsw1-e3-eqiad [12:50:01] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad [12:50:04] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device lsw1-f2-eqiad [12:50:11] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad [12:50:14] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device lsw1-f3-eqiad [12:50:18] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet [12:50:19] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet [12:50:21] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad [12:50:24] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr3-ulsfo [12:50:25] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet [12:50:31] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw [12:50:42] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo [12:50:45] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr4-ulsfo [12:50:57] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo [12:51:00] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr3-eqsin [12:51:35] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin [12:51:37] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr1-drmrs [12:51:51] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs [12:51:54] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr1-esams [12:52:15] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams [12:52:17] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr2-drmrs [12:52:37] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs [12:52:40] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device asw1-bw27-esams [12:53:01] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams [12:53:03] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr2-esams [12:53:07] FIRING: [4x] ProbeDown: Service aqs2011-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:53:13] (03CR) 10Filippo Giunchedi: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) (owner: 10Btullis) [12:53:24] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams [12:53:27] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device asw1-b12-drmrs [12:53:40] (03CR) 10CWilliams: data-persistence: Alert on depooled hosts without silence (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [12:53:46] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs [12:53:49] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device asw1-b13-drmrs [12:54:08] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs [12:54:10] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [12:54:11] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device asw1-by27-esams [12:54:31] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams [12:54:34] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr2-codfw [12:54:45] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm [12:54:49] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw [12:54:52] !log ayounsi@cumin1003 START - Cookbook sre.network.tls for network device cr1-codfw [12:55:07] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw [12:55:07] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [12:55:08] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw [12:55:18] (03PS1) 10Brouberol: airflow: ensure emails are logged but not sent in devenvs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337917 (https://phabricator.wikimedia.org/T416596) [12:55:27] (03CR) 10Filippo Giunchedi: [C:03+1] "LGTM for when the time comes" [puppet] - 10https://gerrit.wikimedia.org/r/1337876 (https://phabricator.wikimedia.org/T437233) (owner: 10Btullis) [12:55:27] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet [12:55:40] !log ayounsi@cumin1003 START - Cookbook sre.network.peering with action 'clear' for AS: 35320 [12:55:52] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320 [12:56:13] !log mvernon@cumin1003 START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet [12:56:27] (03CR) 10Muehlenhoff: "Yes, we have multi-master replication, i.e. changes can be made to either r/w endpoint. The current servers on BDB are doing the same: htt" [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [12:57:38] (03CR) 10CWilliams: data-persistence: Alert on depooled hosts without silence (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [12:58:01] !log ayounsi@cumin1003 START - Cookbook sre.network.peering with action 'email' for AS: 34655 [12:58:07] RESOLVED: [4x] ProbeDown: Service aqs2011-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:58:38] (03CR) 10Marostegui: "db1174 is also not present in dbctl" [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [12:58:40] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655 [13:00:04] Lucas_WMDE, urbanecm, and TheresNoTime: Time to snap out of that daydream and deploy UTC afternoon backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1300). [13:00:04] Tran: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:08] o/ [13:00:27] (03PS4) 10Jelto: service::catalog: Set ipip for echostore, eventstreams, kartotherian eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1337852 (https://phabricator.wikimedia.org/T420436) [13:00:27] I will self-deploy and if no one else needs the window, will also be running a maint script. [13:00:31] (03CR) 10Clément Goubert: [C:03+1] C:mediawiki::tools::cache_warmup remove mobile urls [puppet] - 10https://gerrit.wikimedia.org/r/1337911 (https://phabricator.wikimedia.org/T436781) (owner: 10Slyngshede) [13:01:19] (03CR) 10TrainBranchBot: [C:03+2] "Approved by stran@deploy1003 using scap backport" [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337610 (https://phabricator.wikimedia.org/T436508) (owner: 10STran) [13:01:45] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet [13:01:45] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet [13:01:51] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet [13:02:25] RESOLVED: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:03:07] FIRING: [8x] ProbeDown: Service aqs2011-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:03:24] (03CR) 10Brouberol: [C:03+1] ml: images for Lift Wing Studio [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1333876 (owner: 10Dpogorzelski) [13:04:22] (03PS3) 10Dpogorzelski: ml: images for Lift Wing Studio [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1333876 [13:04:31] (03CR) 10Dpogorzelski: [C:03+2] ml: images for Lift Wing Studio [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1333876 (owner: 10Dpogorzelski) [13:04:33] !log mvernon@cumin2003 START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [13:04:35] (03CR) 10Dpogorzelski: [V:03+2 C:03+2] ml: images for Lift Wing Studio [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1333876 (owner: 10Dpogorzelski) [13:05:00] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for echostore, eventstreams, kartotherian eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1337852 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [13:05:21] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [13:06:16] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet [13:06:17] !log mvernon@cumin1003 END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet [13:06:20] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [13:06:41] !log mvernon@cumin2003 START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm [13:06:53] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet [13:08:07] RESOLVED: [8x] ProbeDown: Service aqs2011-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:08:51] !log ayounsi@cumin1003 START - Cookbook sre.network.peering with action 'email' for AS: 14593 [13:09:02] !log jelto@cumin1003 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [13:09:38] (03Merged) 10jenkins-bot: Split out edit and block-based filters from activity filters [extensions/CheckUser] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337610 (https://phabricator.wikimedia.org/T436508) (owner: 10STran) [13:10:01] !log stran@deploy1003 Started scap sync-world: Backport for [[gerrit:1337610|Split out edit and block-based filters from activity filters (T436508)]] [13:10:04] T436508: Separating combinable and distinct case filters - https://phabricator.wikimedia.org/T436508 [13:10:07] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [13:10:07] !log jelto@cumin1003 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad [13:10:20] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593 [13:11:20] !log ayounsi@cumin1003 START - Cookbook sre.network.peering with action 'email' for AS: 2519 [13:11:29] !log ayounsi@cumin1003 END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519 [13:13:07] FIRING: [8x] ProbeDown: Service aqs2011-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:13:24] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet [13:13:26] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet [13:13:31] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet [13:14:55] FIRING: [2x] SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:15:28] !log ayounsi@cumin1003 START - Cookbook sre.network.peering with action 'email' for AS: 139628 [13:16:57] !log ayounsi@cumin1003 END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628 [13:18:07] RESOLVED: [4x] ProbeDown: Service aqs2012-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:18:33] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet [13:18:44] (03CR) 10Herron: [C:03+1] "LGTM!" [puppet] - 10https://gerrit.wikimedia.org/r/1335233 (https://phabricator.wikimedia.org/T436694) (owner: 10Andrea Denisse) [13:19:56] (03PS1) 10Jelto: service::catalog: Set ipip for mobileapps, proton, push-notifications codfw [puppet] - 10https://gerrit.wikimedia.org/r/1337923 (https://phabricator.wikimedia.org/T420436) [13:19:58] (03PS1) 10Jelto: service::catalog: Set ipip for mobileapps, proton, push-notifications eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1337924 (https://phabricator.wikimedia.org/T420436) [13:20:37] (03CR) 10Herron: [C:03+1] graphite: remove module, references, config [puppet] - 10https://gerrit.wikimedia.org/r/1332767 (https://phabricator.wikimedia.org/T435340) (owner: 10Hnowlan) [13:21:04] !log installing qemu security updates [13:21:05] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:22:48] (03CR) 10Xcollazo: [C:03+1] "LGTM, thanks for this @brouberol@wikimedia.org!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337917 (https://phabricator.wikimedia.org/T416596) (owner: 10Brouberol) [13:22:58] (03CR) 10Brouberol: [C:03+2] airflow: ensure emails are logged but not sent in devenvs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337917 (https://phabricator.wikimedia.org/T416596) (owner: 10Brouberol) [13:23:25] !log mvernon@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage [13:23:28] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet [13:23:29] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet [13:23:35] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet [13:23:37] (03CR) 10Bking: [C:03+1] airflow: ensure emails are logged but not sent in devenvs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337917 (https://phabricator.wikimedia.org/T416596) (owner: 10Brouberol) [13:24:48] (03PS3) 10Jelto: fix helm command in all NOTES.txt [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333699 (https://phabricator.wikimedia.org/T433589) [13:26:07] (03CR) 10Elukey: "Post-merge comment: please don't push images from other registries without having imported them beforehand via Docker files, in production" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1333876 (owner: 10Dpogorzelski) [13:26:18] (03CR) 10Jelto: "I agree! I removed the chart version bumps in patchset 3." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333699 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [13:28:34] (03CR) 10Btullis: "I agree and would like to request a revert, so we can discuss this more widely." [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1333876 (owner: 10Dpogorzelski) [13:28:37] (03CR) 10JMeybohm: [C:03+1] aptrepo: add thirdparty/gvisor to newer distros as well [puppet] - 10https://gerrit.wikimedia.org/r/1328540 (https://phabricator.wikimedia.org/T435758) (owner: 10Giuseppe Lavagetto) [13:28:37] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet [13:28:43] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage [13:28:57] (03CR) 10JMeybohm: [V:03+1 C:03+1] profile::containerd: format according to our style guide [puppet] - 10https://gerrit.wikimedia.org/r/1330128 (owner: 10Giuseppe Lavagetto) [13:29:09] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12297365 (10Fabfur) [13:29:57] !log stran@deploy1003 stran: Backport for [[gerrit:1337610|Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:30:01] T436508: Separating combinable and distinct case filters - https://phabricator.wikimedia.org/T436508 [13:31:31] (03PS2) 10CWilliams: admin: Added dotfiles for cwilliams [puppet] - 10https://gerrit.wikimedia.org/r/1309650 [13:31:58] testing looks good, continuing [13:32:03] !log stran@deploy1003 stran: Continuing with deployment [13:33:04] (03CR) 10JMeybohm: [V:03+1] "PCC SUCCESS (CORE_DIFF 13): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9380/c" [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [13:33:11] (03CR) 10Blake: [C:03+2] mcrouter: Upgrade to 1.3.6. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333856 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [13:33:13] (03PS1) 10Dpogorzelski: Revert "ml: images for Lift Wing Studio" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337925 [13:34:33] (03CR) 10Marostegui: [C:03+1] admin: Added dotfiles for cwilliams [puppet] - 10https://gerrit.wikimedia.org/r/1309650 (owner: 10CWilliams) [13:35:03] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet [13:35:04] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet [13:35:09] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet [13:35:26] (03Merged) 10jenkins-bot: mcrouter: Upgrade to 1.3.6. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333856 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [13:35:45] 06SRE, 10MW-on-K8s, 10Observability-Metrics, 06ServiceOps: Evaluate a Benthos replacement - https://phabricator.wikimedia.org/T437288 (10tappof) 03NEW [13:36:07] (03CR) 10Klausman: [C:03+1] Revert "ml: images for Lift Wing Studio" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337925 (owner: 10Dpogorzelski) [13:36:08] 06SRE, 10MW-on-K8s, 10Observability-Metrics, 06ServiceOps: Evaluate a Benthos replacement - https://phabricator.wikimedia.org/T437288#12297399 (10tappof) Personally, I lean towards giving Bento a chance, both to avoid having to deal with further license changes in the future and to be less tied to the owne... [13:36:21] (03CR) 10JMeybohm: fix helm command in all NOTES.txt (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333699 (https://phabricator.wikimedia.org/T433589) (owner: 10Jelto) [13:36:56] (03CR) 10Dpogorzelski: [C:03+2] Revert "ml: images for Lift Wing Studio" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337925 (owner: 10Dpogorzelski) [13:37:04] (03CR) 10Dpogorzelski: [V:03+2 C:03+2] Revert "ml: images for Lift Wing Studio" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337925 (owner: 10Dpogorzelski) [13:37:15] 06SRE, 10MW-on-K8s, 10Observability-Metrics, 06ServiceOps: Evaluate a Benthos replacement - https://phabricator.wikimedia.org/T437288#12297401 (10tappof) [13:37:20] !log cgoubert@deploy1003 helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply [13:38:43] (03CR) 10CWilliams: [C:03+2] admin: Added dotfiles for cwilliams [puppet] - 10https://gerrit.wikimedia.org/r/1309650 (owner: 10CWilliams) [13:40:12] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet [13:42:16] (03CR) 10Elukey: [C:03+1] "Perfect, thanks for the explanation :)" [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [13:44:01] !log stran@deploy1003 Finished scap sync-world: Backport for [[gerrit:1337610|Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s) [13:44:05] T436508: Separating combinable and distinct case filters - https://phabricator.wikimedia.org/T436508 [13:44:25] (03PS1) 10Clément Goubert: mw-debug: fix staging images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337927 [13:44:47] 06SRE, 10MW-on-K8s, 10Observability-Metrics, 06ServiceOps: Evaluate a Benthos replacement - https://phabricator.wikimedia.org/T437288#12297446 (10Raine) I briefly looked at Bento earlier this year and it looked good, so consider this a +0.5 :D [13:44:50] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12297447 (10ssingh) Approved. [13:45:46] (03CR) 10Kamila Součková: [C:03+1] mw-debug: fix staging images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337927 (owner: 10Clément Goubert) [13:46:31] !log mvernon@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm [13:46:33] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet [13:46:34] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet [13:46:39] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet [13:46:55] !log cdobbins@cumin1003 START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox [13:46:55] 06SRE, 06Infrastructure-Foundations: Migrate remaining container build/report steps from build2001 to build2004 - https://phabricator.wikimedia.org/T417389#12297450 (10MoritzMuehlenhoff) There was some disruption on the Debian archive side for bullseye, some package got accidentally removed a little faster the... [13:47:26] (03CR) 10Clément Goubert: [C:03+2] mw-debug: fix staging images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337927 (owner: 10Clément Goubert) [13:47:42] FIRING: [2x] ProbeDown: Service aqs2012-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:48:15] (03CR) 10Arnaudb: [C:03+1] "lgtm, thanks for the patch!" [puppet] - 10https://gerrit.wikimedia.org/r/1337657 (https://phabricator.wikimedia.org/T437242) (owner: 10AOkoth) [13:49:59] (03Merged) 10jenkins-bot: mw-debug: fix staging images [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337927 (owner: 10Clément Goubert) [13:50:05] !log cgoubert@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply [13:51:48] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet [13:52:42] RESOLVED: [3x] ProbeDown: Service aqs2012-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:52:46] (03PS1) 10Muehlenhoff: Remove Java 8/Bullseye and Java 11 production images [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337929 (https://phabricator.wikimedia.org/T417389) [13:52:57] !log cgoubert@deploy1003 helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply [13:54:28] !log stran@deploy1003 mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # T435066 [13:54:31] T435066: Backfill PROPERTY_SHARED_PAGE_EDITS_COUNT values in cusi_case_property - https://phabricator.wikimedia.org/T435066 [13:54:52] done with everything [13:55:23] FIRING: GnmiInterfaceCountersDrop: ... [13:55:23] asw1-bw27-esams is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=asw1-bw27-esams:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:55:30] !log cgoubert@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply [13:56:07] !log cgoubert@deploy1003 helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply [13:57:44] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet [13:57:46] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet [13:57:50] (03PS2) 10Daniel Kertesz: admin: move dkertesz from ldap_only_users to users; add to ops group [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) [13:57:51] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet [13:57:54] (03CR) 10Clément Goubert: memcached: Clean up pki migration switch. (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1337558 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [13:58:07] !log cgoubert@deploy1003 helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply [13:58:32] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12297478 (10Fabfur) [13:59:25] (03PS9) 10Blake: memcached: Clean up pki migration switch. [puppet] - 10https://gerrit.wikimedia.org/r/1337558 (https://phabricator.wikimedia.org/T353511) [14:00:08] Deploy window Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1400) [14:00:39] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [14:00:48] (03CR) 10Majavah: [C:04-1] "The task has approval for `ops-limited` but this is adding to `ops`?" [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) (owner: 10Daniel Kertesz) [14:02:42] (03PS1) 10Muehlenhoff: Remove golang1.5 production image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337931 (https://phabricator.wikimedia.org/T416452) [14:02:53] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet [14:03:13] !log drain traffic from ssw1-a1-codfw before JunOS upgrade T426197 [14:03:13] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Degraded RAID on an-worker1200 - https://phabricator.wikimedia.org/T437170#12297506 (10Jclark-ctr) a:03Jclark-ctr Dell ticket opened SR231432096 Parts Arrive Wednesday, Sep 9, 5:00 p.m. [14:03:16] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:03:16] T426197: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197 [14:03:39] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12297513 (10Fabfur) [14:07:12] (03CR) 10Slyngshede: [C:03+2] C:mediawiki::tools::cache_warmup remove mobile urls [puppet] - 10https://gerrit.wikimedia.org/r/1337911 (https://phabricator.wikimedia.org/T436781) (owner: 10Slyngshede) [14:07:55] (03PS1) 10Muehlenhoff: Stop the build of Bullseye PHP base images [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337932 (https://phabricator.wikimedia.org/T417389) [14:08:25] (03CR) 10Blake: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337558 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:10:52] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06DC-Ops: Disk (sdf) failed in ms-be2089 - https://phabricator.wikimedia.org/T437293 (10MatthewVernon) 03NEW [14:11:33] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06DC-Ops: Disk (sdf) failed in ms-be2089 - https://phabricator.wikimedia.org/T437293#12297573 (10MatthewVernon) p:05Triage→03High [14:12:14] (03PS2) 10Hnowlan: kafka: migrate tls check to prometheus, per node [puppet] - 10https://gerrit.wikimedia.org/r/1333763 (https://phabricator.wikimedia.org/T407117) [14:12:34] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet [14:12:35] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet [14:12:41] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet [14:13:28] !log urbanecm@deploy1003 mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan 1788220800 --verbose # T437158 [14:13:31] T437158: Revalidate Add Link suggestions on enwiki - https://phabricator.wikimedia.org/T437158 [14:13:41] (03CR) 10Btullis: [C:03+2] ceph: Add a script to compare CephX keyrings with the cluster [puppet] - 10https://gerrit.wikimedia.org/r/1337875 (https://phabricator.wikimedia.org/T437233) (owner: 10Btullis) [14:15:30] (03CR) 10MVernon: "Tagging in Jaime as reviewer, as he's our backup guru :)" [puppet] - 10https://gerrit.wikimedia.org/r/1335952 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:15:55] (03CR) 10JMeybohm: [V:03+1 C:04-1] "I think it's a bit confusing to carry shim options and runsc flags in one structure but its probably an okay tradeoff. What I would like i" [puppet] - 10https://gerrit.wikimedia.org/r/1330129 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:16:36] (03CR) 10MVernon: "Tagging the DBAs rather than myself for review of mariadb stuff." [puppet] - 10https://gerrit.wikimedia.org/r/1335953 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:19:31] (03PS1) 10Majavah: mediawiki::tools: mediawiki-cache-warmup: Remove mobileServer support [puppet] - 10https://gerrit.wikimedia.org/r/1337933 [14:22:44] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet [14:22:45] (03CR) 10Klausman: [C:03+1] amd::gpu: Remove support for bullseye [puppet] - 10https://gerrit.wikimedia.org/r/1337885 (owner: 10Muehlenhoff) [14:23:01] (03CR) 10Brouberol: [C:03+1] Remove Java 8/Bullseye and Java 11 production images [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337929 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [14:23:16] (03PS1) 10Muehlenhoff: Stop building Node 12/14/16 images [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337934 (https://phabricator.wikimedia.org/T416452) [14:23:50] (03PS5) 10Clément Goubert: mediawiki: Redirect /api/ to /w/rest.php [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330505 (https://phabricator.wikimedia.org/T433547) [14:25:00] FIRING: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker2004 has a BGP session which is not in the 'established' state. [14:25:39] (03CR) 10JMeybohm: [V:03+1] "PCC SUCCESS (CORE_DIFF 13): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9381/c" [puppet] - 10https://gerrit.wikimedia.org/r/1330130 (owner: 10Giuseppe Lavagetto) [14:27:49] jouncebot: nowandnext [14:27:49] For the next 0 hour(s) and 2 minute(s): Test Kitchen UI Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1400) [14:27:49] In 0 hour(s) and 2 minute(s): Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1430) [14:28:13] Going to deploy a security pach [14:28:16] *patch [14:28:54] !log klausman@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet [14:28:55] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet [14:28:55] !log klausman@cumin1004 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad [14:28:57] (03CR) 10Clément Goubert: [C:03+1] "Doesn't look like it changes anything major on deployment-prep, lgtm" [puppet] - 10https://gerrit.wikimedia.org/r/1337558 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:29:38] (03CR) 10Blake: memcached: Clean up pki migration switch. (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1337558 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:30:04] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1430) [14:30:25] (03CR) 10Blake: [C:03+2] memcached: Clean up pki migration switch. [puppet] - 10https://gerrit.wikimedia.org/r/1337558 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:30:46] !log btullis@cumin1003 START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw [14:31:12] (03CR) 10Scott French: [C:03+1] "Thanks, Moritz!" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337932 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [14:31:45] (03PS3) 10Samtar: IS: enable wgEnableWatchstarPopover on test.wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1335032 (https://phabricator.wikimedia.org/T436955) [14:32:30] (03CR) 10Jsn.sherman: [C:03+1] "LGTM!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1335032 (https://phabricator.wikimedia.org/T436955) (owner: 10Samtar) [14:32:46] (03PS1) 10Scott French: scap.cfg.erb: Remove mediawiki_runtime_image pin [puppet] - 10https://gerrit.wikimedia.org/r/1337936 (https://phabricator.wikimedia.org/T418200) [14:33:26] !log btullis@cumin1003 END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw [14:34:20] !log btullis@cumin1003 START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet [14:34:35] (03CR) 10Muehlenhoff: "The titan* and cloudcontrol* hosts also use profile::memcached::instance, did you also check that they are migrated?" [puppet] - 10https://gerrit.wikimedia.org/r/1337558 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [14:35:00] (03CR) 10Marostegui: [C:03+1] "PCC looks good, so +1" [puppet] - 10https://gerrit.wikimedia.org/r/1335953 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [14:35:08] ^ blake: my Gerrit comment was a few seconds to late [14:35:13] oh shoot sorry [14:35:16] should i roll this back? [14:35:19] (03CR) 10Elukey: [C:03+1] "LGTM! We should also follow up with https://wikitech.wikimedia.org/wiki/Docker-registry#Deleting_images (the blob layers won't get dropped" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337931 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [14:35:26] (03CR) 10JMeybohm: "I would suggest to enable this in staging-codfw only first (so `hieradata/role/codfw/kubernetes/staging/worker.yaml`), verify everything t" [puppet] - 10https://gerrit.wikimedia.org/r/1330131 (https://phabricator.wikimedia.org/T435796) (owner: 10Giuseppe Lavagetto) [14:36:00] (03PS1) 10Blake: Revert "memcached: Clean up pki migration switch." [puppet] - 10https://gerrit.wikimedia.org/r/1337937 [14:36:12] moritzm: if you don't mind, ^ [14:36:13] If you haven't puppet-merged, maybe roll back so that we can check this safely in PCC first [14:36:17] looking [14:36:31] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1337937 (owner: 10Blake) [14:36:34] +1d [14:36:37] ty [14:36:58] i had puppet merged, so i'll get this in asap, and then explore the titan and cloudcontrol hosts [14:37:01] my apologies! [14:37:11] PROBLEM - BFD status on lsw1-a7-codfw.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [14:37:34] I think one of the gpg keys in pwstore expired? (can't create a new commit, it complains about the key for klausman being invalid) [14:38:29] dcaro: I'd ping moritzm or bring that to -sre which is less likely to get lost [14:38:30] (03CR) 10Blake: [C:03+2] Revert "memcached: Clean up pki migration switch." [puppet] - 10https://gerrit.wikimedia.org/r/1337937 (owner: 10Blake) [14:38:34] yeah, I think EC3D2B2DAC6964AFC7134FB69B91773F42CB71E3 expired. Having a look [14:39:10] pub rsa4096 2020-09-02 [SC] [expired: 2025-09-28] Huh. Almost a year ago, how did I not notice? [14:39:10] FIRING: [2x] BFDdown: BFD session down between lsw1-a7-codfw and 10.192.9.17 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=lsw1-a7-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:39:56] dcaro, klausman: make sure to update your git remotes, the repo on cumin1003 is disabled to prevent people pushing to the old server [14:40:22] I'm using cumin1004, that's the good one right? [14:40:33] dcaro: yes! [14:40:53] (03CR) 10Btullis: [C:03+2] Add two new dse-k8s-worker nodes to conftool-data in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1337629 (https://phabricator.wikimedia.org/T432206) (owner: 10Btullis) [14:41:09] (03CR) 10Btullis: [C:03+2] Add the new dse-k8s-worker nodes to the cluster in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1337630 (https://phabricator.wikimedia.org/T432206) (owner: 10Btullis) [14:41:14] moritzm: fixed my repo and pulled from 1004. Still got an expired key, though [14:41:27] (03CR) 10Marostegui: [C:03+2] mariadb.yaml: Add x4 [puppet] - 10https://gerrit.wikimedia.org/r/1337894 (https://phabricator.wikimedia.org/T437229) (owner: 10Marostegui) [14:42:11] RECOVERY - BFD status on lsw1-a7-codfw.mgmt is OK: UP: 4 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [14:43:08] dcaro: I can-reecrypt just fine, did you run "pws update-keyring"? [14:43:20] !log cmooney@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad [14:43:30] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12297757 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=2feb17c1-1370-41ec-bccd-8b3aadc69c13) set by cmooney@cumin1004 for 1:00:00 o... [14:44:10] RESOLVED: [2x] BFDdown: BFD session down between lsw1-a7-codfw and 10.192.9.17 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=lsw1-a7-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [14:44:21] !log shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw [14:44:23] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [14:44:55] Ok, renewed my key and pushed it to https://keys.openpgp.org/. Hope that's enough [14:45:11] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet [14:45:31] moritzm: I did, klausman let me refetch from the server then [14:46:09] !log dreamyjazz Deployed security patch for T436429 [14:46:11] hmm, did not help [14:46:50] oh, I deleted it from my keyring and re-imported from the server, and it worked now [14:46:50] \o/ [14:47:15] (03CR) 10Marostegui: [C:03+1] "Thank you" [alerts] - 10https://gerrit.wikimedia.org/r/1337633 (https://phabricator.wikimedia.org/T436049) (owner: 10Hnowlan) [14:47:24] not sure if it was just cache on the server side (it was importing it, but saying that it was unchanged) or something [14:47:29] klausman: moritzm thanks! [14:47:40] np, thanks for reminding me of an expire key :) [14:48:14] dcaro, klausman: perfect :-) [14:48:53] dcaro: wait, which server are you referring to? the keys are stored within the repo in the keys/ sub directory since a few years [14:49:16] (03PS23) 10Herron: sre.opensearch.roll-restart-reboot: include checklist items [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) [14:49:16] moritzm: I imported it from keys.openpgp.org [14:49:23] (as klausman pushed it there) [14:49:58] that's not what we are using and doesn't work, keys.opengpg.org doesn't even store signatures [14:50:13] I patched pwstore to read the keys in use from keys/ within the repos [14:50:14] The wikitech page says to publish the key, you send it to the keyserver, but it does not mention the pw repo [14:50:14] interesting [14:50:15] I patched pwstore to read the keys in use from keys/ within the repo [14:50:39] https://office.wikimedia.org/wiki/Pwstore#Updating_your_own_key [14:51:00] klausman: where are these? I can update or delete them [14:51:13] (03PS1) 10Zabe: Use local database for category table in SpecialWantedCategories [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337946 (https://phabricator.wikimedia.org/T437286) [14:51:15] importing from both ends up getting the same key (deleting in between to make sure) [14:51:28] (03PS1) 10Zabe: Use local database for category table in SpecialWantedCategories [core] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337947 (https://phabricator.wikimedia.org/T437286) [14:51:28] https://www.irccloud.com/pastebin/StuxoNgB/ [14:51:37] (03CR) 10Herron: sre.opensearch.roll-restart-reboot: include checklist items (038 comments) [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [14:51:37] and make sure to get the patched script either manually https://office.wikimedia.org/wiki/Pwstore#Installation or otherwise via wmf-laptop [14:52:09] (03CR) 10CI reject: [V:04-1] sre.opensearch.roll-restart-reboot: include checklist items [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) (owner: 10Herron) [14:52:11] aha! exported, committed, pushed [14:53:01] (03PS24) 10Herron: sre.opensearch.roll-restart-reboot: include checklist items [cookbooks] - 10https://gerrit.wikimedia.org/r/1334048 (https://phabricator.wikimedia.org/T435265) [14:53:36] (03CR) 10JMeybohm: [V:03+1 C:04-1] "I don't think this works as intended, it will enable dragonfly in places where it should not." [puppet] - 10https://gerrit.wikimedia.org/r/1330130 (owner: 10Giuseppe Lavagetto) [14:55:09] !log dreamyjazz Deployed security patch for T436429 [14:55:46] moritzm: https://wikitech.wikimedia.org/wiki/PGP_Keys#Renewing_keys [14:56:41] updated the script to the latest version just in case xd, but it was working already [14:57:03] PROBLEM - Host lsw1-b2-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:57:04] FIRING: MediaWikiElevatedUnknownLogins: Elevated number of login successes (source unknown) via mw-web - TODO - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?from=now-6h&orgId=1&to=now&viewPanel=26 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiElevatedUnknownLogins [14:57:07] PROBLEM - Host lsw1-b7-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:57:07] PROBLEM - Host lsw1-b4-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:57:07] PROBLEM - Host lsw1-b3-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:57:07] PROBLEM - Host lsw1-b6-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:57:07] PROBLEM - Host lsw1-b5-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:57:17] PROBLEM - Host lsw1-b8-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:57:27] moritzm: As for the pws instructions, even when using the one from wmf-laptop, I am getting an error about `/home/klausman/.pws-trusted-users` not existing [14:57:27] PROBLEM - Host mr1-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:57:39] PROBLEM - Host rpki2003 is DOWN: PING CRITICAL - Packet loss = 100% [14:57:39] ok, but these are generic PGP instructions, they can stay as-is [14:57:59] klausman: https://office.wikimedia.org/wiki/Pwstore#User_database [14:58:05] (03PS2) 10Blake: memcached: Clean up pki migration switch. [puppet] - 10https://gerrit.wikimedia.org/r/1337942 (https://phabricator.wikimedia.org/T353511) [14:58:10] oops, totally missed that [14:58:15] PROBLEM - Host lsw1-a8-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:15] PROBLEM - Host lsw1-a8-codfw is DOWN: PING CRITICAL - Packet loss = 100% [14:58:15] klausman: yep, you have to manually create it [14:58:19] PROBLEM - Host lsw1-b4-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:19] PROBLEM - Host lsw1-b2-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:19] PROBLEM - Host lsw1-b7-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:19] PROBLEM - Host lsw1-b3-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:19] PROBLEM - Host lsw1-b5-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:20] PROBLEM - Host lsw1-b8-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:20] PROBLEM - Host lsw1-b6-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [14:58:32] uh? [14:58:39] PROBLEM - Host mr1-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [15:00:05] jelto, arnoldokoth, mutante, and arnaudb: OwO what's this, a deployment window?? SRE Collaboration Services office hours. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1500). nyaa~ [15:00:07] PROBLEM - MariaDB Replica IO: ms1 on db1267 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@db2251.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on db2251.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:00:07] PROBLEM - MariaDB Replica IO: pc2 on pc1022 is CRITICAL: CRITICAL slave_io_state Slave_IO_Running: No, Errno: 2003, Errmsg: error reconnecting to master repl2024@pc2022.codfw.wmnet:3306 - retry-time: 60 maximum-retries: 100000 message: Cant connect to server on pc2022.codfw.wmnet (110 Connection timed out) https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:00:18] topranks: ^ ? [15:00:35] bblack: lsw are me yes (downtime seems not to have worked) [15:00:43] ack [15:00:48] (03CR) 10Blake: "The cloudcontrol hosts explicitly disable tls, and are therefore unaffected by this change:" [puppet] - 10https://gerrit.wikimedia.org/r/1337942 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [15:00:53] mr1-codfw.... possibly me, the db replica I don't think so [15:01:21] (03CR) 10Muehlenhoff: "Thanks for the pointer, I will do that tomorrow after merging the various patches" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337931 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [15:01:33] the db replicas are eqiad boxes complaining about connecting to codfw boxes, possibly-related [15:02:04] RESOLVED: MediaWikiElevatedUnknownLogins: Elevated number of login successes (source unknown) via mw-web - TODO - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?from=now-6h&orgId=1&to=now&viewPanel=26 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiElevatedUnknownLogins [15:02:25] FIRING: SystemdUnitFailed: netbox_ganeti_codfw02_sync.service on netbox1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:02:52] ^ ditto that [15:03:58] FIRING: NELHigh: Elevated Network Error Logging events (tcp.timed_out) #page - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELHigh [15:04:09] RECOVERY - Host lsw1-b4-codfw is UP: PING OK - Packet loss = 0%, RTA = 34.81 ms [15:04:09] RECOVERY - Host lsw1-b2-codfw is UP: PING OK - Packet loss = 0%, RTA = 37.69 ms [15:04:09] RECOVERY - Host lsw1-b5-codfw is UP: PING OK - Packet loss = 0%, RTA = 39.17 ms [15:04:09] RECOVERY - Host lsw1-b3-codfw is UP: PING OK - Packet loss = 0%, RTA = 36.15 ms [15:04:09] RECOVERY - Host lsw1-b7-codfw is UP: PING OK - Packet loss = 0%, RTA = 47.13 ms [15:04:10] RECOVERY - Host lsw1-b6-codfw is UP: PING OK - Packet loss = 0%, RTA = 38.43 ms [15:04:11] RECOVERY - Host lsw1-b8-codfw is UP: PING OK - Packet loss = 0%, RTA = 34.43 ms [15:04:23] RECOVERY - Host mr1-codfw is UP: PING OK - Packet loss = 0%, RTA = 36.11 ms [15:04:45] RECOVERY - Host lsw1-a8-codfw is UP: PING OK - Packet loss = 0%, RTA = 38.67 ms [15:04:55] RECOVERY - Host rpki2003 is UP: PING OK - Packet loss = 0%, RTA = 32.02 ms [15:06:05] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:06:07] RECOVERY - MariaDB Replica IO: pc2 on pc1022 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:06:07] RECOVERY - MariaDB Replica IO: ms1 on db1267 is OK: OK slave_io_state Slave_IO_Running: Yes https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [15:07:01] !incidentgs [15:07:03] !incidents [15:07:04] 8328 (ACKED) NELHigh sre (thanos-rule@main tcp.timed_out) [15:08:01] PROBLEM - Check unit status of httpbb_kubernetes_mw-wikifunctions_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-wikifunctions_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:08:29] RECOVERY - Host lsw1-a8-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 74.16 ms [15:08:33] RECOVERY - Host lsw1-b2-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 32.00 ms [15:08:33] RECOVERY - Host lsw1-b4-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 31.88 ms [15:08:33] RECOVERY - Host lsw1-b7-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 31.90 ms [15:08:33] RECOVERY - Host lsw1-b3-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 31.90 ms [15:08:33] RECOVERY - Host lsw1-b5-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 35.47 ms [15:08:34] RECOVERY - Host lsw1-b6-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 37.30 ms [15:08:34] RECOVERY - Host lsw1-b8-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 42.34 ms [15:08:53] RECOVERY - Host mr1-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 32.14 ms [15:09:01] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:09:01] PROBLEM - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:10:03] FIRING: [2x] SystemdUnitFailed: rsync-srv_firmwares.service on cumin2003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:10:13] FIRING: [4x] CoreBGPDown: Core BGP session down between ssw1-e1-codfw and ssw1-a1-codfw (10.192.253.162) - group core - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:10:22] FIRING: [3x] SwitchCoreInterfaceDown: Switch core interface down - ssw1-d8-codfw:et-0/0/29 (Core: ssw1-a1-codfw:et-0/0/29) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:10:32] RESOLVED: NELHigh: Elevated Network Error Logging events (tcp.timed_out) #page - https://wikitech.wikimedia.org/wiki/Network_monitoring#NEL_alerts - https://logstash.wikimedia.org/goto/5c8f4ca1413eda33128e5c5a35da7e28 - https://alerts.wikimedia.org/?q=alertname%3DNELHigh [15:10:59] (03CR) 10JMeybohm: [C:04-1] admin: add support for gVisor RuntimeClass handlers (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330344 (https://phabricator.wikimedia.org/T436212) (owner: 10Giuseppe Lavagetto) [15:11:03] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:12:19] 06SRE, 06Infrastructure-Foundations, 06Release-Engineering-Team (Radar): apt-get broken in docker-registry.wikimedia.org/bullseye:20260830 - https://phabricator.wikimedia.org/T437069#12297979 (10dancy) >>! In T437069#12294166, @MoritzMuehlenhoff wrote: > bullseye-security was restored on the Debian archive s... [15:13:01] PROBLEM - Check unit status of httpbb_kubernetes_mw-jobrunner_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-jobrunner_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:13:47] PROBLEM - BFD status on ssw1-a8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:13:47] PROBLEM - BFD status on ssw1-d1-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:13:47] PROBLEM - BFD status on ssw1-d8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:14:03] PROBLEM - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:14:36] (03CR) 10Elukey: [C:03+1] "LGTM, let's drop these as well via https://wikitech.wikimedia.org/wiki/Docker-registry#Deleting_images" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337929 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [15:15:24] (03CR) 10Clément Goubert: [C:03+1] memcached: Clean up pki migration switch. [puppet] - 10https://gerrit.wikimedia.org/r/1337942 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [15:15:33] FIRING: [2x] SystemdUnitFailed: rsync-srv_firmwares.service on cumin2003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:15:44] FIRING: BFDdown: BFD session down between lsw1-a4-codfw and 10.192.252.1 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=lsw1-a4-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:15:49] FIRING: [5x] CoreBGPDown: Core BGP session down between lsw1-a3-codfw and ssw1-a1-codfw (10.192.252.1) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:15:53] FIRING: [4x] SwitchCoreInterfaceDown: Switch core interface down - lsw1-a3-codfw:et-0/0/55 (Core: ssw1-a1-codfw:et-0/0/2) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:16:50] (03CR) 10Kamila Součková: [C:03+1] "oops, thank you for catching it!" [puppet] - 10https://gerrit.wikimedia.org/r/1337936 (https://phabricator.wikimedia.org/T418200) (owner: 10Scott French) [15:17:25] RESOLVED: [2x] SystemdUnitFailed: rsync-srv_firmwares.service on cumin2003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:18:10] FIRING: [14x] BFDdown: BFD session down between lsw1-a2-codfw and 10.192.252.1 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [15:18:39] FIRING: [20x] CoreBGPDown: Core BGP session down between lsw1-a2-codfw and ssw1-a1-codfw (10.192.252.1) - group EVPN_IBGP - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:18:51] FIRING: [17x] SwitchCoreInterfaceDown: Switch core interface down - lsw1-a2-codfw:et-0/0/55 (Core: ssw1-a1-codfw:et-0/0/1) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [15:19:41] !log btullis@cumin1003 START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet [15:21:34] 06SRE, 06Infrastructure-Foundations, 06Release-Engineering-Team (Radar): apt-get broken in docker-registry.wikimedia.org/bullseye:20260830 - https://phabricator.wikimedia.org/T437069#12298009 (10MoritzMuehlenhoff) This is now reported as https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1147093 [15:22:45] PROBLEM - BFD status on lsw1-c2-codfw.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:26:45] RECOVERY - BFD status on lsw1-c2-codfw.mgmt is OK: UP: 4 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:26:56] (03PS1) 10JMeybohm: vap: Fix variable naming for validationActions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337953 (https://phabricator.wikimedia.org/T436380) [15:28:05] (03PS2) 10JMeybohm: vap: Fix variable naming for validationActions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337953 (https://phabricator.wikimedia.org/T436380) [15:29:21] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet [15:30:37] (03CR) 10Clément Goubert: [C:03+1] "titan and cloudcontrol don't seem to override `profile::memcached::enable_tls` and it defaults to `false` in `hieradata/common/profile/mem" [puppet] - 10https://gerrit.wikimedia.org/r/1337558 (https://phabricator.wikimedia.org/T353511) (owner: 10Blake) [15:30:41] (03CR) 10AOkoth: [C:03+2] phabricator: add skip-ssl option to client config [puppet] - 10https://gerrit.wikimedia.org/r/1337657 (https://phabricator.wikimedia.org/T437242) (owner: 10AOkoth) [15:33:42] (03PS3) 10Slyngshede: site.pp move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1331590 (https://phabricator.wikimedia.org/T436363) [15:35:39] (03PS4) 10Slyngshede: site.pp move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1331590 (https://phabricator.wikimedia.org/T436363) [15:37:00] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in role mariadb [puppet] - 10https://gerrit.wikimedia.org/r/1335953 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:37:02] (03CR) 10CDanis: [C:03+2] Deduplicate the proxy configuration value [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328403 (owner: 10Brouberol) [15:37:06] (03CR) 10CDanis: [C:03+2] "thanks!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328403 (owner: 10Brouberol) [15:37:13] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module query_service [puppet] - 10https://gerrit.wikimedia.org/r/1335930 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:37:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.14% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:37:44] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module phabricator [puppet] - 10https://gerrit.wikimedia.org/r/1335928 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:38:15] (03CR) 10Ssingh: "@taavi@wikimedia.org: Any thoughts on this please? We want to move ahead with the change." [puppet] - 10https://gerrit.wikimedia.org/r/1332813 (owner: 10BCornwall) [15:39:32] (03Merged) 10jenkins-bot: Deduplicate the proxy configuration value [deployment-charts] - 10https://gerrit.wikimedia.org/r/1328403 (owner: 10Brouberol) [15:42:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.38% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:42:59] PROBLEM - BFD status on lsw1-b8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:13] PROBLEM - BFD status on lsw1-b6-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:13] PROBLEM - BFD status on lsw1-a8-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:13] PROBLEM - BFD status on lsw1-a7-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:39] FIRING: [45x] CoreBGPDown: Core BGP session down between cr1-codfw and ssw1-a1-codfw (10.192.254.5) - group Switch - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:43:41] PROBLEM - BFD status on lsw1-a5-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:43] PROBLEM - BFD status on lsw1-b2-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:45] PROBLEM - BFD status on lsw1-a2-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:45] PROBLEM - BFD status on lsw1-b4-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:45] PROBLEM - BFD status on lsw1-b7-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:45] PROBLEM - BFD status on lsw1-a6-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:45] PROBLEM - BFD status on lsw1-a3-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:46] PROBLEM - BFD status on lsw1-b3-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:46] PROBLEM - BFD status on lsw1-b5-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:43:47] PROBLEM - BFD status on lsw1-a4-codfw.mgmt is CRITICAL: Down: 1 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:44:00] !log jhancock@cumin2003 START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART [15:44:33] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module mailman3 [puppet] - 10https://gerrit.wikimedia.org/r/1335927 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:44:52] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module openstack [puppet] - 10https://gerrit.wikimedia.org/r/1335917 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:44:52] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/1/5 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [15:45:17] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in module galera [puppet] - 10https://gerrit.wikimedia.org/r/1335901 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:45:41] PROBLEM - Host ssw1-a1-codfw.mgmt is DOWN: PING CRITICAL - Packet loss = 100% [15:46:33] PROBLEM - Host ssw1-a1-codfw is DOWN: PING CRITICAL - Packet loss = 100% [15:46:33] PROBLEM - Host ssw1-a1-codfw IPv6 is DOWN: PING CRITICAL - Packet loss = 100% [15:47:38] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace unscoped legacy facts in profile wmcs [puppet] - 10https://gerrit.wikimedia.org/r/1335918 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:48:39] FIRING: [45x] CoreBGPDown: Core BGP session down between cr1-codfw and ssw1-a1-codfw (10.192.254.5) - group Switch - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [15:50:24] FIRING: [10x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [15:54:26] !log jhancock@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART [15:54:31] RECOVERY - Host ssw1-a1-codfw.mgmt is UP: PING OK - Packet loss = 0%, RTA = 33.38 ms [15:55:24] FIRING: [10x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [15:55:26] !log btullis@cumin1003 START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet [15:57:45] PROBLEM - BFD status on lsw1-d2-codfw.mgmt is CRITICAL: Down: 2 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:58:01] RECOVERY - Check unit status of httpbb_kubernetes_mw-wikifunctions_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-wikifunctions_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:58:12] !log jhancock@cumin2003 START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART [15:58:59] RECOVERY - BFD status on lsw1-b8-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:01] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-int_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-int_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:59:01] RECOVERY - Check unit status of httpbb_kubernetes_mw-web-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [15:59:13] RECOVERY - BFD status on lsw1-b6-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:13] RECOVERY - BFD status on lsw1-a8-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:13] RECOVERY - BFD status on lsw1-a7-codfw.mgmt is OK: UP: 4 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:41] RECOVERY - BFD status on lsw1-a5-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:43] RECOVERY - BFD status on lsw1-b2-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:45] RECOVERY - BFD status on lsw1-a3-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:45] RECOVERY - BFD status on lsw1-b4-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:45] RECOVERY - BFD status on lsw1-b3-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:45] RECOVERY - BFD status on lsw1-b7-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:45] RECOVERY - BFD status on lsw1-b5-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:46] RECOVERY - BFD status on lsw1-a6-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:46] RECOVERY - BFD status on lsw1-a4-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:47] RECOVERY - BFD status on lsw1-a2-codfw.mgmt is OK: UP: 2 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:49] RECOVERY - BFD status on ssw1-a8-codfw.mgmt is OK: UP: 17 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:49] RECOVERY - BFD status on ssw1-d1-codfw.mgmt is OK: UP: 18 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [15:59:49] RECOVERY - BFD status on ssw1-d8-codfw.mgmt is OK: UP: 18 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [16:00:05] jhathaway and rzl: #bothumor I � Unicode. All rise for Puppet request window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1600). [16:00:05] No Gerrit patches in the queue for this window AFAICS. [16:00:51] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12298420 (10ssingh) [16:01:03] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:01:14] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12298424 (10ssingh) Approval for just ops is fine, since the onboarding checklist specifies a waiting period of a week, which we have met. [16:01:45] RECOVERY - BFD status on lsw1-d2-codfw.mgmt is OK: UP: 4 AdminDown: 0 Down: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23BFD_status [16:01:59] RECOVERY - Host ssw1-a1-codfw IPv6 is UP: PING OK - Packet loss = 0%, RTA = 30.95 ms [16:01:59] RECOVERY - Host ssw1-a1-codfw is UP: PING OK - Packet loss = 0%, RTA = 36.77 ms [16:03:01] RECOVERY - Check unit status of httpbb_kubernetes_mw-jobrunner_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-jobrunner_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:03:10] RESOLVED: [14x] BFDdown: BFD session down between lsw1-a2-codfw and 10.192.252.1 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [16:03:39] RESOLVED: [47x] CoreBGPDown: Core BGP session down between cr1-codfw and ssw1-a1-codfw (10.192.254.5) - group Switch - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [16:03:51] RESOLVED: [17x] SwitchCoreInterfaceDown: Switch core interface down - lsw1-a2-codfw:et-0/0/55 (Core: ssw1-a1-codfw:et-0/0/1) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Switch_interface_down - https://alerts.wikimedia.org/?q=alertname%3DSwitchCoreInterfaceDown [16:04:03] RECOVERY - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:04:49] !log btullis@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet [16:04:52] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr1-codfw:et-1/1/5 () - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [16:05:31] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [16:05:40] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply [16:06:05] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [16:06:51] (03PS1) 10Majavah: P:wmcs::novaproxy: Copy more timeout settings from Toolforge config [puppet] - 10https://gerrit.wikimedia.org/r/1337960 [16:08:41] !log jhancock@cumin2003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART [16:08:47] (03PS1) 10Hashar: tests: Remove newline after header from MassMessageJob [extensions/MassMessage] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337961 [16:09:16] (03CR) 10David Caro: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1337960 (owner: 10Majavah) [16:09:33] (03CR) 10Majavah: [C:03+2] P:wmcs::novaproxy: Copy more timeout settings from Toolforge config [puppet] - 10https://gerrit.wikimedia.org/r/1337960 (owner: 10Majavah) [16:11:42] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:12:28] !log jhancock@cumin2003 START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie [16:12:38] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12298525 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jhancock@cumin2003 for host sretest2013.codfw.wmnet with OS trixie [16:12:41] (03CR) 10Jelto: [C:03+1] "lgtm" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337953 (https://phabricator.wikimedia.org/T436380) (owner: 10JMeybohm) [16:13:57] FIRING: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:16:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:18:57] RESOLVED: ProbeDown: Service text-https:443 has failed probes (http_text-https_ip6) - https://wikitech.wikimedia.org/wiki/Runbook#text-https:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [16:24:44] !log btullis@puppetserver1001 conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet [16:24:49] !log btullis@puppetserver1001 conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet [16:25:01] !log btullis@puppetserver1001 conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet [16:25:06] !log btullis@puppetserver1001 conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet [16:25:08] !log jhancock@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage [16:28:15] jouncebot: nowandnext [16:28:15] For the next 0 hour(s) and 31 minute(s): Puppet request window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1600) [16:28:15] In 0 hour(s) and 31 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1700) [16:28:22] (03CR) 10Zabe: [C:03+2] Use local database for category table in SpecialWantedCategories [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337946 (https://phabricator.wikimedia.org/T437286) (owner: 10Zabe) [16:28:23] (03CR) 10Zabe: [C:03+2] Use local database for category table in SpecialWantedCategories [core] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337947 (https://phabricator.wikimedia.org/T437286) (owner: 10Zabe) [16:28:41] (03CR) 10Jasmine: [C:03+1] "Thanks Scott! I agree. Pasting some of the previous weeks estimates here too, which suggest the projection here is a good starting point." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332725 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [16:29:18] !log jhancock@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage [16:34:27] (03Abandoned) 10Btullis: Update the SSH configuration to add the keys to the agent on first use [debs/wmf-laptop] - 10https://gerrit.wikimedia.org/r/891834 (owner: 10Btullis) [16:34:41] (03CR) 10Btullis: [C:03+2] Promote dse-k8s-worker200[4-5] to their correct role [puppet] - 10https://gerrit.wikimedia.org/r/1337631 (https://phabricator.wikimedia.org/T432206) (owner: 10Btullis) [16:37:00] zabe: Ping to make sure you noticed my message on https://phabricator.wikimedia.org/T437108 [16:38:02] yes, I will take a look after deploying the above fix:) [16:38:09] Great. thanks! [16:38:33] (03CR) 10Ssingh: [C:03+1] "Needs approval of Mark / Kavitha, so let's wait for that before merging." [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) (owner: 10Daniel Kertesz) [16:38:53] (03Merged) 10jenkins-bot: Use local database for category table in SpecialWantedCategories [core] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337946 (https://phabricator.wikimedia.org/T437286) (owner: 10Zabe) [16:38:59] (03CR) 10Ssingh: [C:03+1] "Group has been changed to ops, in the task as well. Per the onboarding docs, a week's wait is fine and we have reached that." [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) (owner: 10Daniel Kertesz) [16:39:01] (03Merged) 10jenkins-bot: Use local database for category table in SpecialWantedCategories [core] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1337947 (https://phabricator.wikimedia.org/T437286) (owner: 10Zabe) [16:41:12] !log zabe@deploy1003 Started scap sync-world: Backport for [[gerrit:1337946|Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947|Use local database for category table in SpecialWantedCategories (T437286)]] [16:41:16] T437286: Wikimedia\Rdbms\DBQueryError: Error 1146: Table 'commonswiki.category' doesn't existFunction: MediaWiki\Specials\SpecialWantedCategories::preprocessResultsQuery: SELECT cat_title,cat_pages FROM `category` WHERE cat_title - https://phabricator.wikimedia.org/T437286 [16:42:42] (03CR) 10Ssingh: [C:03+1] site.pp move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1331590 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [16:42:54] (03CR) 10Majavah: "Is there some additional context for this? The defines mentioned here will not work on the `ensure => present` case either, and I am not a" [puppet] - 10https://gerrit.wikimedia.org/r/1332813 (owner: 10BCornwall) [16:45:44] !log zabe@deploy1003 zabe: Backport for [[gerrit:1337946|Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947|Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [16:46:53] !log zabe@deploy1003 zabe: Continuing with deployment [16:47:12] (03CR) 10Ssingh: [C:03+1] "Please don't merge until @mark@wikimedia.org 's approval" [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) (owner: 10Daniel Kertesz) [16:47:17] (03CR) 10BCornwall: "I believe that throwing an error on `present` when an interfaces file is missing is proper." [puppet] - 10https://gerrit.wikimedia.org/r/1332813 (owner: 10BCornwall) [16:51:32] !log zabe@deploy1003 Finished scap sync-world: Backport for [[gerrit:1337946|Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947|Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s) [16:51:36] T437286: Wikimedia\Rdbms\DBQueryError: Error 1146: Table 'commonswiki.category' doesn't existFunction: MediaWiki\Specials\SpecialWantedCategories::preprocessResultsQuery: SELECT cat_title,cat_pages FROM `category` WHERE cat_title - https://phabricator.wikimedia.org/T437286 [16:51:53] (03PS1) 10Zabe: Do not try to use x4 in beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337969 (https://phabricator.wikimedia.org/T437108) [16:53:27] (03CR) 10Zabe: [C:03+2] Do not try to use x4 in beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337969 (https://phabricator.wikimedia.org/T437108) (owner: 10Zabe) [16:54:23] (03Merged) 10jenkins-bot: Do not try to use x4 in beta cluster [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337969 (https://phabricator.wikimedia.org/T437108) (owner: 10Zabe) [16:54:35] (03PS3) 10Scott French: rest-gateway: Adopt cluster specifier plugin and XWD routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1335285 (https://phabricator.wikimedia.org/T433752) [16:55:00] FIRING: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker2004 has a BGP session which is not in the 'established' state. [16:57:36] (03PS1) 10Majavah: interface::ipip: Guard absenting an address behind ifupdown fact [puppet] - 10https://gerrit.wikimedia.org/r/1337971 [16:57:52] (03CR) 10Scott French: "Thanks in advance for the review. Let me know if you'd like me to go about this in a different way - e.g., incrementally across routes. Se" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1335285 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [16:59:23] (03CR) 10Majavah: "What do you think of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1337971? (assuming `interface::ipip` is the problem usage here, " [puppet] - 10https://gerrit.wikimedia.org/r/1332813 (owner: 10BCornwall) [17:00:00] RESOLVED: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker2004 has a BGP session which is not in the 'established' state. [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1700) [17:00:15] (03CR) 10Scott French: [C:03+2] scap.cfg.erb: Remove mediawiki_runtime_image pin [puppet] - 10https://gerrit.wikimedia.org/r/1337936 (https://phabricator.wikimedia.org/T418200) (owner: 10Scott French) [17:01:14] o/ [17:02:10] once puppet-agent on deploy1003 completes, I'll run a quick smoke test with scap [17:02:15] (03PS1) 10Jforrester: Periodic jobs: Add mainstash_metrics [puppet] - 10https://gerrit.wikimedia.org/r/1334836 (https://phabricator.wikimedia.org/T430940) [17:02:24] (03PS2) 10Jforrester: Periodic jobs: Add mainstash_metrics [puppet] - 10https://gerrit.wikimedia.org/r/1334836 (https://phabricator.wikimedia.org/T430940) [17:04:50] dancy: Could you tell me where the beta cluster update job is nowadays? https://integration.wikimedia.org/ci/view/Beta/job/beta-code-update-eqiad/ says "Not found." [17:05:00] zabe: it is a systemd timer [17:05:17] Is it running on the beta cluster deploy host? [17:05:21] yes [17:06:36] zabe: called wmf-beta-update-all.timer [17:07:03] Thanks! [17:09:35] Zabe: https://beta-update.wmcloud.org/ [17:09:40] or that, if one is for logs [17:09:44] (for logs) [17:09:58] kind of depends what zabe wanted to do :D [17:10:54] !log swfrench@deploy1003 Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup [17:11:00] https://beta-update.wmcloud.org/202609081700.log looks like it is working again [17:11:55] !log dropping links tables from db1247 (s4 replica) - (T437278) [17:11:57] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [17:11:57] T437278: Drop unneeded tables from x4 and s4 - https://phabricator.wikimedia.org/T437278 [17:15:10] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:15:12] !log swfrench@deploy1003 Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s) [17:16:01] * swfrench-wmf is done with what he had planned [17:16:46] Thanks Zabe! [17:17:17] zabe: since you're not busy, I just dropped the tables from db1247, so if you see fatals, that's why :P [17:17:37] 10ops-codfw, 06SRE, 06DC-Ops: Degraded RAID on maps-test2001 - https://phabricator.wikimedia.org/T437082#12299092 (10Jhancock.wm) 05Open→03Declined duplicate [17:17:50] Amir1: did you also drop linktarget and collation? [17:18:07] linktarget yes, collation no. Should I? [17:18:15] I wasn't sure if you backported it [17:19:21] I did, and for linktarget I just overlooked it :) [17:19:42] I do not mind to much, but maybe just for the sake of completness [17:21:29] dropped collation too [17:22:00] linktarget is using the virtual domains, otherwise things would have exploded majestically [17:22:27] https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&refresh=5m&var-server=db1247&var-datasource=000000026&var-cluster=mysql&viewPanel=panel-28&from=2026-09-08T16:54:34.359Z&to=2026-09-08T17:22:11.470Z&timezone=utc [17:22:40] s4 is now 1.03TB smaller, great job zabe. Seriously [17:23:31] Amir1: do you mind if I quickly reboot the deployment host, or are you using it? [17:23:44] 🤯 that's a fair drop :D [17:23:49] Raine: I'm not using it, zabe might be though [17:23:57] Dreamy_Jazz: it's half of its size now [17:24:11] I am not using it :) [17:24:12] (excuse my DBA) fucking half [17:24:21] :D [17:24:25] :D [17:24:38] ok, thanks Amir1 and zabe, off I go [17:25:32] !log kamila@cumin1003 START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet [17:30:38] (03PS1) 10Bking: w[cd]qs: Auto-clean blazegraph tmpfiles [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) [17:30:51] !log kamila@cumin1003 START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet [17:31:06] (03PS1) 10Ssingh: acme_chief: add sretest2013 to unified cert host [puppet] - 10https://gerrit.wikimedia.org/r/1337979 (https://phabricator.wikimedia.org/T436691) [17:31:07] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [17:33:06] (03CR) 10CI reject: [V:04-1] w[cd]qs: Auto-clean blazegraph tmpfiles [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [17:33:28] (03PS2) 10Ssingh: acme_chief: add sretest2013 to unified cert host [puppet] - 10https://gerrit.wikimedia.org/r/1337979 (https://phabricator.wikimedia.org/T436691) [17:34:16] (03PS3) 10Ssingh: acme_chief: add sretest2013 to unified cert host [puppet] - 10https://gerrit.wikimedia.org/r/1337979 (https://phabricator.wikimedia.org/T436691) [17:35:04] (03CR) 10JMeybohm: [C:03+2] vap: Fix variable naming for validationActions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337953 (https://phabricator.wikimedia.org/T436380) (owner: 10JMeybohm) [17:35:48] (03CR) 10JMeybohm: [C:04-1] "For some reason testing the bad-pod test was initially flaky for me. Adding a short sleep between applying the runtimeclass.yaml and start" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333210 (https://phabricator.wikimedia.org/T436655) (owner: 10Giuseppe Lavagetto) [17:36:30] !log cdobbins@cumin1003 END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox [17:36:44] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet [17:36:49] (03CR) 10RLazarus: Periodic jobs: Add mainstash_metrics (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1334836 (https://phabricator.wikimedia.org/T430940) (owner: 10Jforrester) [17:40:13] (03CR) 10Jforrester: "Hey @gtisza@wikimedia.org, as part of AW use of MainStash Amir and I thought it'd be good to track overall what's using MainStash; this ad" [puppet] - 10https://gerrit.wikimedia.org/r/1334836 (https://phabricator.wikimedia.org/T430940) (owner: 10Jforrester) [17:40:38] (03CR) 10Dzahn: [C:03+2] zookeeper: add spec tests [puppet] - 10https://gerrit.wikimedia.org/r/1329228 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [17:40:58] (03CR) 10Dzahn: [C:03+1] zookeeper: add spec tests [puppet] - 10https://gerrit.wikimedia.org/r/1329228 (https://phabricator.wikimedia.org/T435503) (owner: 10Hashar) [17:41:26] (03PS2) 10Bking: w[cd]qs: Auto-clean blazegraph tmpfiles [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) [17:42:49] (03CR) 10JMeybohm: [C:04-1] "I forgot: This requires a chart version bump" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333210 (https://phabricator.wikimedia.org/T436655) (owner: 10Giuseppe Lavagetto) [17:43:46] !log kamila@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet [17:44:35] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [17:44:37] (03PS2) 10Sbisson: ArticleGuidance: Add the redirect configuration keys [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334945 (https://phabricator.wikimedia.org/T434487) [17:44:46] (03PS2) 10Sbisson: ArticleGuidance: Remove the experiment configuration keys [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334946 (https://phabricator.wikimedia.org/T434487) [17:44:57] (03PS1) 10JMeybohm: vap: Makefile improvements [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337982 [17:45:24] (03PS1) 10Sbisson: Replace experiment with instrument and config-driven redirect [extensions/ArticleGuidance] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337983 (https://phabricator.wikimedia.org/T434487) [17:46:15] (03Merged) 10jenkins-bot: vap: Fix variable naming for validationActions [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337953 (https://phabricator.wikimedia.org/T436380) (owner: 10JMeybohm) [17:47:50] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 09 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#de" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1334945 (https://phabricator.wikimedia.org/T434487) (owner: 10Sbisson) [17:48:14] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Wednesday, September 09 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#de" [extensions/ArticleGuidance] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337983 (https://phabricator.wikimedia.org/T434487) (owner: 10Sbisson) [17:48:37] (03PS2) 10Slyngshede: debug.json: order codfw (primary) DC backends first [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319818 (https://phabricator.wikimedia.org/T433363) [17:50:01] (03CR) 10Slyngshede: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1337979 (https://phabricator.wikimedia.org/T436691) (owner: 10Ssingh) [17:51:43] (03CR) 10Ssingh: [C:03+2] acme_chief: add sretest2013 to unified cert host [puppet] - 10https://gerrit.wikimedia.org/r/1337979 (https://phabricator.wikimedia.org/T436691) (owner: 10Ssingh) [17:55:23] FIRING: GnmiInterfaceCountersDrop: ... [17:55:23] asw1-bw27-esams is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=asw1-bw27-esams:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [17:58:23] (03CR) 10Krinkle: "See also https://gerrit.wikimedia.org/r/c/operations/software/thumbor-plugins/+/1315693. It looks like that may've been when "make test" b" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1337608 (owner: 10Ladsgroup) [17:59:45] (03PS3) 10Bking: w[cd]qs: Auto-clean blazegraph tmpfiles [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) [18:00:02] 10ops-codfw, 06SRE, 10SRE-swift-storage, 06DC-Ops: Disk (sdf) failed in ms-be2089 - https://phabricator.wikimedia.org/T437293#12299285 (10Jhancock.wm) 05Open→03Resolved a:03Jhancock.wm @MatthewVernon disk has been replaced. please reopen if anything is awry! [18:00:05] dduvall and dancy: How many deployers does it take to do MediaWiki train - Utc-7 Version deploy? (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T1800). [18:00:18] (03CR) 10Krinkle: "Does this do anything other than swap around the order of two non-default options in a dropdown menu? If yes, why?" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319818 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [18:01:09] 10ops-codfw, 06SRE, 06DBA, 06DC-Ops: db2196 crashed - https://phabricator.wikimedia.org/T437015#12299301 (10Jhancock.wm) @Marostegui all the firmware is upgraded and the reports have been uploaded. Sorry for the wait! [18:02:07] (03PS2) 10Aaron Schulz: Add wmf-analytics-commons external module to commonswiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326425 (https://phabricator.wikimedia.org/T434927) [18:03:33] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 08 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326425 (https://phabricator.wikimedia.org/T434927) (owner: 10Aaron Schulz) [18:04:14] dancy: o/ [18:04:20] o/ [18:04:37] [18:05:36] (03CR) 10Btullis: "It seems that there is a systemd::tmpfile defined type that you could use instead of a file and a separate source." [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [18:05:39] dduvall: I have some outstanding SpiderPig changes to be able to add notes to error log messages. Unfortunately https://phabricator.wikimedia.org/T437069 is blocking delivery. [18:06:06] oh that sounds nice [18:06:33] yeah that issue is blocking a lot of image builds as well [18:08:19] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.19 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337985 (https://phabricator.wikimedia.org/T430838) [18:08:23] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by dduvall@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337985 (https://phabricator.wikimedia.org/T430838) (owner: 10TrainBranchBot) [18:09:28] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.19 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337985 (https://phabricator.wikimedia.org/T430838) (owner: 10TrainBranchBot) [18:10:32] (03CR) 10Krinkle: "Discussion at T289745." [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1319818 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [18:10:32] (03CR) 10Herron: [C:03+1] "nice, thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1333763 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [18:18:44] !log dduvall@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs T430838 [18:18:47] T430838: 1.47.0-wmf.19 deployment blockers - https://phabricator.wikimedia.org/T430838 [18:19:01] !log swfrench@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet [18:19:05] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet [18:19:39] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet [18:20:01] !log swfrench@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie [18:20:16] !log swfrench@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275 [18:20:45] !log swfrench@cumin1003 START - Cookbook sre.dns.netbox [18:25:20] !log swfrench@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003" [18:28:24] swfrench@cumin1003 renumber-node (PID 3795502) is awaiting input [18:28:35] ^ investigating a surprise diff [18:29:08] PROBLEM - HAProxy HTTPS wikiworkshop.org ECDSA on sretest2013 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [18:29:08] PROBLEM - haproxy process on sretest2013 is CRITICAL: PROCS CRITICAL: 0 processes with command name haproxy https://wikitech.wikimedia.org/wiki/HAProxy [18:29:36] PROBLEM - HAProxy HTTPS wikipedia25.org ECDSA on sretest2013 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [18:29:36] PROBLEM - HAProxy HTTPS wikipedia.org ECDSA on sretest2013 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [18:31:13] yes yes [18:31:18] (03PS1) 10JHathaway: rspamd: use new repo fork [puppet] - 10https://gerrit.wikimedia.org/r/1337990 (https://phabricator.wikimedia.org/T435225) [18:31:40] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337990 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [18:33:09] !log swfrench@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003" [18:33:09] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [18:33:10] !log swfrench@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [18:33:13] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [18:33:13] !log swfrench@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275 [18:34:06] !log swfrench@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275 [18:34:06] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275 [18:35:30] (03PS1) 10Ssingh: Add sretest2013 to unified cert host (expand) and late_command [puppet] - 10https://gerrit.wikimedia.org/r/1337993 (https://phabricator.wikimedia.org/T436691) [18:37:55] (03CR) 10BBlack: [C:03+1] Add sretest2013 to unified cert host (expand) and late_command [puppet] - 10https://gerrit.wikimedia.org/r/1337993 (https://phabricator.wikimedia.org/T436691) (owner: 10Ssingh) [18:38:41] (03CR) 10Ssingh: [C:03+2] Add sretest2013 to unified cert host (expand) and late_command [puppet] - 10https://gerrit.wikimedia.org/r/1337993 (https://phabricator.wikimedia.org/T436691) (owner: 10Ssingh) [18:39:56] (03CR) 10Ssingh: admin: move dkertesz from ldap_only_users to users; add to ops group [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) (owner: 10Daniel Kertesz) [18:41:10] (03PS2) 10Reedy: InitialiseSettings: Enable 2FA enforcement on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324753 (https://phabricator.wikimedia.org/T428103) [18:43:03] !log sukhe@cumin1003 START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie [18:43:20] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12299524 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie [18:43:21] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12299525 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jhancock@cumin2003 for host sretest2013.codfw.wmnet with OS trixie executed wit... [18:43:43] (03CR) 10Ssingh: "Sorry for the confusion. Let's rework this patch to ops-limited for now." [puppet] - 10https://gerrit.wikimedia.org/r/1337863 (https://phabricator.wikimedia.org/T437271) (owner: 10Daniel Kertesz) [18:45:11] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12299529 (10ssingh) On further discussion, we will go with ops-limited for now and then do ops in a later patch. Sorry for the confusion: it stemmed from this role being... [18:45:32] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to production for dkertesz - https://phabricator.wikimedia.org/T437271#12299532 (10ssingh) [18:48:34] (03PS1) 10JHathaway: WIP - stdlib [puppet] - 10https://gerrit.wikimedia.org/r/1337995 [18:49:03] (03PS2) 10JHathaway: WIP - stdlib [puppet] - 10https://gerrit.wikimedia.org/r/1337995 [18:49:12] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337995 (owner: 10JHathaway) [18:50:26] (03PS1) 10SBassett: Filter out non-http(s) license urls [extensions/CommonsMetadata] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337996 (https://phabricator.wikimedia.org/T435999) [18:50:44] (03PS1) 10SBassett: Filter out non-http(s) license urls [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337997 (https://phabricator.wikimedia.org/T435999) [18:50:55] (03PS1) 10SBassett: Filter out non-http(s) license urls [extensions/MediaSearch] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337998 (https://phabricator.wikimedia.org/T435999) [18:50:58] (03PS1) 10DErenrich: Enable discord preview extension code on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337999 (https://phabricator.wikimedia.org/T437344) [18:51:28] (03CR) 10Ssingh: "This change if applied for a given site, will indeed cease ns2 announcements for that particular site since we will stop advertising the V" [cookbooks] - 10https://gerrit.wikimedia.org/r/1335049 (owner: 10Ssingh) [18:52:11] (03CR) 10BCornwall: [C:03+1] interface::ipip: Guard absenting an address behind ifupdown fact [puppet] - 10https://gerrit.wikimedia.org/r/1337971 (owner: 10Majavah) [18:52:46] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 08 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337997 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [18:53:00] !log swfrench@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage [18:53:18] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 08 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [extensions/MediaSearch] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337998 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [18:53:36] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 08 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337996 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [18:56:01] (03CR) 10Bking: "Yeah, I saw that...it doesn't actually expose the config fields explained in the man page:" [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [18:56:11] PROBLEM - Host wikikube-worker1275 is DOWN: PING CRITICAL - Packet loss = 100% [18:57:02] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12299590 (10ssingh) Hi @RobH. Any update from their end? [18:58:23] (03CR) 10BCornwall: [C:03+2] interface::ipip: Guard absenting an address behind ifupdown fact [puppet] - 10https://gerrit.wikimedia.org/r/1337971 (owner: 10Majavah) [18:58:51] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage [18:59:11] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12299596 (10RobH) This had remote hands update that it was done, double checking the work now. [18:59:17] (03PS4) 10Bking: w[cd]qs: Auto-clean blazegraph tmpfiles [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) [19:01:12] RECOVERY - Host wikikube-worker1275 is UP: PING OK - Packet loss = 0%, RTA = 0.35 ms [19:01:34] (03CR) 10Eric Gardner: Enable discord preview extension code on testwiki (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337999 (https://phabricator.wikimedia.org/T437344) (owner: 10DErenrich) [19:02:11] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [19:02:53] (03CR) 10BCornwall: "The change is a welcome improvement! However, I'm still seeing the same issue:" [puppet] - 10https://gerrit.wikimedia.org/r/1332813 (owner: 10BCornwall) [19:04:07] (03PS1) 10Majavah: interface::ipip: Add additional check for absenting on networkd hosts [puppet] - 10https://gerrit.wikimedia.org/r/1338004 [19:04:16] (03CR) 10JHathaway: [C:03+1] w[cd]qs: Auto-clean blazegraph tmpfiles [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [19:04:32] !log sukhe@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie [19:04:44] !log sukhe@cumin1003 START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie [19:04:45] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12299635 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie executed with e... [19:04:58] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12299636 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie [19:05:05] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12299637 (10RobH) [19:06:14] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12299642 (10RobH) a:05RobH→03ssingh @ssingh: The new memory is dectected, same speed as the old just I can compare the serial output to the photos, throwing the photos h... [19:07:32] (03CR) 10BCornwall: [C:03+2] interface::ipip: Add additional check for absenting on networkd hosts [puppet] - 10https://gerrit.wikimedia.org/r/1338004 (owner: 10Majavah) [19:09:19] 10ops-codfw, 06SRE, 06DC-Ops, 06Infrastructure-Foundations, 10netops: codfw: pod AB switches upgrade (2026) - https://phabricator.wikimedia.org/T426197#12299657 (10cmooney) [19:10:08] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 08 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1335706 (https://phabricator.wikimedia.org/T436426) (owner: 10Hamish) [19:10:17] sukhe@cumin1003 reimage (PID 3802056) is awaiting input [19:12:04] !log sukhe@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie [19:12:16] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12299674 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie executed with e... [19:12:17] !log sukhe@cumin1003 START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie [19:12:29] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12299687 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie [19:17:24] !log sukhe@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie [19:17:34] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12299743 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie executed with e... [19:19:58] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie [19:20:18] (03CR) 10Bking: [C:03+2] w[cd]qs: Auto-clean blazegraph tmpfiles [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [19:21:06] (03PS2) 10DErenrich: Enable discord preview extension code on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337999 (https://phabricator.wikimedia.org/T437344) [19:21:27] (03PS3) 10DErenrich: Enable discord preview extension code on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337999 (https://phabricator.wikimedia.org/T437344) [19:22:10] (03PS4) 10DErenrich: Enable discord preview extension code on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337999 (https://phabricator.wikimedia.org/T437344) [19:22:16] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploy v1.5.4 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337862 (https://phabricator.wikimedia.org/T421813) [19:22:43] (03PS2) 10Santiago Faci: Test Kitchen UI: Deploy v1.5.4 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1337862 (https://phabricator.wikimedia.org/T421813) [19:23:39] !log sukhe@cumin1003 START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie [19:23:53] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12299785 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie [19:24:03] (03PS1) 10BCornwall: puppet-merge: Fix broken newline [puppet] - 10https://gerrit.wikimedia.org/r/1338013 [19:29:07] (03CR) 10Eric Gardner: Enable discord preview extension code on testwiki (031 comment) [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337999 (https://phabricator.wikimedia.org/T437344) (owner: 10DErenrich) [19:30:29] (03CR) 10Ssingh: puppet-merge: Fix broken newline (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338013 (owner: 10BCornwall) [19:31:40] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet [19:31:41] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet [19:31:43] !log swfrench@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet [19:31:47] (03CR) 10Ssingh: [C:03+1] puppet-merge: Fix broken newline (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338013 (owner: 10BCornwall) [19:32:08] (03CR) 10BCornwall: puppet-merge: Fix broken newline (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338013 (owner: 10BCornwall) [19:33:25] (03CR) 10Ssingh: [C:03+1] puppet-merge: Fix broken newline (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338013 (owner: 10BCornwall) [19:37:12] (03CR) 10BCornwall: [C:03+2] puppet-merge: Fix broken newline [puppet] - 10https://gerrit.wikimedia.org/r/1338013 (owner: 10BCornwall) [19:37:23] !log swfrench@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet [19:37:26] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet [19:37:57] !log eevans@cumin1003 START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART [19:37:58] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet [19:38:33] !log swfrench@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie [19:39:00] !log swfrench@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305 [19:39:16] !log swfrench@cumin1003 START - Cookbook sre.dns.netbox [19:40:25] PROBLEM - Host aqs1023 is DOWN: PING CRITICAL - Packet loss = 100% [19:41:07] FIRING: [2x] ProbeDown: Service aqs1023-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:43:14] !log swfrench@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003" [19:43:18] !log swfrench@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003" [19:43:18] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [19:43:19] !log swfrench@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [19:43:21] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [19:43:22] !log swfrench@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305 [19:44:08] !log swfrench@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305 [19:44:08] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305 [19:45:05] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART [19:45:11] 10ops-eqiad, 06SRE, 06Collaboration-Services, 06DC-Ops, 13Patch-For-Review: Repurpose ganeti102[3456] for Zuul migration - https://phabricator.wikimedia.org/T427353#12299906 (10Dzahn) I have been out-of-office and back now. The regex in the preseed file is: `'zuul[1-2]00[4-7]'` so I don't see a differen... [19:45:27] RECOVERY - Host aqs1023 is UP: PING OK - Packet loss = 0%, RTA = 0.32 ms [19:46:07] FIRING: [4x] ProbeDown: Service aqs1023-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:47:57] 06SRE, 06serviceops-deprecated, 07Wikimedia-Performance-recommendation: Evaluate using igbinary for MW php-apcu at WMF - https://phabricator.wikimedia.org/T225074#12299913 (10Krinkle) a:03MGoncalves-WMF [19:51:07] FIRING: [4x] ProbeDown: Service aqs1023-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:51:08] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm [19:55:39] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [19:56:07] FIRING: [4x] ProbeDown: Service aqs1023-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: That opportune time for a UTC late backport window deploy is upon us again. Don't be afraid. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T2000). [20:00:05] AaronSchulz, sbassett, aranyap, and hamishcz: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:10] :) [20:00:14] im here again [20:01:13] !log eevans@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage [20:02:23] o/ [20:02:31] o/ [20:03:09] ok, I can go first [20:03:10] (03PS3) 10Brouberol: global_config: add the LVS VIPs to the public druid external service endpoints [puppet] - 10https://gerrit.wikimedia.org/r/1338021 (https://phabricator.wikimedia.org/T437310) [20:03:25] (03PS1) 10SBassett: Filter out non-http(s) license urls [extensions/MultimediaViewer] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338023 (https://phabricator.wikimedia.org/T435999) [20:03:35] (03PS1) 10SBassett: Filter out non-http(s) license urls [extensions/MediaSearch] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338024 (https://phabricator.wikimedia.org/T435999) [20:03:38] (03CR) 10TrainBranchBot: [C:03+2] "Approved by aaron@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326425 (https://phabricator.wikimedia.org/T434927) (owner: 10Aaron Schulz) [20:03:41] (03PS2) 10SBassett: Filter out non-http(s) license urls [extensions/MediaSearch] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338024 (https://phabricator.wikimedia.org/T435999) [20:04:02] (03PS1) 10SBassett: Filter out non-http(s) license urls [extensions/CommonsMetadata] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338025 (https://phabricator.wikimedia.org/T435999) [20:04:16] wow.. [20:04:52] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [20:04:53] (03Merged) 10jenkins-bot: Add wmf-analytics-commons external module to commonswiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1326425 (https://phabricator.wikimedia.org/T434927) (owner: 10Aaron Schulz) [20:04:54] AaronSchulz: Sounds good. I just picked the 3 patches I’m deploying to wmf.18 as well, which I think I can still run all together via spiderpig. [20:05:09] looks like I have to wait for a while [20:05:17] !log aaron@deploy1003 Started scap sync-world: Backport for [[gerrit:1326425|Add wmf-analytics-commons external module to commonswiki (T434927)]] [20:05:21] T434927: Add "Wikimedia Commons Impact Metrics API" to Commons project - https://phabricator.wikimedia.org/T434927 [20:05:27] !log swfrench@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage [20:05:37] hamishcz: should just be two syncs from Aaron and I [20:05:54] can can, not a problem [20:08:14] mine is just config for the restsandbox, so it will be quick [20:09:23] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage [20:09:56] !log aaron@deploy1003 aaron: Backport for [[gerrit:1326425|Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:10:59] !log aaron@deploy1003 aaron: Continuing with deployment [20:13:14] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage [20:15:33] !log aaron@deploy1003 Finished scap sync-world: Backport for [[gerrit:1326425|Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s) [20:15:37] T434927: Add "Wikimedia Commons Impact Metrics API" to Commons project - https://phabricator.wikimedia.org/T434927 [20:15:38] 10ops-ulsfo, 06SRE, 06DC-Ops: ulsfo: OOB migration from copper to fiber - https://phabricator.wikimedia.org/T436499#12300102 (10RobH) Turns out the reply may not have included all the right parties so I did another reply today and included the additional parties on the DR side. > Digital Realty, > > I've... [20:16:37] AaronSchulz: all done? [20:16:48] yep [20:17:24] sukhe@cumin1003 reimage (PID 3806860) is awaiting input [20:17:35] (03PS1) 10Ahmon Dancy: scap.cfg.erb: mediawiki_runtime_image: Use php8.5 in beta [puppet] - 10https://gerrit.wikimedia.org/r/1338027 (https://phabricator.wikimedia.org/T432989) [20:17:35] Ok, I’ll run mine now... [20:18:07] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338023 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:18:08] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337997 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:18:08] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/MediaSearch] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338024 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:18:09] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/MediaSearch] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337998 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:18:10] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338025 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:18:15] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337996 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:19:18] (03Merged) 10jenkins-bot: Filter out non-http(s) license urls [extensions/MultimediaViewer] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338023 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:19:36] (03Merged) 10jenkins-bot: Filter out non-http(s) license urls [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337997 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:19:42] (03Merged) 10jenkins-bot: Filter out non-http(s) license urls [extensions/MediaSearch] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338024 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:19:48] (03Merged) 10jenkins-bot: Filter out non-http(s) license urls [extensions/MediaSearch] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337998 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:19:53] (03Merged) 10jenkins-bot: Filter out non-http(s) license urls [extensions/CommonsMetadata] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338025 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:22:36] (03Merged) 10jenkins-bot: Filter out non-http(s) license urls [extensions/CommonsMetadata] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1337996 (https://phabricator.wikimedia.org/T435999) (owner: 10SBassett) [20:23:06] !log sbassett@deploy1003 Started scap sync-world: Backport for [[gerrit:1338023|Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997|Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024|Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998|Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025|Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996|Filter out non- [20:23:06] http(s) license urls (T435999)]] [20:27:00] (03CR) 10Southparkfan: [C:03+1] scap.cfg.erb: mediawiki_runtime_image: Use php8.5 in beta [puppet] - 10https://gerrit.wikimedia.org/r/1338027 (https://phabricator.wikimedia.org/T432989) (owner: 10Ahmon Dancy) [20:27:20] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm [20:27:31] !log sbassett@deploy1003 sbassett: Backport for [[gerrit:1338023|Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997|Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024|Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998|Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025|Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996|Filter out non-http(s) license [20:27:31] urls (T435999)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:28:39] !log sbassett@deploy1003 sbassett: Continuing with deployment [20:30:39] (03CR) 10Ahmon Dancy: "Tested good in beta." [puppet] - 10https://gerrit.wikimedia.org/r/1338027 (https://phabricator.wikimedia.org/T432989) (owner: 10Ahmon Dancy) [20:32:47] Hey all - I need to stop the above, there’s a strtolower error happening... [20:33:11] (03PS1) 10Urbanecm: [Growth] Enable iterative Add Link task pool population on all wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1338031 (https://phabricator.wikimedia.org/T392944) [20:33:31] (03PS1) 10SBassett: Revert "Filter out non-http(s) license urls" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338032 [20:33:44] (03PS1) 10SBassett: Revert "Filter out non-http(s) license urls" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1338033 [20:34:21] (03PS1) 10SBassett: Revert "Filter out non-http(s) license urls" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338034 [20:34:22] !log sukhe@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie [20:34:31] (03PS1) 10SBassett: Revert "Filter out non-http(s) license urls" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1338035 [20:34:32] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12300205 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie executed with errors: - sretest2013 (... [20:34:50] (03PS1) 10SBassett: Revert "Filter out non-http(s) license urls" [extensions/MediaSearch] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338036 [20:35:07] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie [20:35:51] (03PS1) 10SBassett: Revert "Filter out non-http(s) license urls" [extensions/MediaSearch] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1338037 [20:36:46] 10SRE-swift-storage, 10Cloud-VPS (Debian Bullseye Deprecation): Migrate swift away from Debian Bullseye to Bookworm/Trixie - https://phabricator.wikimedia.org/T435783#12300211 (10komla) @MatthewVernon any luck with progress on this? [20:37:36] Running deploy of 6 revert patches now… [20:37:42] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/MediaSearch] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1338037 (owner: 10SBassett) [20:37:43] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338034 (owner: 10SBassett) [20:37:43] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1338035 (owner: 10SBassett) [20:37:45] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/MediaSearch] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338036 (owner: 10SBassett) [20:37:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338032 (owner: 10SBassett) [20:37:53] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbassett@deploy1003 using scap backport" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1338033 (owner: 10SBassett) [20:38:03] (03CR) 10Ahmon Dancy: [V:03+1 C:03+1] "This is ready to merge" [puppet] - 10https://gerrit.wikimedia.org/r/1338027 (https://phabricator.wikimedia.org/T432989) (owner: 10Ahmon Dancy) [20:43:41] (03CR) 10RLazarus: [C:03+1] mediawiki: Redirect /api/ to /w/rest.php [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330505 (https://phabricator.wikimedia.org/T433547) (owner: 10Clément Goubert) [20:44:03] (03CR) 10RLazarus: [C:03+1] "Resolving." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330505 (https://phabricator.wikimedia.org/T433547) (owner: 10Clément Goubert) [20:44:22] (03CR) 10Dzahn: [C:03+2] scap.cfg.erb: mediawiki_runtime_image: Use php8.5 in beta [puppet] - 10https://gerrit.wikimedia.org/r/1338027 (https://phabricator.wikimedia.org/T432989) (owner: 10Ahmon Dancy) [20:48:34] (03CR) 10Btullis: [C:03+1] w[cd]qs: Auto-clean blazegraph tmpfiles [puppet] - 10https://gerrit.wikimedia.org/r/1337978 (https://phabricator.wikimedia.org/T437298) (owner: 10Bking) [20:48:54] (03PS1) 10JHathaway: Puppet 8: Replace unscoped legacy facts in module opensearch_dashboards [puppet] - 10https://gerrit.wikimedia.org/r/1338040 (https://phabricator.wikimedia.org/T435225) [20:49:13] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1338040 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [20:50:43] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet [20:50:44] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet [20:50:46] !log swfrench@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet [20:51:54] (03Merged) 10jenkins-bot: Revert "Filter out non-http(s) license urls" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338034 (owner: 10SBassett) [20:51:55] (03Merged) 10jenkins-bot: Revert "Filter out non-http(s) license urls" [extensions/CommonsMetadata] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1338035 (owner: 10SBassett) [20:52:37] (03PS1) 10Bking: matomo: Allow dse-k8s pods to access database [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) [20:53:05] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) (owner: 10Bking) [20:53:23] (03CR) 10CI reject: [V:04-1] matomo: Allow dse-k8s pods to access database [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) (owner: 10Bking) [20:53:28] !log sukhe@cumin1003 START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie [20:53:40] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12300260 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie [20:54:02] sbassett: I think you got 6 patches but 2 merged? [20:54:44] (03PS2) 10Bking: matomo: Allow dse-k8s pods to access database [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) [20:55:27] hamishcz: 4 now, just very slow, sorry [20:55:44] the master patches all merged fine, so this was a little unexpected [20:55:51] (03CR) 10Bking: [C:03+1] global_config: add the LVS VIPs to the public druid external service endpoints [puppet] - 10https://gerrit.wikimedia.org/r/1338021 (https://phabricator.wikimedia.org/T437310) (owner: 10Brouberol) [20:55:54] can, no worries [20:56:16] (03PS1) 10JHathaway: Puppet 8: Replace unscoped legacy facts in profile lists [puppet] - 10https://gerrit.wikimedia.org/r/1338044 (https://phabricator.wikimedia.org/T435225) [20:56:18] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) (owner: 10Bking) [20:56:46] I've tried to wait all over the window in the morning but no one came to help me :| [20:57:23] How can I be able to deploy the patch by myself, btw? [20:57:25] (03Merged) 10jenkins-bot: Revert "Filter out non-http(s) license urls" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.18) - 10https://gerrit.wikimedia.org/r/1338032 (owner: 10SBassett) [20:57:44] or it does need advance permission? [20:58:18] (03Merged) 10jenkins-bot: Revert "Filter out non-http(s) license urls" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.19) - 10https://gerrit.wikimedia.org/r/1338033 (owner: 10SBassett) [20:58:48] !log sbassett@deploy1003 Started scap sync-world: Backport for [[gerrit:1338037|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033|Revert "Filter out n [20:58:49] on-http(s) license urls"]] [20:58:51] hamishcz: Often the backport windows are self-serve, with spiderpig. But if you need someone to help deploy, we can arrange that. [20:59:01] !log eevans@cumin1003 START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet [20:59:30] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1338044 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [21:00:04] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260908T2100) [21:00:31] sbassett: That requires the production access I think, from the manual on Wikitech.. [21:01:07] RESOLVED: [4x] ProbeDown: Service aqs1023-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:02:07] FIRING: [2x] ProbeDown: Service aqs1023-a:9042 has failed probes (tcp_cassandra_a_cql_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:03:23] !log sbassett@deploy1003 sbassett: Backport for [[gerrit:1338037|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033|Revert "Filter out non-http(s) lice [21:03:23] nse urls"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:03:41] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12300268 (10ssingh) ` in-target /usr/sbin/nvme id-ns /dev/nvme0n1 Sep 8 21:02:05 in-target: lbaf 0 : ms:0 lbads:9 rp:0x2 (in use) Sep 8 21:02:05 in-target: lbaf 1... [21:04:27] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet [21:04:30] (03PS1) 10Ssingh: sretest2013: set LBA_FORMAT_NUMBER to 1 [puppet] - 10https://gerrit.wikimedia.org/r/1338045 (https://phabricator.wikimedia.org/T436691) [21:04:54] !log sukhe@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie [21:05:03] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12300272 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by sukhe@cumin1003 for host sretest2013.codfw.wmnet with OS trixie executed with e... [21:05:38] !log sbassett@deploy1003 sbassett: Continuing with deployment [21:05:49] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm [21:05:52] (03PS2) 10JHathaway: Puppet 8: Replace unscoped legacy facts in module opensearch_dashboards [puppet] - 10https://gerrit.wikimedia.org/r/1338040 (https://phabricator.wikimedia.org/T435225) [21:05:53] (03PS3) 10Andrea Denisse: klaxon: Switch to GitLab origin [puppet] - 10https://gerrit.wikimedia.org/r/1335233 (https://phabricator.wikimedia.org/T436694) [21:06:08] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1338040 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [21:06:22] RESOLVED: [4x] ProbeDown: Service aqs1023-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:06:38] (03CR) 10Andrea Denisse: klaxon: Switch to GitLab origin (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1335233 (https://phabricator.wikimedia.org/T436694) (owner: 10Andrea Denisse) [21:06:46] (03CR) 10BCornwall: [C:03+1] sretest2013: set LBA_FORMAT_NUMBER to 1 [puppet] - 10https://gerrit.wikimedia.org/r/1338045 (https://phabricator.wikimedia.org/T436691) (owner: 10Ssingh) [21:06:59] (03CR) 10Andrea Denisse: [C:03+2] klaxon: Switch to GitLab origin [puppet] - 10https://gerrit.wikimedia.org/r/1335233 (https://phabricator.wikimedia.org/T436694) (owner: 10Andrea Denisse) [21:08:32] (03CR) 10Ryan Kemper: matomo: Allow dse-k8s pods to access database (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) (owner: 10Bking) [21:08:45] (03CR) 10Ssingh: [C:03+2] sretest2013: set LBA_FORMAT_NUMBER to 1 [puppet] - 10https://gerrit.wikimedia.org/r/1338045 (https://phabricator.wikimedia.org/T436691) (owner: 10Ssingh) [21:09:04] (03CR) 10JHathaway: [C:03+1] Setup a syncrepl cluster on trixie/MDB [puppet] - 10https://gerrit.wikimedia.org/r/1335825 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [21:10:12] !log sbassett@deploy1003 Finished scap sync-world: Backport for [[gerrit:1338037|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032|Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033|Revert "Filter out [21:10:12] non-http(s) license urls"]] (duration: 11m 23s) [21:11:01] (03PS3) 10Bking: matomo: Allow dse-k8s pods to access database [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) [21:11:22] FIRING: [8x] ProbeDown: Service aqs1023-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:11:30] (03PS4) 10Bking: matomo: Allow dse-k8s pods to access database [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) [21:13:39] hamishcz: my issue should be resolved now. i can probably deploy your config change, but it looks like there were issues with previous deployment attempts? [21:13:43] (03CR) 10Reedy: [C:03+2] InitialiseSettings: Enable 2FA enforcement on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324753 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [21:14:20] Ah yes [21:14:32] but we can try again I think [21:14:40] (03Merged) 10jenkins-bot: InitialiseSettings: Enable 2FA enforcement on remaining private wikis [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1324753 (https://phabricator.wikimedia.org/T428103) (owner: 10Reedy) [21:14:44] As what I said on Phab, it's LGTM [21:14:54] Ok, I can try it, give me a second [21:15:10] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:15:13] !log reedy@deploy1003 Started scap sync-world: Backport for [[gerrit:1324753|InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] [21:15:17] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [21:15:44] !log eevans@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage [21:16:45] (03CR) 10Bking: matomo: Allow dse-k8s pods to access database (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) (owner: 10Bking) [21:17:05] hamishcz: ok, we’ll have to wait on reedy’s deployment here for a second [21:17:16] (03CR) 10Ryan Kemper: [C:03+1] matomo: Allow dse-k8s pods to access database [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) (owner: 10Bking) [21:17:32] Ah okay.. [21:18:40] !log swfrench@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet [21:18:44] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet [21:19:08] (03CR) 10Bking: [C:03+2] matomo: Allow dse-k8s pods to access database [puppet] - 10https://gerrit.wikimedia.org/r/1338042 (https://phabricator.wikimedia.org/T436003) (owner: 10Bking) [21:19:20] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet [21:19:24] !log reedy@deploy1003 reedy: Backport for [[gerrit:1324753|InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:19:32] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage [21:19:44] !log swfrench@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie [21:19:53] !log reedy@deploy1003 reedy: Continuing with deployment [21:20:11] !log swfrench@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306 [21:20:21] !log swfrench@cumin1003 START - Cookbook sre.dns.netbox [21:24:25] !log reedy@deploy1003 Finished scap sync-world: Backport for [[gerrit:1324753|InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s) [21:24:29] T428103: Enforce 2FA for all users on private wikis in WMF production - https://phabricator.wikimedia.org/T428103 [21:25:19] !log swfrench@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003" [21:28:23] swfrench@cumin1003 renumber-node (PID 3821322) is awaiting input [21:30:21] !log swfrench@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003" [21:30:21] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [21:30:22] !log swfrench@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:30:25] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [21:30:25] !log swfrench@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306 [21:31:03] !log swfrench@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306 [21:31:03] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306 [21:33:50] !log brett@cumin2003 START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie [21:33:56] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12300442 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by brett@cumin2003 for host cp6008.drmrs.wmnet with OS trixie [21:34:02] hey folks, there's a readers deployment window happening now, I'd like to do a portal deploy, just checking that nothing else is going on now [21:35:58] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm [21:39:33] !log jdrewniak@deploy1003 Started scap sync-world: Backport for [[gerrit:1338049|Assets build - 2026-09-08 21:25:19+00:00]] [21:40:51] !log jdrewniak@deploy1003 portalsbuilder, jdrewniak: Backport for [[gerrit:1338049|Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:43:15] !log jdrewniak@deploy1003 portalsbuilder, jdrewniak: Continuing with deployment [21:44:17] (03PS2) 10SBassett: thwikibooks: update tagline and wordmark [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1335706 (https://phabricator.wikimedia.org/T436426) (owner: 10Hamish) [21:45:00] !log jdrewniak@deploy1003 Finished scap sync-world: Backport for [[gerrit:1338049|Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s) [21:47:07] FIRING: [4x] ProbeDown: Service aqs1024-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:47:54] (03CR) 10Eric Gardner: [C:03+1] Enable discord preview extension code on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337999 (https://phabricator.wikimedia.org/T437344) (owner: 10DErenrich) [21:48:35] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [21:49:36] (03PS1) 10Jdrewniak: Bumping portals to master [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1338053 (https://phabricator.wikimedia.org/T128546) [21:50:39] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdrewniak@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1338053 (https://phabricator.wikimedia.org/T128546) (owner: 10Jdrewniak) [21:51:22] RESOLVED: [4x] ProbeDown: Service aqs1024-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:51:29] (03Merged) 10jenkins-bot: Bumping portals to master [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1338053 (https://phabricator.wikimedia.org/T128546) (owner: 10Jdrewniak) [21:51:46] !log swfrench@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage [21:51:49] !log jdrewniak@deploy1003 Started scap sync-world: Backport for [[gerrit:1338053|Bumping portals to master (T128546)]] [21:51:52] T128546: [Recurring Task] Update Wikipedia and sister projects portals statistics - https://phabricator.wikimedia.org/T128546 [21:52:07] FIRING: [5x] ProbeDown: Service aqs1024-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:52:46] !log brett@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage [21:55:38] FIRING: GnmiInterfaceCountersDrop: ... [21:55:38] asw1-bw27-esams is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=asw1-bw27-esams:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [21:56:18] !log jdrewniak@deploy1003 jdrewniak: Backport for [[gerrit:1338053|Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:56:22] FIRING: [8x] ProbeDown: Service aqs1024-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:57:03] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage [21:57:06] !log jdrewniak@deploy1003 jdrewniak: Continuing with deployment [21:57:32] (03PS3) 10Hamish: thwikibooks: update tagline and wordmark [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1335706 (https://phabricator.wikimedia.org/T436426) [21:58:39] !log drop links tables from db2210 (T437278) [21:58:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:58:43] T437278: Drop unneeded tables from x4 and s4 - https://phabricator.wikimedia.org/T437278 [21:59:49] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [22:00:04] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [22:00:56] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [22:00:58] PROBLEM - HAProxy HTTPS measure-eqiad.wikimedia.org ECDSA on cp6008 is CRITICAL: SSL CRITICAL - failed to connect or SSL handshake:Connection refused https://wikitech.wikimedia.org/wiki/HTTPS [22:01:32] (03CR) 10JHathaway: "All the depends on commits have been merged!" [puppet] - 10https://gerrit.wikimedia.org/r/1329284 (https://phabricator.wikimedia.org/T435917) (owner: 10Hashar) [22:01:43] !log jdrewniak@deploy1003 Finished scap sync-world: Backport for [[gerrit:1338053|Bumping portals to master (T128546)]] (duration: 09m 53s) [22:01:46] T128546: [Recurring Task] Update Wikipedia and sister projects portals statistics - https://phabricator.wikimedia.org/T128546 [22:03:17] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage [22:18:36] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie [22:21:22] FIRING: [4x] ProbeDown: Service aqs1025-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [22:22:07] RESOLVED: [4x] ProbeDown: Service aqs1025-a:7000 has failed probes (tcp_cassandra_a_ssl_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [22:26:30] RECOVERY - HAProxy HTTPS measure-eqiad.wikimedia.org ECDSA on cp6008 is OK: SSL OK - Certificate measure-eqiad.wikimedia.org contains all required SANs:Certificate measure-eqiad.wikimedia.org (ECDSA) valid until 2026-10-03 14:54:24 +0000 (expires in 24 days) https://wikitech.wikimedia.org/wiki/HTTPS [22:27:17] !log brett@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie [22:27:23] 10ops-drmrs, 06DC-Ops, 06Traffic: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12300565 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by brett@cumin2003 for host cp6008.drmrs.wmnet with OS trixie completed: - cp6008 (**PASS**) -... [22:30:24] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet [22:30:25] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet [22:30:27] !log swfrench@cumin1003 END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet [22:31:44] !log swfrench@cumin1003 START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet [22:31:48] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet [22:32:23] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet [22:32:40] !log swfrench@cumin1003 START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie [22:33:08] !log swfrench@cumin1003 START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313 [22:33:21] !log swfrench@cumin1003 START - Cookbook sre.dns.netbox [22:37:13] !log swfrench@cumin1003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003" [22:37:17] !log swfrench@cumin1003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003" [22:37:17] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [22:37:17] !log swfrench@cumin1003 START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:37:20] !log swfrench@cumin1003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [22:37:20] !log swfrench@cumin1003 START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313 [22:38:00] jouncebot: nowandnext [22:38:00] No deployments scheduled for the next 3 hour(s) and 21 minute(s) [22:38:00] In 3 hour(s) and 21 minute(s): Automatic deployment of MediaWiki to pretrain wikis - see mw:Pretrain (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260909T0200) [22:38:23] cool, going to sneak in a beta-only [22:38:43] !log swfrench@cumin1003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313 [22:38:43] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313 [22:38:48] (03CR) 10Milazg: [C:03+1] Page: Clean up tests of PageProps service class [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327671 (owner: 10Krinkle) [22:44:10] 10ops-codfw, 06SRE, 06Collaboration-Services, 06DC-Ops: lists2001 has multiple bus errors - https://phabricator.wikimedia.org/T423159#12300614 (10Dzahn) @Jhancock.wm are the errors gone? [22:44:33] (03Abandoned) 10Krinkle: Page: Clean up tests of PageProps service class [core] (wmf/1.47.0-wmf.16) - 10https://gerrit.wikimedia.org/r/1327671 (owner: 10Krinkle) [22:47:43] (03CR) 10TrainBranchBot: [C:03+2] "Approved by thcipriani@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1336153 (https://phabricator.wikimedia.org/T436465) (owner: 10Southparkfan) [22:49:11] (03Merged) 10jenkins-bot: LabsServices: adjust Mathoid URI to use service RR [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1336153 (https://phabricator.wikimedia.org/T436465) (owner: 10Southparkfan) [22:57:34] !log swfrench@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage [23:02:58] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [23:03:20] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [23:03:50] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage [23:04:29] (03PS4) 10Andrea Denisse: grafana: Add tamirsuliman-weathermap-panel to sandbox [puppet] - 10https://gerrit.wikimedia.org/r/1338056 (https://phabricator.wikimedia.org/T436075) [23:04:30] (03CR) 10Andrea Denisse: "This plugin was requested by cdanis on IRC." [puppet] - 10https://gerrit.wikimedia.org/r/1338056 (https://phabricator.wikimedia.org/T436075) (owner: 10Andrea Denisse) [23:05:57] !log dropped 57 tables on db1260 (T437278) [23:05:59] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [23:06:00] T437278: Drop unneeded tables from x4 and s4 - https://phabricator.wikimedia.org/T437278 [23:06:18] that's a lot of weight, you're going to break it :( [23:07:15] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [23:07:45] !log eevans@cumin1003 START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART [23:14:55] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART [23:15:42] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [23:19:10] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [23:19:24] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [23:22:25] (03PS1) 10Andrea Denisse: aptrepo: Add grafana-plugins to apt-staging [puppet] - 10https://gerrit.wikimedia.org/r/1338064 (https://phabricator.wikimedia.org/T436056) [23:22:25] (03CR) 10Andrea Denisse: [V:03+1] "Now we have a GitLab CI pipeline that builds Deb package [1] that is then imported into APT Staging [2]. This change would allow us to pro" [puppet] - 10https://gerrit.wikimedia.org/r/1338064 (https://phabricator.wikimedia.org/T436056) (owner: 10Andrea Denisse) [23:25:11] !log swfrench@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie [23:25:14] rzl: https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&refresh=5m&var-server=db1260&var-datasource=000000026&var-cluster=mysql&viewPanel=panel-28&from=now-2d&to=now&timezone=utc :P [23:25:21] (03PS2) 10Andrea Denisse: aptrepo: Add grafana-plugins to main apt repo [puppet] - 10https://gerrit.wikimedia.org/r/1338064 (https://phabricator.wikimedia.org/T436056) [23:25:48] Amir1: you can't fool me, that's not a table, it's a line chart [23:26:12] oh I just understand what you meant [23:26:17] * swfrench-wmf shakes head at absolutely wild disk usage graph [23:26:27] dc-ops problem [23:28:42] oh yes sorry, didn't mean to actually worry you [23:29:04] they can deal with the tables there [23:30:13] swfrench-wmf: https://grafana.wikimedia.org/d/000000273/mysql?from=now-2d&to=now&timezone=utc&var-job=$__all&var-server=db1261&var-port=9104&refresh=1m&viewPanel=panel-13 [23:30:27] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [23:30:37] cache efficiency went from two 9s to 3 9s :D [23:30:46] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [23:31:05] disk read is at one third [23:31:34] damn [23:33:02] credit to z.abe [23:34:45] that's amazing :) [23:35:20] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [23:37:35] !log swfrench@cumin1003 START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet [23:37:36] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet [23:37:37] !log swfrench@cumin1003 END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet [23:39:48] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [23:41:39] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1338067 [23:41:39] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1338067 (owner: 10TrainBranchBot) [23:48:59] !log eevans@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage [23:50:19] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1338067 (owner: 10TrainBranchBot) [23:51:52] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage [23:55:39] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired