[00:21:10] FIRING: BFDdown: BFD session down between cr2-eqsin and fe80::5e5e:ab07:93d:81a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [00:26:10] RESOLVED: BFDdown: BFD session down between cr2-eqsin and fe80::5e5e:ab07:93d:81a4 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [00:29:59] (03PS2) 10Divec: Change feedback URLs for EditCheck TextMatch on ruwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313951 (https://phabricator.wikimedia.org/T426271) (owner: 10Esanders) [00:34:07] (03CR) 10Divec: [C:03+1] Change feedback URLs for EditCheck TextMatch on ruwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313951 (https://phabricator.wikimedia.org/T426271) (owner: 10Esanders) [00:42:44] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12348116 (10ZauberViolino) Reported on zhwiki. https://zh.wikipedia.org/wiki/Project:%E4%BA%92%E5%8A%A9%E5%AE%A2%E6%A0%88/%E6%8A%80%E6%9C%AF#关于最近编辑频繁会话丢失的情况 [01:06:38] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [01:08:33] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 712368584 and 57 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [01:10:17] (03PS1) 10TrainBranchBot: Branch commit for wmf/1.47.0-wmf.21 [core] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343660 (https://phabricator.wikimedia.org/T438217) [01:10:20] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/1.47.0-wmf.21 [core] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343660 (https://phabricator.wikimedia.org/T438217) (owner: 10TrainBranchBot) [01:11:14] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1343661 [01:11:14] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1343661 (owner: 10TrainBranchBot) [01:11:33] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 104832 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [01:18:10] (03Merged) 10jenkins-bot: Branch commit for wmf/1.47.0-wmf.21 [core] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343660 (https://phabricator.wikimedia.org/T438217) (owner: 10TrainBranchBot) [01:20:22] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 22 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342825 (https://phabricator.wikimedia.org/T436652) (owner: 10Codename Noreste) [01:20:32] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1343661 (owner: 10TrainBranchBot) [01:35:10] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [01:35:31] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12348164 (10Soda) [01:53:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:53:41] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [01:59:18] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=logging-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [02:00:05] Deploy window Automatic branching of MediaWiki, extensions, skins, and vendor – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T0200) [02:00:05] Deploy window Automatic deployment of MediaWiki to pretrain wikis - see mw:Pretrain (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T0200) [02:00:58] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:02:13] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:06:26] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:06:37] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b4-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [02:08:29] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 30s) [02:12:13] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:17:13] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:32:39] PROBLEM - dump of s4 in codfw on backupmon1001 is CRITICAL: Last dump for s4 at codfw (db2239) taken on 2026-09-22 01:02:49 is 180 GiB, but the previous one was 242 GiB, a change of -25.5 % https://wikitech.wikimedia.org/wiki/MariaDB/Backups%23Rerun_a_failed_backup [02:38:30] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 22 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313951 (https://phabricator.wikimedia.org/T426271) (owner: 10Esanders) [02:44:10] FIRING: BFDdown: BFD session down between cr2-eqsin and fe80::6687:8807:ef2:7018 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:49:10] RESOLVED: BFDdown: BFD session down between cr2-eqsin and fe80::6687:8807:ef2:7018 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-eqsin:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [02:56:39] (03PS1) 10Andrew Bogott: cloudvps-project-usage.py: Add a few quota metrics [puppet] - 10https://gerrit.wikimedia.org/r/1343670 (https://phabricator.wikimedia.org/T436275) [03:00:05] Deploy window Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see Heterogeneous deployment/Train deploys (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T0300) [03:01:59] (03PS1) 10TrainBranchBot: testwikis to 1.47.0-wmf.21 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343671 (https://phabricator.wikimedia.org/T438217) [03:02:02] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by mwpresync@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343671 (https://phabricator.wikimedia.org/T438217) (owner: 10TrainBranchBot) [03:02:53] (03Merged) 10jenkins-bot: testwikis to 1.47.0-wmf.21 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343671 (https://phabricator.wikimedia.org/T438217) (owner: 10TrainBranchBot) [03:03:15] !log mwpresync@deploy1003 Started scap sync-world: testwikis to 1.47.0-wmf.21 refs T438217 [03:03:18] T438217: 1.47.0-wmf.21 deployment blockers - https://phabricator.wikimedia.org/T438217 [03:21:27] PROBLEM - Improperly owned -0:0- files in /srv/mediawiki-staging on deploy2003 is CRITICAL: Improperly owned (0:0) files in /srv/mediawiki-staging https://wikitech.wikimedia.org/wiki/Monitoring/bad_directory_owner [03:31:27] RECOVERY - Improperly owned -0:0- files in /srv/mediawiki-staging on deploy2003 is OK: Files ownership is ok. https://wikitech.wikimedia.org/wiki/Monitoring/bad_directory_owner [03:39:07] !log mwpresync@deploy1003 Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs T438217 (duration: 35m 52s) [03:39:10] T438217: 1.47.0-wmf.21 deployment blockers - https://phabricator.wikimedia.org/T438217 [04:00:05] Deploy window Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T0400) [04:02:34] !log mwpresync@deploy1003 Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s) [04:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [04:25:28] (03PS1) 10Hashar: jenkins: exit the JVM on OutOfMemoryError [puppet] - 10https://gerrit.wikimedia.org/r/1343675 (https://phabricator.wikimedia.org/T435791) [04:26:37] (03PS2) 10Hashar: jenkins: exit the JVM on OutOfMemoryError [puppet] - 10https://gerrit.wikimedia.org/r/1343675 (https://phabricator.wikimedia.org/T435791) [05:06:38] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [05:30:59] (03PS4) 10Giuseppe Lavagetto: external_clouds_vendors: stop generating the datafile [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) [05:35:10] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [05:43:52] (03PS1) 10Marostegui: installserver: Do not reimage db1284 [puppet] - 10https://gerrit.wikimedia.org/r/1343841 [05:48:02] (03CR) 10Marostegui: [C:03+2] installserver: Do not reimage db1284 [puppet] - 10https://gerrit.wikimedia.org/r/1343841 (owner: 10Marostegui) [05:49:25] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (NOOP 1 CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [05:51:33] (03CR) 10Giuseppe Lavagetto: [V:03+1 C:03+2] external_clouds_vendors: stop generating the datafile [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [05:53:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:53:41] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [05:59:18] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=logging-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [06:00:05] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T0600) [06:00:05] marostegui, cezmunsta, and federico3: #bothumor I � Unicode. All rise for Primary database switchover deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T0600). [06:06:38] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b4-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [06:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:17:28] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:19:13] PROBLEM - Host wikikube-worker1152 is DOWN: PING CRITICAL - Packet loss = 100% [06:23:30] (03CR) 10Arnaudb: [C:03+2] "reviewed, lgtm, deploying" [puppet] - 10https://gerrit.wikimedia.org/r/1343675 (https://phabricator.wikimedia.org/T435791) (owner: 10Hashar) [06:23:32] FIRING: KubernetesCalicoDown: wikikube-worker1152.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s&var-instance=wikikube-worker1152.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [06:31:09] (03PS1) 10Muehlenhoff: Remove the monthly_ganeti_rebalance timer [puppet] - 10https://gerrit.wikimedia.org/r/1343848 [06:32:07] (03CR) 10CI reject: [V:04-1] Remove the monthly_ganeti_rebalance timer [puppet] - 10https://gerrit.wikimedia.org/r/1343848 (owner: 10Muehlenhoff) [06:34:43] (03PS2) 10Muehlenhoff: Remove the monthly_ganeti_rebalance timer [puppet] - 10https://gerrit.wikimedia.org/r/1343848 [06:38:31] (03PS1) 10Muehlenhoff: Record LDAP access for eniroomand and lizf [puppet] - 10https://gerrit.wikimedia.org/r/1343852 [06:39:25] (03PS12) 10Arnaudb: mesh: add opt-in websocket support in configuration 1.17.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338751 (https://phabricator.wikimedia.org/T436657) [06:39:31] (03PS6) 10Arnaudb: modules: Prepare mesh.configuration minor version bump [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338959 (https://phabricator.wikimedia.org/T436657) [06:42:51] (03CR) 10Muehlenhoff: [C:03+2] Record LDAP access for eniroomand and lizf [puppet] - 10https://gerrit.wikimedia.org/r/1343852 (owner: 10Muehlenhoff) [06:48:34] (03PS1) 10Arthur taylor: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343853 (https://phabricator.wikimedia.org/T427589) [06:59:18] divec: I could scap the two config changes together, if that's okay with you? [06:59:34] Ok yes, thanks! [07:00:05] Amir1, urbanecm, and awight: #bothumor My software never has bugs. It just develops random features. Rise for UTC morning backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T0700). [07:00:05] awight and divec: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:24] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12348527 (10MoritzMuehlenhoff) [07:02:04] !log installing pyasn1 security updates [07:02:05] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:02:23] (03CR) 10Awight: [C:03+1] Change feedback URLs for EditCheck TextMatch on ruwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313951 (https://phabricator.wikimedia.org/T426271) (owner: 10Esanders) [07:02:41] (03CR) 10TrainBranchBot: [C:03+2] "Approved by awight@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343347 (https://phabricator.wikimedia.org/T438463) (owner: 10Seanleong-wmde) [07:02:42] (03CR) 10TrainBranchBot: [C:03+2] "Approved by awight@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313951 (https://phabricator.wikimedia.org/T426271) (owner: 10Esanders) [07:03:37] (03Merged) 10jenkins-bot: Config change for launch of stopping sending LL notifications. [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343347 (https://phabricator.wikimedia.org/T438463) (owner: 10Seanleong-wmde) [07:03:41] (03Merged) 10jenkins-bot: Change feedback URLs for EditCheck TextMatch on ruwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1313951 (https://phabricator.wikimedia.org/T426271) (owner: 10Esanders) [07:04:58] !log awight@deploy1003 Started scap sync-world: Backport for [[gerrit:1343347|Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951|Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] [07:05:03] T438463: Create config change for launch of stopping sending LL notifications for other language wikis - https://phabricator.wikimedia.org/T438463 [07:05:03] T426271: Change links in EditCheck Suggestion for ruwiki - https://phabricator.wikimedia.org/T426271 [07:09:27] !log awight@deploy1003 seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347|Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951|Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [07:09:58] (aww, I like saying that.) [07:12:44] Thanks! It works for me on ruwiki [07:13:19] ack [07:13:28] (03PS1) 10Muehlenhoff: use_linux612_on_bookworm: Bump kernel to 6.12.107 [puppet] - 10https://gerrit.wikimedia.org/r/1343856 [07:15:32] !log awight@deploy1003 seanleong-wmde, esanders, awight: Continuing with deployment [07:16:13] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09), 07Essential-Work: Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12348564 (10brouberol) Seems like https://phabricator.wikimedia.org/T438744 is another occurence. [07:22:44] !log awight@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343347|Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951|Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s) [07:22:49] T438463: Create config change for launch of stopping sending LL notifications for other language wikis - https://phabricator.wikimedia.org/T438463 [07:22:49] T426271: Change links in EditCheck Suggestion for ruwiki - https://phabricator.wikimedia.org/T426271 [07:23:14] !log UTC morning deployment window complete [07:23:15] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:23:45] (03CR) 10Trueg: "Where in https://quarkus.io/guides/telemetry-micrometer/#configuration-reference does it say that it is always enabled?" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343643 (https://phabricator.wikimedia.org/T438789) (owner: 10Lerickson) [07:24:09] (03CR) 10Volans: cookbooks/idm: Add user-cleanup cookbook for DPE SRE hosts (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1342235 (https://phabricator.wikimedia.org/T437615) (owner: 10Klausman) [07:24:12] (03CR) 10Mahmoud-abdelsattar: [C:03+1] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343853 (https://phabricator.wikimedia.org/T427589) (owner: 10Arthur taylor) [07:26:57] (03CR) 10Klausman: [C:03+1] dse-k8s: Delegate reverse DNS for the second eqiad Pod range [dns] - 10https://gerrit.wikimedia.org/r/1343650 (https://phabricator.wikimedia.org/T430658) (owner: 10Btullis) [07:30:03] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12348589 (10Joe) Hi, the increase in session lost in edit counts and other places seems to relate to the time when, due to hardware failures, we decided to depool the session storage... [07:33:10] (03CR) 10Ayounsi: [C:03+1] Remove the monthly_ganeti_rebalance timer [puppet] - 10https://gerrit.wikimedia.org/r/1343848 (owner: 10Muehlenhoff) [07:34:13] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12348599 (10Novem_Linguae) Relevant grafana graph (thanks Joe for providing): https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&from=now-7d&to=now&timezone=utc&viewPanel=... [07:34:51] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12348600 (10Novem_Linguae) [07:44:41] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for shellbox* in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1343539 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [07:45:20] (03PS1) 10Muehlenhoff: Only install python3-docker-report on active report host [puppet] - 10https://gerrit.wikimedia.org/r/1343906 (https://phabricator.wikimedia.org/T435314) [07:46:10] (03PS2) 10Gmodena: wdqs: add wikidata prefixes config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343616 (https://phabricator.wikimedia.org/T438476) [07:47:29] (03CR) 10Filippo Giunchedi: [C:03+1] "LGTM, thank you" [puppet] - 10https://gerrit.wikimedia.org/r/1343587 (https://phabricator.wikimedia.org/T431307) (owner: 10Andrew Bogott) [07:51:36] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343906 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [07:57:35] jouncebot: nowandnext [07:57:35] For the next 0 hour(s) and 2 minute(s): UTC morning backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T0700) [07:57:35] In 2 hour(s) and 2 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1000) [07:58:21] (03CR) 10Arthur taylor: [C:03+2] wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343853 (https://phabricator.wikimedia.org/T427589) (owner: 10Arthur taylor) [07:59:26] !log ayounsi@cumin1004 START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, T437984] [07:59:28] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, T437984] [07:59:29] T437984: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984 [08:00:19] (03CR) 10Volans: [C:03+1] "LGTM couple of questions inline" [puppet] - 10https://gerrit.wikimedia.org/r/1343547 (https://phabricator.wikimedia.org/T438731) (owner: 10Majavah) [08:00:54] (03Merged) 10jenkins-bot: wikidata-query-gui: bump to latest version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343853 (https://phabricator.wikimedia.org/T427589) (owner: 10Arthur taylor) [08:01:04] !log ayounsi@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot [08:01:23] !log arthurtaylor@deploy1003 helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply [08:01:37] !log ayounsi@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot [08:01:53] (03CR) 10Volans: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) (owner: 10Majavah) [08:01:56] !log arthurtaylor@deploy1003 helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply [08:03:25] 10SRE-SLO, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th): page_change SLO windows - https://phabricator.wikimedia.org/T438054#12348692 (10APizzata-WMF) a:05APizzata-WMF→03None [08:04:37] !log arthurtaylor@deploy1003 helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply [08:04:59] !log arthurtaylor@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply [08:05:04] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw [08:05:08] !log arthurtaylor@deploy1003 helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply [08:05:27] !log arthurtaylor@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply [08:06:18] (03CR) 10Muehlenhoff: [C:03+2] Remove the monthly_ganeti_rebalance timer [puppet] - 10https://gerrit.wikimedia.org/r/1343848 (owner: 10Muehlenhoff) [08:06:29] (03CR) 10Majavah: P:wmcs::etcd: Add profile to allow backing up etcd cluster data (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1343547 (https://phabricator.wikimedia.org/T438731) (owner: 10Majavah) [08:09:53] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [08:10:52] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [08:10:52] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw [08:11:05] (03PS2) 10Jelto: service::catalog: Set ipip for shellbox* in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343540 (https://phabricator.wikimedia.org/T420436) [08:11:23] RESOLVED: [2x] GnmiInterfaceCountersDrop: asw1-b4-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [08:13:29] (03CR) 10Volans: [C:03+1] "LGTM. suggestion for the tests" [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [08:18:35] !log installing gst-plugins-base1.0 security updates [08:18:36] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:19:50] (03CR) 10Volans: [C:03+1] P:wmcs::etcd: Add profile to allow backing up etcd cluster data (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1343547 (https://phabricator.wikimedia.org/T438731) (owner: 10Majavah) [08:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [08:22:00] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for shellbox* in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343540 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [08:22:34] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploy v2.0.0 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343912 (https://phabricator.wikimedia.org/T421814) [08:22:38] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [08:23:08] (03PS1) 10Trueg: WDQS: wdqs-proxy version bump to 0.9.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343913 [08:24:20] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [08:24:27] !log dcausse@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [08:25:58] (03PS1) 10Muehlenhoff: Stop installing component/jdk21 on build2002 [puppet] - 10https://gerrit.wikimedia.org/r/1343916 [08:26:22] !log ayounsi@cumin1004 START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27 [08:26:33] !log jelto@cumin1004 END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad [08:27:21] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27 [08:27:42] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343916 (owner: 10Muehlenhoff) [08:28:22] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [08:29:52] (03CR) 10Brouberol: [C:03+1] WDQS: wdqs-proxy version bump to 0.9.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343913 (owner: 10Trueg) [08:29:56] !log asw1-bw27-esams> request system reboot - T437984 [08:29:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:29:59] T437984: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984 [08:30:06] (03CR) 10Trueg: [C:03+2] WDQS: wdqs-proxy version bump to 0.9.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343913 (owner: 10Trueg) [08:30:43] !log jelto@cumin1004 END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad [08:30:45] (03PS1) 10Seanleong-wmde: Config change for launch of stopping sending LL notifications. [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343917 (https://phabricator.wikimedia.org/T437577) [08:31:30] (03PS2) 10Seanleong-wmde: Config change for launch of stopping sending LL notifications into group1. [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343917 (https://phabricator.wikimedia.org/T437577) [08:32:15] !log installig zip security updates [08:32:16] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:32:21] PROBLEM - Host bast3007 is DOWN: PING CRITICAL - Packet loss = 100% [08:32:43] (03Merged) 10jenkins-bot: WDQS: wdqs-proxy version bump to 0.9.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343913 (owner: 10Trueg) [08:32:51] PROBLEM - Host doh3006 is DOWN: PING CRITICAL - Packet loss = 100% [08:32:51] PROBLEM - Host doh3005 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:01] PROBLEM - Host durum3006 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:01] PROBLEM - Host hcaptcha-proxy3002 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:01] PROBLEM - Host durum3005 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:01] PROBLEM - Host hcaptcha-proxy3001 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:03] PROBLEM - Host ncredir3005 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:15] PROBLEM - Host tcp-proxy3001 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:15] PROBLEM - Host prometheus3004 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:15] PROBLEM - Host tcp-proxy3002 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:17] PROBLEM - Host ncredir3006 is DOWN: PING CRITICAL - Packet loss = 100% [08:33:23] PROBLEM - Router interfaces on mr1-esams is CRITICAL: CRITICAL: host 185.15.59.130, interfaces up: 34, down: 1, dormant: 0, excluded: 0, unused: 0: https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [08:33:27] PROBLEM - Host netflow3004 is DOWN: PING CRITICAL - Packet loss = 100% [08:34:19] 06SRE, 10Continuous-Integration-Infrastructure, 10observability, 05Goal, 06Release-Engineering-Team (Seen): Add Prometheus exporter to Jenkins instances - https://phabricator.wikimedia.org/T182759#12348805 (10ABran-WMF) [08:35:21] RECOVERY - Router interfaces on mr1-esams is OK: OK: host 185.15.59.130, interfaces up: 35, down: 0, dormant: 0, excluded: 0, unused: 0 https://wikitech.wikimedia.org/wiki/Network_monitoring%23Router_interface_down [08:35:57] RECOVERY - Host netflow3004 is UP: PING OK - Packet loss = 0%, RTA = 78.38 ms [08:35:59] RECOVERY - Host durum3006 is UP: PING OK - Packet loss = 0%, RTA = 78.19 ms [08:35:59] RECOVERY - Host durum3005 is UP: PING OK - Packet loss = 0%, RTA = 89.04 ms [08:36:02] switch is already back up [08:36:03] yeah [08:36:21] RECOVERY - Host doh3005 is UP: PING OK - Packet loss = 0%, RTA = 88.21 ms [08:36:21] RECOVERY - Host doh3006 is UP: PING OK - Packet loss = 0%, RTA = 81.98 ms [08:36:29] RECOVERY - Host hcaptcha-proxy3002 is UP: PING OK - Packet loss = 0%, RTA = 78.72 ms [08:36:31] RECOVERY - Host hcaptcha-proxy3001 is UP: PING OK - Packet loss = 0%, RTA = 89.46 ms [08:36:31] RECOVERY - Host ncredir3005 is UP: PING OK - Packet loss = 0%, RTA = 88.31 ms [08:36:41] RECOVERY - Host bast3007 is UP: PING OK - Packet loss = 0%, RTA = 88.25 ms [08:36:43] RECOVERY - Host prometheus3004 is UP: PING OK - Packet loss = 0%, RTA = 78.38 ms [08:36:43] RECOVERY - Host tcp-proxy3001 is UP: PING OK - Packet loss = 0%, RTA = 78.37 ms [08:36:43] RECOVERY - Host tcp-proxy3002 is UP: PING OK - Packet loss = 0%, RTA = 88.35 ms [08:36:45] RECOVERY - Host ncredir3006 is UP: PING OK - Packet loss = 0%, RTA = 88.28 ms [08:37:04] !log trueg@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [08:37:19] !log trueg@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [08:38:26] FIRING: [20x] ProbeDown: Ripe Atlas anchor atlas3001:80 is not returning HTTP 200 OK on port 80 - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:39:08] !log trueg@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [08:39:24] !log trueg@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [08:41:29] (03PS6) 10Filippo Giunchedi: wmcs: export prometheus metrics from wmcs-backup [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) [08:41:29] (03PS4) 10Filippo Giunchedi: backy2: run wmcs-backup metrics once an hour [puppet] - 10https://gerrit.wikimedia.org/r/1343348 (https://phabricator.wikimedia.org/T428893) [08:41:54] (03CR) 10Filippo Giunchedi: wmcs: export prometheus metrics from wmcs-backup (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [08:42:51] (03PS1) 10Kevin Bazira: ml-services: update tts-section-generator isvc to fold latin characters espeak cannot pronounce before synthesis [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343919 (https://phabricator.wikimedia.org/T438647) [08:43:38] (03CR) 10Awight: [C:03+1] Config change for launch of stopping sending LL notifications into group1. [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343917 (https://phabricator.wikimedia.org/T437577) (owner: 10Seanleong-wmde) [08:44:30] !log ayounsi@cumin1004 START - Cookbook sre.hosts.remove-downtime for 13 hosts [08:44:38] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts [08:44:39] (03PS2) 10Kevin Bazira: ml-services: update tts-section-generator service to fold latin characters espeak cannot pronounce before synthesis [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343919 (https://phabricator.wikimedia.org/T438647) [08:44:56] !log ayounsi@cumin1004 START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt [08:44:58] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt [08:45:15] 10ops-eqiad, 06DC-Ops, 10Prod-Kubernetes, 06ServiceOps, 07Kubernetes: wikikube-worker1152.eqiad.wmnet crashed - https://phabricator.wikimedia.org/T438819 (10Jelto) 03NEW [08:45:53] !log ayounsi@cumin1004 START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, T437984] [08:45:55] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, T437984] [08:45:56] T437984: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984 [08:48:59] (03CR) 10Volans: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [08:49:14] !log trueg@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply [08:50:23] (03PS1) 10Slyngshede: Service::catalog exclude rest-gateway and Sophroid from DC switch [puppet] - 10https://gerrit.wikimedia.org/r/1343920 (https://phabricator.wikimedia.org/T435443) [08:50:26] (03PS1) 10Slyngshede: service::catalog: exclude sessionstore from DC switchover [puppet] - 10https://gerrit.wikimedia.org/r/1343921 (https://phabricator.wikimedia.org/T435443) [08:51:41] (03CR) 10Filippo Giunchedi: [C:03+2] wmcs: export prometheus metrics from wmcs-backup [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [08:51:56] (03PS2) 10Gmodena: wdqs: add wikidata prefixes config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343616 (https://phabricator.wikimedia.org/T438476) [08:51:56] (03CR) 10Gmodena: "tagging as WIP because I did not fully test this yet (was waiting on proxy 0.9.)." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343616 (https://phabricator.wikimedia.org/T438476) (owner: 10Gmodena) [08:52:00] (03CR) 10Filippo Giunchedi: [C:03+2] backy2: run wmcs-backup metrics once an hour [puppet] - 10https://gerrit.wikimedia.org/r/1343348 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [08:52:36] !log trueg@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply [08:53:26] FIRING: [20x] ProbeDown: Ripe Atlas anchor atlas3001:80 is not returning HTTP 200 OK on port 80 - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:54:38] (03PS1) 10Ayounsi: Routed ganeti: add depool strategy [puppet] - 10https://gerrit.wikimedia.org/r/1343923 (https://phabricator.wikimedia.org/T327300) [08:55:13] !log trueg@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply [08:55:42] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1343923 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [08:56:12] (03CR) 10Cathal Mooney: [C:03+1] Routed ganeti: add depool strategy [puppet] - 10https://gerrit.wikimedia.org/r/1343923 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [08:57:10] (03CR) 10Ayounsi: [C:03+2] Routed ganeti: add depool strategy [puppet] - 10https://gerrit.wikimedia.org/r/1343923 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [08:57:29] (03CR) 10Elukey: [C:03+1] Only install python3-docker-report on active report host [puppet] - 10https://gerrit.wikimedia.org/r/1343906 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [08:58:48] !log trueg@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply [08:59:03] RESOLVED: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=logging-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [08:59:27] 10ops-eqiad, 06DC-Ops, 10Prod-Kubernetes, 06ServiceOps, 07Kubernetes: wikikube-worker1152.eqiad.wmnet crashed - https://phabricator.wikimedia.org/T438819#12348904 (10Jelto) Management interface is reachable but `getsel` has no recent entries, most recent entry is from 2025. But it has quite a lot of entr... [09:02:32] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1 [09:04:10] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1 [09:04:27] (03CR) 10Trueg: wdqs: add wikidata prefixes config (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343616 (https://phabricator.wikimedia.org/T438476) (owner: 10Gmodena) [09:04:29] (03CR) 10Vgutierrez: [C:04-1] D:prometheus::trafficserver_exporter: Metrics for cache evacuation (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1343309 (https://phabricator.wikimedia.org/T438627) (owner: 10Slyngshede) [09:05:13] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3 [09:06:38] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [09:07:45] 06SRE, 06Infrastructure-Foundations: Disable TUN/TAP outside of virtualisation roles - https://phabricator.wikimedia.org/T438459#12348939 (10fgiunchedi) Confirmed no `tap` module is loaded across cloudvps: ` root@cloudcumin1001:~# cumin -x -p 0 'O{*}' 'lsmod | grep ^tap' ... OK to proceed on 797 hosts? Enter... [09:08:19] fceratto@cumin1004 prepare (PID 27330) is awaiting input [09:09:48] !log ayounsi@cumin1004 START - Cookbook sre.network.peering with action 'email' for AS: 52320 [09:10:23] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12348948 (10Novem_Linguae) [09:10:25] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320 [09:11:12] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3 [09:11:25] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12348951 (10Novem_Linguae) Should we change the SRE tag to a more specific tag? Do the sessionstore servers have a team that maintains them? [09:11:38] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4 [09:12:54] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4 [09:13:32] (03PS1) 10Tiziano Fogli: Revert "kafka-logging1006: disable icinga notifications" [puppet] - 10https://gerrit.wikimedia.org/r/1343930 [09:14:11] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6 [09:15:16] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6 [09:15:41] 10ops-eqiad, 06DC-Ops, 10Prod-Kubernetes, 06ServiceOps, 07Kubernetes: wikikube-worker1152.eqiad.wmnet crashed - https://phabricator.wikimedia.org/T438819#12348990 (10Jelto) I powercycled the host. It restarted but the issues were not resolved. Some troubleshooting using the management interface revealed... [09:15:54] 10ops-eqiad, 06DC-Ops, 10Prod-Kubernetes, 06ServiceOps, 07Kubernetes: wikikube-worker1152.eqiad.wmnet networking issue - https://phabricator.wikimedia.org/T438819#12348991 (10Jelto) [09:16:38] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7 [09:20:43] !log jelto@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues [09:20:54] 10ops-eqiad, 06DC-Ops, 10Prod-Kubernetes, 06ServiceOps, 07Kubernetes: wikikube-worker1152.eqiad.wmnet networking issue - https://phabricator.wikimedia.org/T438819#12349025 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=68aa574d-c114-4b8b-92b4-6084ef33f327) set by jelto@cumin1004 for... [09:21:05] 06SRE, 06ServiceOps, 07Datacenter-Switchover: Services without a service IP cannot automatically be switched by the switchdc cookbook - https://phabricator.wikimedia.org/T285707#12349036 (10Blake) a:05Blake→03None This task is still valid. At the moment, we're using the 'exclude_from_switchover' field in... [09:22:26] (03CR) 10Btullis: [C:03+2] dse-k8s: Delegate reverse DNS for the second eqiad Pod range [dns] - 10https://gerrit.wikimedia.org/r/1343650 (https://phabricator.wikimedia.org/T430658) (owner: 10Btullis) [09:23:06] (03CR) 10Blake: [C:03+2] golang: add trixie-based golang-1.26 image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341880 (https://phabricator.wikimedia.org/T423851) (owner: 10Blake) [09:23:15] (03CR) 10Blake: [V:03+2 C:03+2] golang: add trixie-based golang-1.26 image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341880 (https://phabricator.wikimedia.org/T423851) (owner: 10Blake) [09:23:29] !log jelto@cumin1004 conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet [09:23:33] !log btullis@dns1004 START - running authdns-update [09:24:33] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [09:24:41] (03CR) 10Tiziano Fogli: [C:03+2] Revert "kafka-logging1006: disable icinga notifications" [puppet] - 10https://gerrit.wikimedia.org/r/1343930 (owner: 10Tiziano Fogli) [09:25:23] 06SRE, 10Data-Persistence-Backup, 10database-backups: Put db2201 back into backup production as a backup source - https://phabricator.wikimedia.org/T437411#12349052 (10Marostegui) ` | 42613 | dump.s5.2026-09-22--00-00-02 | finished | db2201.codfw.wmnet:3315 | dbprov2005.codfw.wmnet | dump | s5 |... [09:26:00] !log btullis@dns1004 END - running authdns-update [09:26:19] fceratto@cumin1004 prepare (PID 49841) is awaiting input [09:26:37] (03CR) 10Ozge: [C:03+1] ml-services: update tts-section-generator service to fold latin characters espeak cannot pronounce before synthesis [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343919 (https://phabricator.wikimedia.org/T438647) (owner: 10Kevin Bazira) [09:27:33] (03PS1) 10Marostegui: db2250: Remove s5 [puppet] - 10https://gerrit.wikimedia.org/r/1343932 (https://phabricator.wikimedia.org/T437411) [09:27:34] (03PS1) 10Elukey: CHANGELOG: add changelogs for release v13.3.0 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1343933 [09:28:23] !log jelto@cumin1004 END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad [09:28:41] !log jelto@cumin1004 conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet [09:28:41] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs1019 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [09:29:26] (03CR) 10Kevin Bazira: [C:03+2] ml-services: update tts-section-generator service to fold latin characters espeak cannot pronounce before synthesis [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343919 (https://phabricator.wikimedia.org/T438647) (owner: 10Kevin Bazira) [09:29:35] ^ thats me, cookbook is blocked by network issues on wikikube-worker1152, I'm on it [09:29:44] (03CR) 10Muehlenhoff: [C:03+2] Only install python3-docker-report on active report host [puppet] - 10https://gerrit.wikimedia.org/r/1343906 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [09:29:50] fceratto@cumin1004 prepare (PID 49841) is awaiting input [09:30:25] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [09:31:26] (03CR) 10Elukey: [C:03+2] CHANGELOG: add changelogs for release v13.3.0 [software/spicerack] - 10https://gerrit.wikimedia.org/r/1343933 (owner: 10Elukey) [09:31:57] !log marostegui@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations [09:32:29] (03Merged) 10jenkins-bot: ml-services: update tts-section-generator service to fold latin characters espeak cannot pronounce before synthesis [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343919 (https://phabricator.wikimedia.org/T438647) (owner: 10Kevin Bazira) [09:32:40] !log marostegui@cumin1004 START - Cookbook sre.mysql.depool depool es1035: issues [09:32:41] !log jelto@cumin1004 END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad [09:32:41] !log marostegui@cumin1004 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues [09:33:37] (03CR) 10Tiziano Fogli: [C:03+2] kafka-logging100[78]: disable icinga notifications [puppet] - 10https://gerrit.wikimedia.org/r/1342553 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [09:33:39] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for dpislaru - https://phabricator.wikimedia.org/T438827 (10DPislaru-WMF) 03NEW [09:34:13] PROBLEM - Check if Pybal has been restarted after pybal.conf was changed on lvs1020 is CRITICAL: CRITICAL: Service pybal.service has not been restarted after /etc/pybal/pybal.conf was changed (gt 1h). https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [09:34:53] (03PS1) 10Elukey: Upstream release v13.3.0 [software/spicerack] (debian) - 10https://gerrit.wikimedia.org/r/1343934 [09:35:10] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:36:38] (03CR) 10Elukey: [V:03+2 C:03+2] Upstream release v13.3.0 [software/spicerack] (debian) - 10https://gerrit.wikimedia.org/r/1343934 (owner: 10Elukey) [09:38:02] (03PS1) 10Federico Ceratto: switchdc: add sleep before checking replication [cookbooks] - 10https://gerrit.wikimedia.org/r/1343935 (https://phabricator.wikimedia.org/T436500) [09:39:20] elukey@cumin1004 reimage (PID 87165) is awaiting input [09:39:24] !log fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - T438828 [09:39:26] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:39:27] T438828: Upgrade HAProxy to 3.2.23 on cp hosts - https://phabricator.wikimedia.org/T438828 [09:39:50] (03PS2) 10Federico Ceratto: switchdc: add sleep before checking replication [cookbooks] - 10https://gerrit.wikimedia.org/r/1343935 (https://phabricator.wikimedia.org/T436500) [09:40:19] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7 [09:40:22] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1343486 (https://phabricator.wikimedia.org/T438034) (owner: 10Tiziano Fogli) [09:40:29] (03PS1) 10Jelto: remove wikikube-worker1152 from wikikube eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343936 (https://phabricator.wikimedia.org/T438819) [09:41:25] (03PS1) 10Elukey: profile::docker_registry: remove check for port 5001 [puppet] - 10https://gerrit.wikimedia.org/r/1343937 [09:41:28] !log elukey@cumin1004 START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie [09:41:52] !log kevinbazira@deploy1003 helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [09:42:53] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7 [09:42:53] (03CR) 10Marostegui: [C:03+1] switchdc: add sleep before checking replication [cookbooks] - 10https://gerrit.wikimedia.org/r/1343935 (https://phabricator.wikimedia.org/T436500) (owner: 10Federico Ceratto) [09:42:57] (03CR) 10JMeybohm: [C:03+1] profile::docker_registry: remove check for port 5001 [puppet] - 10https://gerrit.wikimedia.org/r/1343937 (owner: 10Elukey) [09:43:17] !log kevinbazira@deploy1003 helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . [09:44:07] !log kevinbazira@deploy1003 helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [09:44:23] !log vgutierrez@cumin1004 START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P{cp[7010,7016].*} and A:cp - 3.2.23 upgrade (T438828) [09:44:43] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7 [09:45:21] (03CR) 10Elukey: [C:03+2] profile::docker_registry: remove check for port 5001 [puppet] - 10https://gerrit.wikimedia.org/r/1343937 (owner: 10Elukey) [09:45:29] (03CR) 10Elukey: [C:03+2] profile::k8s::deployment_server: add python3-docker-report [puppet] - 10https://gerrit.wikimedia.org/r/1341899 (https://phabricator.wikimedia.org/T437297) (owner: 10Elukey) [09:45:46] (03PS1) 10Muehlenhoff: Explain build2002/2004 in site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1343938 [09:46:58] !log marostegui@cumin1004 START - Cookbook sre.mysql.pool pool es1035: issues [09:47:01] !log marostegui@cumin1004 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues [09:47:16] (03CR) 10Urbanecm: [C:03+1] Growth: Set minimum registration date for A/B test [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343611 (https://phabricator.wikimedia.org/T432126) (owner: 10Michael Große) [09:47:17] !log uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia [09:47:18] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:47:41] (03PS16) 10Tiziano Fogli: kafka-logging: add kafka-logging100[7-8] to eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) [09:47:41] (03PS4) 10Tiziano Fogli: kafka-logging: remove kafka-logging100[12] [puppet] - 10https://gerrit.wikimedia.org/r/1342557 (https://phabricator.wikimedia.org/T432444) [09:47:41] (03PS1) 10Tiziano Fogli: kafka/logging: set up legacy iptables instead of nftables [puppet] - 10https://gerrit.wikimedia.org/r/1343939 (https://phabricator.wikimedia.org/T432444) [09:50:58] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6 [09:51:03] (03CR) 10Tiziano Fogli: [C:03+2] kafka/logging: set up legacy iptables instead of nftables [puppet] - 10https://gerrit.wikimedia.org/r/1343939 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [09:51:16] !log install spicerack 13.3.0 on cumin1004 and cumin2003 [09:51:17] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:51:38] (03CR) 10Majavah: [C:03+2] P:wmcs::etcd: Add profile to allow backing up etcd cluster data [puppet] - 10https://gerrit.wikimedia.org/r/1343547 (https://phabricator.wikimedia.org/T438731) (owner: 10Majavah) [09:52:39] (03CR) 10Filippo Giunchedi: [C:03+2] "Thank you for review!" [puppet] - 10https://gerrit.wikimedia.org/r/1338132 (https://phabricator.wikimedia.org/T437272) (owner: 10Filippo Giunchedi) [09:52:56] (03CR) 10Majavah: [C:03+2] O:wmcs::toolforge: Add role for etcd backups [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) (owner: 10Majavah) [09:53:18] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6 [09:53:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:53:47] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Automate PDU Deployment Process - https://phabricator.wikimedia.org/T403173#12349269 (10ayounsi) p:05Medium→03Low I feel like as this is something not done at a high frequency, the priority to fully automate them will be lower than other proje... [09:55:00] 06SRE, 06Infrastructure-Foundations, 06Traffic, 13Patch-For-Review: GeoIP mapping experiments - https://phabricator.wikimedia.org/T332024#12349276 (10LSobanski) @CDanis are you still planning to work on this or should it be unassigned? [09:55:31] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5 [09:55:53] 07Puppet, 06Infrastructure-Foundations: Puppetmaster volatile data not synced to all puppet frontends for a month and a half (2024-04-27 to 2024-06-10) - https://phabricator.wikimedia.org/T367113#12349284 (10LSobanski) @CDanis are you still planning to work on this or should it be unassigned? [09:56:07] 07Puppet, 06Infrastructure-Foundations: Puppetmaster volatile data not synced to all puppet frontends for a month and a half (2024-04-27 to 2024-06-10) - https://phabricator.wikimedia.org/T367113#12349285 (10LSobanski) [09:56:48] (03PS1) 10Muehlenhoff: Fix and remove broken Hiera variable profile::docker::builder::docker_pkg [puppet] - 10https://gerrit.wikimedia.org/r/1343941 (https://phabricator.wikimedia.org/T417389) [09:56:52] !log vgutierrez@cumin1004 END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P{cp[7010,7016].*} and A:cp - 3.2.23 upgrade (T438828) [09:56:55] T438828: Upgrade HAProxy to 3.2.23 on cp hosts - https://phabricator.wikimedia.org/T438828 [09:57:33] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5 [09:57:53] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343941 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [09:58:06] !log elukey@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage [09:58:26] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2 [09:59:17] (03PS1) 10Filippo Giunchedi: wmcs: fix wmcs-backup interval [puppet] - 10https://gerrit.wikimedia.org/r/1343942 (https://phabricator.wikimedia.org/T428893) [10:00:05] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1000) [10:00:10] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2 [10:01:06] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3 [10:02:48] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3 [10:02:52] !log elukey@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage [10:02:55] (03CR) 10Filippo Giunchedi: [C:03+2] "Trivial, self-merging" [puppet] - 10https://gerrit.wikimedia.org/r/1343942 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [10:03:50] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7 [10:03:55] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: eqiad row A&B host migration details request for Infrastructure Foundations - https://phabricator.wikimedia.org/T432646#12349335 (10LSobanski) a:05LSobanski→03None [10:05:43] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7 [10:06:06] (03CR) 10Ayounsi: [C:03+1] Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [10:06:30] !log installing libcap2 security updates [10:06:31] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:06:44] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8 [10:08:18] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8 [10:08:58] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [10:09:07] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [10:09:09] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4 [10:09:47] (03CR) 10Tiziano Fogli: [C:03+2] admin/data: grant access to hany (analytics_privatedata_users l1) [puppet] - 10https://gerrit.wikimedia.org/r/1343486 (https://phabricator.wikimedia.org/T438034) (owner: 10Tiziano Fogli) [10:09:54] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [10:10:06] (03PS1) 10Cathal Mooney: Disable HE paths to eqsin due to significant packet loss [homer/public] - 10https://gerrit.wikimedia.org/r/1343944 (https://phabricator.wikimedia.org/T438835) [10:10:06] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [10:10:44] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to Superset Dashboard for Hany EL Mokadem - https://phabricator.wikimedia.org/T438034#12349377 (10tappof) 05Open→03Resolved a:03tappof Patch merged. Access granted. [10:10:56] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4 [10:11:51] !log fceratto@cumin1004 START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1 [10:11:57] (03PS3) 10Blake: bird: add a new prometheus-exporter for bird. [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1343943 [10:13:35] !log fceratto@cumin1004 END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1 [10:14:22] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for AHan-WMF - https://phabricator.wikimedia.org/T438538#12349398 (10tappof) [10:14:27] (03PS4) 10Blake: bird: add a new prometheus-exporter for bird. [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1343943 (https://phabricator.wikimedia.org/T423851) [10:17:28] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:18:19] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for AHan-WMF - https://phabricator.wikimedia.org/T438538#12349418 (10tappof) @AHan-WMF, do you also need a kerberos principal? https://wikitech.wikimedia.org/wiki/Data_Platform/Data_access#Access_Levels Thanks [10:18:26] FIRING: [3x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://wikitech.wikimedia.org/wiki/Docker - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:18:45] (03CR) 10Federico Ceratto: [C:03+2] switchdc: add sleep before checking replication [cookbooks] - 10https://gerrit.wikimedia.org/r/1343935 (https://phabricator.wikimedia.org/T436500) (owner: 10Federico Ceratto) [10:20:22] !log elukey@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie [10:21:25] 06SRE, 06Infrastructure-Foundations, 10netops: HE Packet Loss Issues - https://phabricator.wikimedia.org/T438836 (10cmooney) 03NEW p:05Triage→03High [10:21:33] 06SRE, 06Infrastructure-Foundations, 10netops: HE Packet Loss Issues - https://phabricator.wikimedia.org/T438836#12349437 (10cmooney) [10:23:26] RESOLVED: [3x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://wikitech.wikimedia.org/wiki/Docker - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:26:50] (03PS3) 10Hnowlan: karma: consistently order some labels first [puppet] - 10https://gerrit.wikimedia.org/r/1310519 (https://phabricator.wikimedia.org/T431980) [10:27:32] (03PS2) 10Hnowlan: karma: strip some useless labels from panel display [puppet] - 10https://gerrit.wikimedia.org/r/1311029 (https://phabricator.wikimedia.org/T431980) [10:27:38] (03CR) 10Hnowlan: "Done!" [puppet] - 10https://gerrit.wikimedia.org/r/1311029 (https://phabricator.wikimedia.org/T431980) (owner: 10Hnowlan) [10:28:09] (03CR) 10Hnowlan: karma: consistently order some labels first (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1310519 (https://phabricator.wikimedia.org/T431980) (owner: 10Hnowlan) [10:29:56] (03CR) 10Btullis: [C:03+2] k8s: Let kube-proxy detect local Pod traffic by interface name [puppet] - 10https://gerrit.wikimedia.org/r/1343023 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [10:34:37] (03PS1) 10Brouberol: ceph/osd: add a script allowing the mapping of OSD to physical disk location [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) [10:35:31] (03CR) 10Brouberol: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9462/co" [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [10:35:42] (03CR) 10CI reject: [V:04-1] ceph/osd: add a script allowing the mapping of OSD to physical disk location [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [10:35:48] (03CR) 10Brouberol: "Example of the script running on cephosd1002:" [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [10:37:38] (03PS2) 10Brouberol: ceph/osd: add a script allowing the mapping of OSD to physical disk location [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) [10:38:29] (03CR) 10Cathal Mooney: [C:03+2] Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [10:41:45] (03CR) 10Btullis: [C:03+1] "Looks good to me. Just one suggestion for a comment in the puppet part." [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [10:41:58] (03Merged) 10jenkins-bot: Add alert on percentage of timeouts from POPs in blackbox pings [alerts] - 10https://gerrit.wikimedia.org/r/1328168 (https://phabricator.wikimedia.org/T435592) (owner: 10Cathal Mooney) [10:44:54] jouncebot: nowandnext [10:44:54] For the next 0 hour(s) and 15 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1000) [10:44:54] In 1 hour(s) and 15 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1200) [10:45:14] I'd like to sync a wmf.21 patch [10:45:29] (03PS1) 10Kosta Harlan: AbuseReview: Allow interaction with verdict buttons on closed rows [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343952 (https://phabricator.wikimedia.org/T438808) [10:48:49] (03PS1) 10A-pizzata: Add weekly full snapshot of the actor table [puppet] - 10https://gerrit.wikimedia.org/r/1343954 (https://phabricator.wikimedia.org/T437961) [10:49:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy1003 using scap backport" [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343952 (https://phabricator.wikimedia.org/T438808) (owner: 10Kosta Harlan) [10:50:14] (03CR) 10Btullis: [C:03+2] hieradata: Detect local Pod traffic by interface name on dse-k8s-eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343024 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [10:51:00] (03Merged) 10jenkins-bot: AbuseReview: Allow interaction with verdict buttons on closed rows [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343952 (https://phabricator.wikimedia.org/T438808) (owner: 10Kosta Harlan) [10:51:15] (03CR) 10CI reject: [V:04-1] Add weekly full snapshot of the actor table [puppet] - 10https://gerrit.wikimedia.org/r/1343954 (https://phabricator.wikimedia.org/T437961) (owner: 10A-pizzata) [10:51:27] !log kharlan@deploy1003 Started scap sync-world: Backport for [[gerrit:1343952|AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] [10:51:30] T438808: AbuseReview: Allow verdict buttons to be interactive without expanding the row - https://phabricator.wikimedia.org/T438808 [10:52:31] !log enable rule cache-upload/eqsin_originals_scraper_20260922 [10:52:32] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:52:33] (03PS1) 10Hnowlan: zuul: migrate icinga checks to prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1343955 (https://phabricator.wikimedia.org/T384939) [10:53:15] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [10:53:28] !log gmodena@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [10:54:15] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [10:54:17] (03CR) 10Slyngshede: [C:03+1] Stop installing component/jdk21 on build2002 [puppet] - 10https://gerrit.wikimedia.org/r/1343916 (owner: 10Muehlenhoff) [10:54:24] !log gmodena@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [10:57:51] (03CR) 10Hnowlan: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343955 (https://phabricator.wikimedia.org/T384939) (owner: 10Hnowlan) [11:01:02] (03PS3) 10Jelto: sre.loadbalancer.migrate-service-ipip: skip depooled host in validation [cookbooks] - 10https://gerrit.wikimedia.org/r/1343947 (https://phabricator.wikimedia.org/T438819) [11:01:44] (03PS1) 10Kosta Harlan: AbuseReview: Hide Echo banner when user cannot see personal info [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343960 (https://phabricator.wikimedia.org/T438477) [11:02:06] (03PS2) 10Daniel Kertesz: cache::haproxy: implement support for persistent stats [puppet] - 10https://gerrit.wikimedia.org/r/1343039 (https://phabricator.wikimedia.org/T343000) [11:05:00] FIRING: PingLossPercent: Blackbox probe packet loss 5.417% from codfw to eqsin - https://wikitech.wikimedia.org/wiki/Network_monitoring#PingLossPercent - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/?from=now-3h&to=now&var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DPingLossPercent [11:05:33] (03PS2) 10Hnowlan: zuul: migrate icinga checks to prometheus [puppet] - 10https://gerrit.wikimedia.org/r/1343955 (https://phabricator.wikimedia.org/T384939) [11:05:56] (03CR) 10Hnowlan: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343955 (https://phabricator.wikimedia.org/T384939) (owner: 10Hnowlan) [11:06:01] (03PS1) 10Majavah: P:wmcs::etcd: Rename nginx upstream [puppet] - 10https://gerrit.wikimedia.org/r/1343961 (https://phabricator.wikimedia.org/T410721) [11:06:04] (03PS1) 10Majavah: P:etcd::tlsproxy: Move hardcoded ACL exceptions to Hiera [puppet] - 10https://gerrit.wikimedia.org/r/1343962 (https://phabricator.wikimedia.org/T410721) [11:06:06] (03PS1) 10Majavah: P:etcd::tlsproxy: Add support for acme-chief certificates [puppet] - 10https://gerrit.wikimedia.org/r/1343963 (https://phabricator.wikimedia.org/T410721) [11:06:09] (03PS1) 10Majavah: P:etcd::tlsproxy: Support client cert auth to upstream [puppet] - 10https://gerrit.wikimedia.org/r/1343964 (https://phabricator.wikimedia.org/T410721) [11:07:00] (03PS2) 10A-pizzata: sqoop_mediawiki: Add weekly full snapshot of the actor table [puppet] - 10https://gerrit.wikimedia.org/r/1343954 (https://phabricator.wikimedia.org/T437961) [11:07:50] (03CR) 10CI reject: [V:04-1] P:etcd::tlsproxy: Add support for acme-chief certificates [puppet] - 10https://gerrit.wikimedia.org/r/1343963 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [11:08:30] (03CR) 10JMeybohm: sre.loadbalancer.migrate-service-ipip: skip depooled host in validation (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1343947 (https://phabricator.wikimedia.org/T438819) (owner: 10Jelto) [11:08:31] (03CR) 10CI reject: [V:04-1] P:etcd::tlsproxy: Support client cert auth to upstream [puppet] - 10https://gerrit.wikimedia.org/r/1343964 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [11:09:50] (03CR) 10Cathal Mooney: [C:03+2] Disable HE paths to eqsin due to significant packet loss [homer/public] - 10https://gerrit.wikimedia.org/r/1343944 (https://phabricator.wikimedia.org/T438835) (owner: 10Cathal Mooney) [11:09:51] (03PS2) 10Majavah: P:etcd::tlsproxy: Add support for acme-chief certificates [puppet] - 10https://gerrit.wikimedia.org/r/1343963 (https://phabricator.wikimedia.org/T410721) [11:09:51] (03PS2) 10Majavah: P:etcd::tlsproxy: Support client cert auth to upstream [puppet] - 10https://gerrit.wikimedia.org/r/1343964 (https://phabricator.wikimedia.org/T410721) [11:11:01] (03Merged) 10jenkins-bot: Disable HE paths to eqsin due to significant packet loss [homer/public] - 10https://gerrit.wikimedia.org/r/1343944 (https://phabricator.wikimedia.org/T438835) (owner: 10Cathal Mooney) [11:12:03] !log kharlan@deploy1003 kharlan: Backport for [[gerrit:1343952|AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:12:07] T438808: AbuseReview: Allow verdict buttons to be interactive without expanding the row - https://phabricator.wikimedia.org/T438808 [11:12:51] (03CR) 10CI reject: [V:04-1] P:etcd::tlsproxy: Add support for acme-chief certificates [puppet] - 10https://gerrit.wikimedia.org/r/1343963 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [11:13:05] !log kharlan@deploy1003 kharlan: Continuing with deployment [11:13:49] (03CR) 10CI reject: [V:04-1] P:etcd::tlsproxy: Support client cert auth to upstream [puppet] - 10https://gerrit.wikimedia.org/r/1343964 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [11:15:00] (03PS3) 10Majavah: P:etcd::tlsproxy: Add support for acme-chief certificates [puppet] - 10https://gerrit.wikimedia.org/r/1343963 (https://phabricator.wikimedia.org/T410721) [11:15:00] (03PS3) 10Majavah: P:etcd::tlsproxy: Support client cert auth to upstream [puppet] - 10https://gerrit.wikimedia.org/r/1343964 (https://phabricator.wikimedia.org/T410721) [11:16:13] (03PS2) 10JMeybohm: gVisor: disable on bookworm in containerd [puppet] - 10https://gerrit.wikimedia.org/r/1342961 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [11:16:17] (03CR) 10JMeybohm: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342961 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [11:16:25] (03PS2) 10JMeybohm: gVisor: use kvm as a platform when hardware virtualization is available [puppet] - 10https://gerrit.wikimedia.org/r/1342962 (https://phabricator.wikimedia.org/T438287) (owner: 10Giuseppe Lavagetto) [11:19:03] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [11:22:04] (03CR) 10Atsuko: [C:03+1] admin_ng: pin the ceph-csi charts to the versions in production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343130 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [11:22:27] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [11:22:33] (03CR) 10Atsuko: [C:03+1] admin_ng: move dse-k8s-codfw to ceph-csi-rbd v3.14.2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343624 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [11:22:47] RECOVERY - Check if Pybal has been restarted after pybal.conf was changed on lvs1019 is OK: OK: pybal.service was restarted after /etc/pybal/pybal.conf was changed. https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [11:23:19] RECOVERY - Check if Pybal has been restarted after pybal.conf was changed on lvs1020 is OK: OK: pybal.service was restarted after /etc/pybal/pybal.conf was changed. https://wikitech.wikimedia.org/wiki/PyBal%23Pybal_service_has_not_been_restarted [11:24:30] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [11:24:30] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad [11:24:36] !log kharlan@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343952|AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s) [11:24:39] T438808: AbuseReview: Allow verdict buttons to be interactive without expanding the row - https://phabricator.wikimedia.org/T438808 [11:24:48] continuing with another backport [11:25:20] (03CR) 10Atsuko: [C:03+1] admin_ng: move dse-k8s-codfw to ceph-csi-cephfs v3.14.2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343625 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [11:25:33] (03Abandoned) 10Jelto: remove wikikube-worker1152 from wikikube eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343936 (https://phabricator.wikimedia.org/T438819) (owner: 10Jelto) [11:25:40] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy1003 using scap backport" [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343960 (https://phabricator.wikimedia.org/T438477) (owner: 10Kosta Harlan) [11:26:09] (03CR) 10Gmodena: wdqs: add wikidata prefixes config (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343616 (https://phabricator.wikimedia.org/T438476) (owner: 10Gmodena) [11:26:51] (03Merged) 10jenkins-bot: AbuseReview: Hide Echo banner when user cannot see personal info [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343960 (https://phabricator.wikimedia.org/T438477) (owner: 10Kosta Harlan) [11:27:17] !log kharlan@deploy1003 Started scap sync-world: Backport for [[gerrit:1343960|AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] [11:27:21] T438477: AbuseReview: User without access to personal info tag can trigger error when using Echo notification query param - https://phabricator.wikimedia.org/T438477 [11:29:08] (03PS7) 10Giuseppe Lavagetto: haproxy: modularize hiddenparma support [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) [11:30:00] RESOLVED: PingLossPercent: Blackbox probe packet loss 3.542% from codfw to eqsin - https://wikitech.wikimedia.org/wiki/Network_monitoring#PingLossPercent - https://grafana.wikimedia.org/d/c4eb2910-ee2b-459b-8df6-369803085a1b/?from=now-3h&to=now&var-site=eqsin - https://alerts.wikimedia.org/?q=alertname%3DPingLossPercent [11:30:03] (03PS2) 10Kosta Harlan: AbuseReview: Hide recently saved revisions from the vandalism queue [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343953 (https://phabricator.wikimedia.org/T438235) [11:30:33] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (CORE_DIFF 2): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9464/co" [puppet] - 10https://gerrit.wikimedia.org/r/1309128 (https://phabricator.wikimedia.org/T422235) (owner: 10Giuseppe Lavagetto) [11:31:59] (03PS3) 10JMeybohm: gVisor: use kvm as a platform when hardware virtualization is available [puppet] - 10https://gerrit.wikimedia.org/r/1342962 (https://phabricator.wikimedia.org/T438287) (owner: 10Giuseppe Lavagetto) [11:32:11] (03PS2) 10JMeybohm: kubernetes::node: add gvisor label based on hardware virtualization [puppet] - 10https://gerrit.wikimedia.org/r/1342963 (https://phabricator.wikimedia.org/T438287) (owner: 10Giuseppe Lavagetto) [11:33:38] !log kharlan@deploy1003 kharlan: Backport for [[gerrit:1343960|AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [11:33:41] T438477: AbuseReview: User without access to personal info tag can trigger error when using Echo notification query param - https://phabricator.wikimedia.org/T438477 [11:34:07] !log kharlan@deploy1003 kharlan: Continuing with deployment [11:36:32] (03CR) 10Volans: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1343961 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [11:40:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 18.75% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:41:04] !log kharlan@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343960|AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s) [11:41:07] T438477: AbuseReview: User without access to personal info tag can trigger error when using Echo notification query param - https://phabricator.wikimedia.org/T438477 [11:44:19] and one more... [11:44:37] (03CR) 10TrainBranchBot: [C:03+2] "Approved by kharlan@deploy1003 using scap backport" [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343953 (https://phabricator.wikimedia.org/T438235) (owner: 10Kosta Harlan) [11:46:08] (03Merged) 10jenkins-bot: AbuseReview: Hide recently saved revisions from the vandalism queue [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343953 (https://phabricator.wikimedia.org/T438235) (owner: 10Kosta Harlan) [11:46:31] !log kharlan@deploy1003 Started scap sync-world: Backport for [[gerrit:1343953|AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] [11:46:38] T438235: AbuseReview: Implement filter for time delay view of edits - https://phabricator.wikimedia.org/T438235 [11:50:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.18% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:52:26] !log jnuche@deploy1003 Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): T435791 to backup host [11:52:32] !log btullis@cumin1004 START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{dse-k8s-worker1001.eqiad.wmnet} and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad) [11:52:35] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet [11:52:54] !log jnuche@deploy1003 Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): T435791 to backup host (duration: 01m 01s) [11:54:12] !log jnuche@deploy1003 Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): T435791 to production host [11:54:50] !log jnuche@deploy1003 Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): T435791 to production host (duration: 00m 54s) [11:55:10] (03CR) 10Btullis: [C:03+2] admin_ng: pin the ceph-csi charts to the versions in production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343130 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [11:56:37] (03PS3) 10Brouberol: ceph/osd: add a script allowing the mapping of OSD to physical disk location [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) [11:56:38] (03CR) 10Brouberol: ceph/osd: add a script allowing the mapping of OSD to physical disk location (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [11:57:25] (03CR) 10Muehlenhoff: [C:03+2] Stop installing component/jdk21 on build2002 [puppet] - 10https://gerrit.wikimedia.org/r/1343916 (owner: 10Muehlenhoff) [11:57:34] (03CR) 10Btullis: [C:03+1] ceph/osd: add a script allowing the mapping of OSD to physical disk location [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [11:57:52] (03CR) 10Brouberol: [C:03+2] ceph/osd: add a script allowing the mapping of OSD to physical disk location [puppet] - 10https://gerrit.wikimedia.org/r/1343949 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:00:05] Deploy window Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1200) [12:00:39] (03PS1) 10Cathal Mooney: Urldownloader: expand squid ACL to all production nets + analytics [puppet] - 10https://gerrit.wikimedia.org/r/1343970 (https://phabricator.wikimedia.org/T437560) [12:01:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.28% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:05:20] (03Merged) 10jenkins-bot: admin_ng: pin the ceph-csi charts to the versions in production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343130 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [12:06:13] (03PS1) 10Brouberol: cephosd1002: absent the osd associated with the faulty /dev/sdo device [puppet] - 10https://gerrit.wikimedia.org/r/1343971 (https://phabricator.wikimedia.org/T438823) [12:06:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.55% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [12:06:46] !log kharlan@deploy1003 kharlan: Backport for [[gerrit:1343953|AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [12:06:49] T438235: AbuseReview: Implement filter for time delay view of edits - https://phabricator.wikimedia.org/T438235 [12:07:05] (03CR) 10Brouberol: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9465/co" [puppet] - 10https://gerrit.wikimedia.org/r/1343971 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:08:03] !log kharlan@deploy1003 kharlan: Continuing with deployment [12:15:29] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:16:06] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:16:16] (03CR) 10Btullis: [C:03+1] cephosd1002: absent the osd associated with the faulty /dev/sdo device [puppet] - 10https://gerrit.wikimedia.org/r/1343971 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:16:23] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [12:16:51] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [12:17:47] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [12:18:02] (03CR) 10Brouberol: [V:03+1 C:03+2] cephosd1002: absent the osd associated with the faulty /dev/sdo device [puppet] - 10https://gerrit.wikimedia.org/r/1343971 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:19:19] PROBLEM - Host pki1002 is DOWN: PING CRITICAL - Packet loss = 100% [12:19:33] !log kharlan@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343953|AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s) [12:19:36] T438235: AbuseReview: Implement filter for time delay view of edits - https://phabricator.wikimedia.org/T438235 [12:19:54] (03CR) 10Muehlenhoff: [C:03+2] Explain build2002/2004 in site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1343938 (owner: 10Muehlenhoff) [12:20:58] (03CR) 10Elukey: [C:03+1] Fix and remove broken Hiera variable profile::docker::builder::docker_pkg [puppet] - 10https://gerrit.wikimedia.org/r/1343941 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [12:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [12:22:53] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet [12:23:26] FIRING: [44x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [12:26:27] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 28 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343917 (https://phabricator.wikimedia.org/T437577) (owner: 10Seanleong-wmde) [12:28:11] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 22 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343611 (https://phabricator.wikimedia.org/T432126) (owner: 10Michael Große) [12:28:37] (03CR) 10Muehlenhoff: [C:03+2] Fix and remove broken Hiera variable profile::docker::builder::docker_pkg [puppet] - 10https://gerrit.wikimedia.org/r/1343941 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [12:29:00] (03PS1) 10Brouberol: ceph::osd: fix missing semicolon in ceph-osd-remove exec [puppet] - 10https://gerrit.wikimedia.org/r/1343974 (https://phabricator.wikimedia.org/T438823) [12:29:25] (03CR) 10Klausman: [C:03+1] ceph::osd: fix missing semicolon in ceph-osd-remove exec [puppet] - 10https://gerrit.wikimedia.org/r/1343974 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:30:02] (03CR) 10CI reject: [V:04-1] ceph::osd: fix missing semicolon in ceph-osd-remove exec [puppet] - 10https://gerrit.wikimedia.org/r/1343974 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:31:01] (03PS2) 10Brouberol: ceph::osd: fix missing semicolon in ceph-osd-remove exec [puppet] - 10https://gerrit.wikimedia.org/r/1343974 (https://phabricator.wikimedia.org/T438823) [12:31:04] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet [12:31:05] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet [12:31:05] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{dse-k8s-worker1001.eqiad.wmnet} and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad) [12:32:13] FIRING: [2x] JobUnavailable: Reduced availability for job cfssl in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [12:32:23] (03CR) 10Brouberol: [C:03+2] ceph::osd: fix missing semicolon in ceph-osd-remove exec [puppet] - 10https://gerrit.wikimedia.org/r/1343974 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:32:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [12:40:17] (03CR) 10CWilliams: [C:03+2] db2250: Remove s5 [puppet] - 10https://gerrit.wikimedia.org/r/1343932 (https://phabricator.wikimedia.org/T437411) (owner: 10Marostegui) [12:40:18] (03CR) 10Cathal Mooney: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343970 (https://phabricator.wikimedia.org/T437560) (owner: 10Cathal Mooney) [12:41:02] (03CR) 10CWilliams: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1343932 (https://phabricator.wikimedia.org/T437411) (owner: 10Marostegui) [12:41:23] RESOLVED: CertAlmostExpired: gNMI TLS certificate for lsw1-c5-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [12:43:25] (03CR) 10Marostegui: [C:03+2] db2250: Remove s5 [puppet] - 10https://gerrit.wikimedia.org/r/1343932 (https://phabricator.wikimedia.org/T437411) (owner: 10Marostegui) [12:43:35] !log marostegui@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations [12:44:00] !log Stop mariadb on db2250:s5 T437411 T437279 [12:44:03] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:44:04] T437411: Put db2201 back into backup production as a backup source - https://phabricator.wikimedia.org/T437411 [12:44:04] T437279: Setup x4 backups - https://phabricator.wikimedia.org/T437279 [12:44:48] (03CR) 10Volans: [C:03+1] "LGTM, PCC looks good too:" [puppet] - 10https://gerrit.wikimedia.org/r/1343962 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [12:45:22] (03PS1) 10Brouberol: ceph::osd: fix if condition in remove-osd exec [puppet] - 10https://gerrit.wikimedia.org/r/1343977 (https://phabricator.wikimedia.org/T438823) [12:45:31] 06SRE, 10Data-Persistence-Backup, 10database-backups, 13Patch-For-Review: Put db2201 back into backup production as a backup source - https://phabricator.wikimedia.org/T437411#12349905 (10Marostegui) 05Open→03Resolved a:03Marostegui Mariadb has been stopped on db2250. I am going to consider this... [12:46:01] (03CR) 10Btullis: [C:03+1] ceph::osd: fix if condition in remove-osd exec [puppet] - 10https://gerrit.wikimedia.org/r/1343977 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:46:39] (03CR) 10Brouberol: [C:03+2] ceph::osd: fix if condition in remove-osd exec [puppet] - 10https://gerrit.wikimedia.org/r/1343977 (https://phabricator.wikimedia.org/T438823) (owner: 10Brouberol) [12:47:38] jouncebot: nowandnext [12:47:38] For the next 0 hour(s) and 12 minute(s): Mobileapps/RESTBase/Wikifeeds (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1200) [12:47:38] In 1 hour(s) and 12 minute(s): Southward Datacenter Switchover: Services + Traffic (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1400) [12:47:45] FIRING: [2x] WidespreadPuppetFailure: Puppet has failed in esams - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [12:47:56] (03CR) 10Volans: [C:03+1] "LGTM, PCC clean too:" [puppet] - 10https://gerrit.wikimedia.org/r/1343963 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [12:48:54] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet [12:49:03] 06SRE, 06Infrastructure-Foundations, 06Traffic, 13Patch-For-Review: GeoIP mapping experiments - https://phabricator.wikimedia.org/T332024#12349921 (10ssingh) >>! In T332024#12349276, @LSobanski wrote: > @CDanis are you still planning to work on this or should it be unassigned? @CDobbins has been working o... [12:51:10] (03CR) 10Volans: [C:03+1] "LGTM, PCC shows an empty line diff (see inline)" [puppet] - 10https://gerrit.wikimedia.org/r/1343964 (https://phabricator.wikimedia.org/T410721) (owner: 10Majavah) [12:52:40] (03PS7) 10Fabfur: profile:haproxy: kapow score setup [puppet] - 10https://gerrit.wikimedia.org/r/1328543 (https://phabricator.wikimedia.org/T435799) [12:52:45] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [12:53:55] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet [12:54:01] (03CR) 10KartikMistry: [C:03+2] machinetranslation: staging: Update to 2026-09-21-112314-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343502 (https://phabricator.wikimedia.org/T437213) (owner: 10KartikMistry) [12:55:23] (03CR) 10Lerickson: "The first line in that table with heading Configuration Property lists the default value to true, and we do not (to my knowledge!) change " [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343643 (https://phabricator.wikimedia.org/T438789) (owner: 10Lerickson) [12:56:22] (03Merged) 10jenkins-bot: machinetranslation: staging: Update to 2026-09-21-112314-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343502 (https://phabricator.wikimedia.org/T437213) (owner: 10KartikMistry) [12:56:48] Hello, I need to merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327496 and then run a deploy, since it modifies files under "/etc/helmfile-defaults/mediawiki". Can I proceed? [13:01:15] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1197 - https://phabricator.wikimedia.org/T438615#12349954 (10Jclark-ctr) [13:01:17] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1194 - https://phabricator.wikimedia.org/T438744#12349956 (10Jclark-ctr) →14Duplicate dup:03T438615 [13:02:44] 10ops-eqiad, 06SRE, 06DC-Ops, 10Prod-Kubernetes, and 3 others: wikikube-worker1152.eqiad.wmnet networking issue - https://phabricator.wikimedia.org/T438819#12349962 (10Jclark-ctr) Replaced cable no change in status for link. [13:03:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:05:42] 10ops-eqiad, 06SRE, 06DC-Ops, 10Prod-Kubernetes, and 3 others: wikikube-worker1152.eqiad.wmnet networking issue - https://phabricator.wikimedia.org/T438819#12349976 (10Jclark-ctr) updating a few firmwares on server [13:08:50] (03PS1) 10Dpogorzelski: ml-serve: repartition ml-serve1013 and 1014 to 16x96GB [puppet] - 10https://gerrit.wikimedia.org/r/1343981 (https://phabricator.wikimedia.org/T436928) [13:10:03] (03CR) 10Tiziano Fogli: [C:03+2] kafka-logging: add kafka-logging100[7-8] to eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [13:10:09] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12350013 (10Jclark-ctr) 05Open→03Resolved Received Replacement license from Dell and installed on server sessionstore1005 [13:10:44] (03PS1) 10Kosta Harlan: AbuseReview: Add warning indicating alpha test to vandalism queue [extensions/WikimediaAntiAbuse] (wmf/1.47.0-wmf.21) - 10https://gerrit.wikimedia.org/r/1343982 (https://phabricator.wikimedia.org/T438467) [13:11:24] (03CR) 10Ssingh: [C:03+1] sretest2013: remove manual references to this host [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) (owner: 10CDobbins) [13:12:19] 06SRE, 10observability, 06Traffic, 13Patch-For-Review: HAProxy metrics go down on config reload - https://phabricator.wikimedia.org/T343000#12350029 (10dkertesz) @Vgutierrez brought up an interesting point: after a crash of HAProxy, the new instance will load a stale stats file which will result in weird m... [13:19:55] (03CR) 10Atsuko: [C:03+1] "I run an approximate test for the whole stack on `dse-k8s-codfw`, https://phabricator.wikimedia.org/P96493. LGTM based on the diff, with a" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343123 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [13:20:25] I mean I can see if https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1342825 can be deployed early, I'm not busy right now [13:20:48] (03CR) 10Klausman: [C:03+1] ml-serve: repartition ml-serve1013 and 1014 to 16x96GB [puppet] - 10https://gerrit.wikimedia.org/r/1343981 (https://phabricator.wikimedia.org/T436928) (owner: 10Dpogorzelski) [13:21:18] !log tappof@deploy1003 helmfile [eqiad] START helmfile.d/admin 'sync'. [13:21:33] !log tappof@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'sync'. [13:22:46] !log tappof@deploy1003 helmfile [codfw] START helmfile.d/admin 'sync'. [13:23:22] !log tappof@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'sync'. [13:24:47] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for AHan-WMF - https://phabricator.wikimedia.org/T438538#12350070 (10AHan-WMF) @tappof Yes, I also need kerberos principal. [13:24:50] (03CR) 10Atsuko: [C:03+1] "LGTM based on the diff with a full stack, https://phabricator.wikimedia.org/P96492, with two questions" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343124 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [13:33:07] 10ops-eqiad, 06SRE, 06DC-Ops, 10Prod-Kubernetes, and 3 others: wikikube-worker1152.eqiad.wmnet networking issue - https://phabricator.wikimedia.org/T438819#12350122 (10Jclark-ctr) @Jelto It looks like the NICs on the motherboard have failed. We could possibly see if we have any PCI cards we can put into th... [13:34:22] FIRING: GnmiInterfaceCountersDrop: ... [13:34:22] asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=asw1-b3-magru:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:35:43] hello? I want to see if I can deploy a patch earlier... [13:37:24] (03PS3) 10Giuseppe Lavagetto: gVisor: disable on bookworm in containerd [puppet] - 10https://gerrit.wikimedia.org/r/1342961 (https://phabricator.wikimedia.org/T436649) [13:37:24] (03PS4) 10Giuseppe Lavagetto: gVisor: use kvm as a platform when hardware virtualization is available [puppet] - 10https://gerrit.wikimedia.org/r/1342962 (https://phabricator.wikimedia.org/T438287) [13:37:24] (03PS3) 10Giuseppe Lavagetto: kubernetes::node: add gvisor label based on hardware virtualization [puppet] - 10https://gerrit.wikimedia.org/r/1342963 (https://phabricator.wikimedia.org/T438287) [13:38:59] (03CR) 10Giuseppe Lavagetto: [V:03+1] "PCC SUCCESS (NOOP 1 CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/" [puppet] - 10https://gerrit.wikimedia.org/r/1342961 (https://phabricator.wikimedia.org/T436649) (owner: 10Giuseppe Lavagetto) [13:43:48] FIRING: PuppetFailure: Puppet has failed on netflow1003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:46:10] (03PS1) 10Tiziano Fogli: Revert "kafka-logging: add kafka-logging100[7-8] to eqiad cluster" [puppet] - 10https://gerrit.wikimedia.org/r/1343988 [13:47:10] (03CR) 10Tiziano Fogli: [C:03+2] Revert "kafka-logging: add kafka-logging100[7-8] to eqiad cluster" [puppet] - 10https://gerrit.wikimedia.org/r/1343988 (owner: 10Tiziano Fogli) [13:47:47] (03PS2) 10Cathal Mooney: Urldownloader: expand squid ACL to all production nets + analytics [puppet] - 10https://gerrit.wikimedia.org/r/1343970 (https://phabricator.wikimedia.org/T437560) [13:48:10] (03CR) 10Btullis: [C:03+2] "Acknowledged" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343123 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [13:48:48] FIRING: [2x] PuppetFailure: Puppet has failed on netflow1003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:50:44] PROBLEM - Host an-worker1194 is DOWN: PING CRITICAL - Packet loss = 100% [13:51:09] (03PS3) 10Daniel Kertesz: cache::haproxy: implement support for persistent stats [puppet] - 10https://gerrit.wikimedia.org/r/1343039 (https://phabricator.wikimedia.org/T343000) [13:51:41] !log elukey@cumin1004 START - Cookbook sre.hosts.powercycle for host pki1002 [13:51:58] (03CR) 10CI reject: [V:04-1] cache::haproxy: implement support for persistent stats [puppet] - 10https://gerrit.wikimedia.org/r/1343039 (https://phabricator.wikimedia.org/T343000) (owner: 10Daniel Kertesz) [13:52:32] (03CR) 10Btullis: [C:03+2] ceph-csi-cephfs: rebase the chart on upstream v3.14.2 (032 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343124 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [13:53:01] !log elukey@cumin1004 END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002 [13:53:48] FIRING: [4x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:53:56] (03CR) 10Dpogorzelski: [C:03+1] profiles/amd_gpu: Add GPUs ettings verification script (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341714 (https://phabricator.wikimedia.org/T431553) (owner: 10Klausman) [13:54:02] (03CR) 10CDobbins: [C:03+2] sretest2013: remove manual references to this host [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) (owner: 10CDobbins) [13:55:21] (03CR) 10Cathal Mooney: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343970 (https://phabricator.wikimedia.org/T437560) (owner: 10Cathal Mooney) [13:55:57] !log tappof@deploy1003 helmfile [eqiad] START helmfile.d/admin 'sync'. [13:56:13] !log tappof@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'sync'. [13:56:33] !log tappof@deploy1003 helmfile [codfw] START helmfile.d/admin 'sync'. [13:56:46] RECOVERY - Host pki1002 is UP: PING OK - Packet loss = 0%, RTA = 0.36 ms [13:57:07] !log tappof@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'sync'. [13:57:51] !log btullis@cumin1004 START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{dse-k8s-worker10[02-28].eqiad.wmnet} and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad) [13:57:55] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet [13:58:26] RESOLVED: [44x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:58:35] (03CR) 10Klausman: cookbooks/idm: Add user-cleanup cookbook for DPE SRE hosts (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1342235 (https://phabricator.wikimedia.org/T437615) (owner: 10Klausman) [13:58:48] FIRING: [5x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:58:51] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12350225 (10elukey) 05Resolved→03Open a:05elukey→03None Sadly it re-happened again, OEM events registered in `rac... [13:59:09] (03PS4) 10Daniel Kertesz: cache::haproxy: implement support for persistent stats [puppet] - 10https://gerrit.wikimedia.org/r/1343039 (https://phabricator.wikimedia.org/T343000) [13:59:20] (03Merged) 10jenkins-bot: ceph-csi-rbd: rebase the chart on upstream v3.14.2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343123 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [13:59:57] !log btullis@cumin1004 END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet [14:00:05] slyngs: #bothumor Q:How do functions break up? A:They stop calling each other. Rise for Southward Datacenter Switchover: Services + Traffic deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1400). [14:01:20] yess! go slyngs! [14:02:03] 🍿 [14:02:13] FIRING: [2x] JobUnavailable: Reduced availability for job cfssl in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:02:55] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [14:03:06] (03Merged) 10jenkins-bot: ceph-csi-cephfs: rebase the chart on upstream v3.14.2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343124 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [14:03:40] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1194 - https://phabricator.wikimedia.org/T438526#12350258 (10VRiley-WMF) 05Open→03Resolved Updated firmware and it cleared the issue closing the ticket [14:03:48] FIRING: [7x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:03:48] FIRING: PuppetFailure: Puppet has failed on cumin1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:05:36] (03PS5) 10Klausman: cookbooks/idm: Add user-cleanup cookbook for DPE SRE hosts [cookbooks] - 10https://gerrit.wikimedia.org/r/1342235 (https://phabricator.wikimedia.org/T437615) [14:06:35] (03CR) 10TChin: [C:03+1] article-feature-counts: add deployment chart files [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343044 (https://phabricator.wikimedia.org/T437000) (owner: 10AKhatun) [14:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:08:48] FIRING: [7x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:10:40] (03CR) 10Andrew Bogott: [C:03+2] Nova policy.yaml: disable instance pause, lock and suspend [puppet] - 10https://gerrit.wikimedia.org/r/1343587 (https://phabricator.wikimedia.org/T431307) (owner: 10Andrew Bogott) [14:12:30] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations, 06Release-Engineering-Team: Test the Docker Registry's GC with Releng/MediaWiki images - https://phabricator.wikimedia.org/T437297#12350288 (10MatthewVernon) [14:12:45] RESOLVED: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [14:12:51] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12350289 (10Jclark-ctr) This is cabled with 10g. @cmooney i missed that also. I would think it should work being setup as 10g. but both of the pci NICs are no... [14:13:48] FIRING: [7x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:13:48] RESOLVED: PuppetFailure: Puppet has failed on cumin1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:14:02] (03CR) 10Dpogorzelski: [C:03+2] ml-serve: repartition ml-serve1013 and 1014 to 16x96GB [puppet] - 10https://gerrit.wikimedia.org/r/1343981 (https://phabricator.wikimedia.org/T436928) (owner: 10Dpogorzelski) [14:14:21] (03CR) 10MVernon: [C:03+2] Pontoon: stack-specific changes for the swift stack [puppet] - 10https://gerrit.wikimedia.org/r/1305183 (https://phabricator.wikimedia.org/T429630) (owner: 10MVernon) [14:15:09] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet [14:15:10] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet [14:15:18] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet [14:17:20] !log btullis@cumin1004 END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet [14:17:43] (03CR) 10Clare Ming: [C:03+2] Test Kitchen UI: Deploy v2.0.0 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343912 (https://phabricator.wikimedia.org/T421814) (owner: 10Santiago Faci) [14:17:44] (03PS1) 10Ayounsi: depool-rack: add option to remove downtime on "pool" action [cookbooks] - 10https://gerrit.wikimedia.org/r/1343990 (https://phabricator.wikimedia.org/T327300) [14:18:32] RESOLVED: KubernetesCalicoDown: dse-k8s-worker1002.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-dse&var-instance=dse-k8s-worker1002.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [14:19:11] (03PS2) 10Slyngshede: service::catalog: exclude sessionstore from DC switchover [puppet] - 10https://gerrit.wikimedia.org/r/1343921 (https://phabricator.wikimedia.org/T435443) [14:19:41] (03CR) 10Jasmine: [C:03+1] service::catalog: exclude sessionstore from DC switchover [puppet] - 10https://gerrit.wikimedia.org/r/1343921 (https://phabricator.wikimedia.org/T435443) (owner: 10Slyngshede) [14:20:04] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v2.0.0 release to production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343912 (https://phabricator.wikimedia.org/T421814) (owner: 10Santiago Faci) [14:21:02] (03CR) 10Slyngshede: [C:03+2] service::catalog: exclude sessionstore from DC switchover [puppet] - 10https://gerrit.wikimedia.org/r/1343921 (https://phabricator.wikimedia.org/T435443) (owner: 10Slyngshede) [14:21:02] (03PS2) 10Klausman: profiles/amd_gpu: Add GPUs ettings verification script [puppet] - 10https://gerrit.wikimedia.org/r/1341714 (https://phabricator.wikimedia.org/T431553) [14:21:15] (03CR) 10Klausman: profiles/amd_gpu: Add GPUs ettings verification script (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1341714 (https://phabricator.wikimedia.org/T431553) (owner: 10Klausman) [14:23:45] 10ops-eqiad, 06DC-Ops: Firmware update on an-worker1[187-208] - https://phabricator.wikimedia.org/T438856 (10VRiley-WMF) 03NEW [14:23:48] FIRING: [4x] PuppetFailure: Puppet has failed on netflow1003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:24:38] 10ops-eqiad, 06DC-Ops: Firmware update on an-worker1[187-208] - https://phabricator.wikimedia.org/T438856#12350377 (10VRiley-WMF) a:03VRiley-WMF [14:24:47] 10ops-eqiad, 06DC-Ops: Firmware update on an-worker1[187-208] - https://phabricator.wikimedia.org/T438856#12350381 (10VRiley-WMF) p:05Triage→03Low [14:25:17] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12350388 (10cmooney) >>! In T438161#12350289, @Jclark-ctr wrote: > This is cabled with 10g. @cmooney i missed that also. I would think it should work being set... [14:25:19] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet [14:25:20] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet [14:25:25] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet [14:26:13] (03CR) 10Volans: depool-rack: add option to remove downtime on "pool" action (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1343990 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [14:28:24] FIRING: SystemdUnitFailed: amd-smi-gpu-partition.service on ml-serve1013:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:28:48] RESOLVED: [2x] PuppetFailure: Puppet has failed on netflow1003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:29:24] PROBLEM - Host an-worker1197 is DOWN: PING CRITICAL - Packet loss = 100% [14:29:29] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1197 - https://phabricator.wikimedia.org/T438615#12350417 (10VRiley-WMF) 05Open→03Resolved Updated firmware, and that corrected the issue. Sending drive back [14:30:14] (03CR) 10Ayounsi: depool-rack: add option to remove downtime on "pool" action (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1343990 (https://phabricator.wikimedia.org/T327300) (owner: 10Ayounsi) [14:31:31] (03PS1) 10Ayounsi: depool-rack: add policy "manual" for manual pool/depool [cookbooks] - 10https://gerrit.wikimedia.org/r/1343995 (https://phabricator.wikimedia.org/T327300) [14:32:45] (03PS3) 10Herron: grafana-ldap-users-sync: handle 412 duplicate email address [puppet] - 10https://gerrit.wikimedia.org/r/1343991 (https://phabricator.wikimedia.org/T438855) [14:34:09] (03PS1) 10Btullis: dse-k8s-eqiad: Add a second ipv4 pool to calico [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343998 (https://phabricator.wikimedia.org/T430658) [14:35:45] !log cdobbins@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [14:36:21] !log cdobbins@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [14:36:27] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12350493 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by cdobbins@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed... [14:40:20] (03PS1) 10Slyngshede: service::catalog: exclude k8s-ingress-aux [puppet] - 10https://gerrit.wikimedia.org/r/1344000 (https://phabricator.wikimedia.org/T435443) [14:42:33] !log cdobbins@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [14:42:58] !log cdobbins@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [14:43:07] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12350549 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by cdobbins@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed... [14:45:12] (03CR) 10Jasmine: [C:03+1] service::catalog: exclude k8s-ingress-aux [puppet] - 10https://gerrit.wikimedia.org/r/1344000 (https://phabricator.wikimedia.org/T435443) (owner: 10Slyngshede) [14:45:39] (03CR) 10Slyngshede: [C:03+2] service::catalog: exclude k8s-ingress-aux [puppet] - 10https://gerrit.wikimedia.org/r/1344000 (https://phabricator.wikimedia.org/T435443) (owner: 10Slyngshede) [14:47:10] (03CR) 10Vgutierrez: "I like this approach more cause we don't mess with the current behavior set by `KillMode=mixed` by adding `ExecStop=` stanzas" [puppet] - 10https://gerrit.wikimedia.org/r/1343039 (https://phabricator.wikimedia.org/T343000) (owner: 10Daniel Kertesz) [14:47:16] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host cirrussearch1120 [14:47:16] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120 [14:51:41] !log bking@cumin2003 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN T436571 [14:51:44] T436571: Re-IP Search Platform-owned EQIAD hosts to new per-rack vlans/subnets - https://phabricator.wikimedia.org/T436571 [14:53:58] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host cirrussearch1120 [14:54:07] !log slyngshede@cumin1004 START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified] [14:54:13] !log slyngshede@cumin1004 END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified] [14:54:26] !log dancy@deploy1003 Installing scap version "4.290.0" for 155 host(s) [14:54:56] We're now ready to depool eqiad. Hang on everyone! [14:55:07] \o/ [14:55:09] yayay [14:55:22] \i/ [14:55:29] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet [14:55:30] !log slyngshede@cumin1004 START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - T435443 [14:55:34] T435443: Day 1 - Source datacenter depooling (Tue. Sept. 22nd - 2026) - https://phabricator.wikimedia.org/T435443 [14:56:36] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12350717 (10elukey) This is what I see from `/redfish/v1/Systems/1/Oem/Supermicro/FixedBootOrder`: ` 'UEFINetwork': ['(B81/D0/F0) UEFI HTTP IPv4 Intel(R) Ethernet Contro... [14:59:28] !log dancy@deploy1003 Installing scap version "4.290.0" for 3 host(s) [15:00:05] slyngs: Time to do the Southward Datacenter Switchover: Services + Traffic deploy. Don't look at me like that. You signed up for it. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1400). [15:00:05] jelto, arnoldokoth, mutante, and arnaudb: I, the Bot under the Fountain, call upon thee, The Deployer, to do SRE Collaboration Services office hours deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1500). [15:01:17] !log dancy@deploy1003 Installation of scap version "4.290.0" completed for 3 hosts [15:03:39] (03PS1) 10Elukey: profile::pki::multirootca: disable mTLS requirements [puppet] - 10https://gerrit.wikimedia.org/r/1344004 (https://phabricator.wikimedia.org/T436809) [15:03:39] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet [15:03:40] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet [15:03:41] (03PS1) 10Elukey: profile::pki::client: disable client auth via mTLS [puppet] - 10https://gerrit.wikimedia.org/r/1344005 (https://phabricator.wikimedia.org/T436809) [15:03:46] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet [15:03:53] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344004 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [15:04:02] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344005 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [15:04:17] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet [15:04:40] 10ops-codfw, 06DC-Ops: Alert for device ps1-a2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438865 (10phaultfinder) 03NEW [15:09:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 7.529% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:09:18] 06SRE, 06Infrastructure-Foundations: Ignore Valid-Until for WMF container images - https://phabricator.wikimedia.org/T438866 (10MoritzMuehlenhoff) 03NEW [15:11:15] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet [15:11:16] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet [15:11:21] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet [15:11:52] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet [15:14:15] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at codfw: 24.31% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:14:20] !log bking@cumin2003 END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120 [15:14:20] FIRING: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [15:14:24] 06SRE, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Site: EQIAD VM request for Kerberos - https://phabricator.wikimedia.org/T438229#12350803 (10MoritzMuehlenhoff) Why another VM? krb1002 was bought for this and is under warranty for another 1.5 years, so we can expect Supermic... [15:14:37] 10ops-codfw, 06DC-Ops: Alert for device ps1-c4-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438868 (10phaultfinder) 03NEW [15:14:38] 10ops-codfw, 06DC-Ops: Alert for device ps1-b2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438869 (10phaultfinder) 03NEW [15:15:10] PROBLEM - Host ml-serve1013 is DOWN: PING CRITICAL - Packet loss = 100% [15:15:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: codfw mw-web releases routed via main (k8s) 827.4ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [15:16:27] !log bking@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120 [15:16:40] !log bking@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120 [15:18:24] RESOLVED: SystemdUnitFailed: amd-smi-gpu-partition.service on ml-serve1013:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:18:56] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet [15:18:57] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet [15:19:03] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet [15:19:15] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at codfw: 22.9% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:22:10] (03CR) 10BCornwall: [C:03+1] add discovery record for codesearch [dns] - 10https://gerrit.wikimedia.org/r/1342805 (https://phabricator.wikimedia.org/T268199) (owner: 10Dzahn) [15:22:19] !log slyngshede@cumin1004 END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - T435443 [15:22:25] T435443: Day 1 - Source datacenter depooling (Tue. Sept. 22nd - 2026) - https://phabricator.wikimedia.org/T435443 [15:23:32] FIRING: KubernetesCalicoDown: ml-serve1013.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1013.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [15:23:40] RECOVERY - Host ml-serve1013 is UP: PING OK - Packet loss = 0%, RTA = 0.49 ms [15:23:49] (03PS2) 10Elukey: profile::pki::client: disable client auth via mTLS [puppet] - 10https://gerrit.wikimedia.org/r/1344005 (https://phabricator.wikimedia.org/T436809) [15:24:01] (03CR) 10Elukey: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344005 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [15:24:15] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at codfw: 22.9% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:24:36] 10ops-codfw, 06DC-Ops: Alert for device ps1-c5-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438872 (10phaultfinder) 03NEW [15:27:16] !log ayounsi@cumin1004 START - Cookbook sre.dns.netbox [15:28:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 4.153% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:28:32] RESOLVED: KubernetesCalicoDown: ml-serve1013.eqiad.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=eqiad%20prometheus%2Fk8s-mlserve&var-instance=ml-serve1013.eqiad.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [15:28:33] !ack [15:28:34] 8357 (ACKED) PHPFPMTooBusy sre (mw-web main codfw) [15:28:36] (03PS1) 10Dpogorzelski: Revert "ml-serve: repartition ml-serve1013 and 1014 to 16x96GB" [puppet] - 10https://gerrit.wikimedia.org/r/1344010 [15:29:01] <_joe_> uh codfw [15:29:07] <_joe_> what is even going on here [15:29:12] (03CR) 10Dpogorzelski: [C:03+2] Revert "ml-serve: repartition ml-serve1013 and 1014 to 16x96GB" [puppet] - 10https://gerrit.wikimedia.org/r/1344010 (owner: 10Dpogorzelski) [15:29:15] (03CR) 10Dpogorzelski: [V:03+2 C:03+2] Revert "ml-serve: repartition ml-serve1013 and 1014 to 16x96GB" [puppet] - 10https://gerrit.wikimedia.org/r/1344010 (owner: 10Dpogorzelski) [15:29:22] might that have been the depool? [15:29:29] went up 25 mins ago [15:29:44] <_joe_> bjensen: we haven't depooled mediawiki in eqiad, right? [15:30:00] no, not to my knowledge [15:30:08] just services today [15:30:14] <_joe_> uh no actually we did [15:30:17] <_joe_> sigh [15:30:24] ah [15:30:58] <_joe_> ok [15:31:07] do we need to repool mw-web or do we need to bump codfw replicas or both? [15:31:23] <_joe_> the codfw stuff seems worse than replicas [15:31:24] 10SRE-SLO, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th): page_change SLO windows - https://phabricator.wikimedia.org/T438054#12350943 (10elukey) @APizzata-WMF I think that Sloth's base alert will complain that no metrics is present during that time, so the ideal state would be those daily metrics... [15:31:33] <_joe_> it seems to be the databases not keeping up with the work [15:31:41] Need help? [15:31:51] I'm afk but I can be in my laptop in a few [15:32:18] <_joe_> so yeah we can't sustain the current read-only traffic in one DC I would say [15:32:39] Eqiad is depooled for DBs? [15:32:41] ayounsi@cumin1004 netbox (PID 224393) is awaiting input [15:32:53] <_joe_> marostegui: is depooled for read mw traffic atm [15:33:02] Right [15:33:03] what's with the SSL connection rate? https://grafana.wikimedia.org/goto/swcdst?orgId=default [15:33:29] <_joe_> marostegui: checking how hot the dbs in codfw have been running would be useful [15:33:37] <_joe_> for now I am going to repool eqiad for mw-web [15:33:42] Ok I'll get to my desk [15:34:15] FIRING: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at codfw: 21.57% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:34:17] the query throughput and looks consistent [15:35:02] !log oblivian@cumin1004 START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance [15:35:02] !log oblivian@cumin1004 END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance [15:35:10] We've doubled their traffic in codfw [15:35:12] ofc [15:35:17] s2 is spiking a bit but nothing too strange https://grafana.wikimedia.org/goto/sbbrjj?orgId=default [15:35:34] !log oblivian@cumin1004 START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance [15:35:35] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12350954 (10MoritzMuehlenhoff) [15:36:01] <_joe_> ok so this is a showstopper for the dc switchover [15:36:05] <_joe_> as I sadly expected [15:36:13] <_joe_> slyngs: the switchover is on hold from now on [15:36:21] !log ayounsi@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004" [15:36:24] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004" [15:36:24] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [15:36:25] Got it [15:36:35] So far they are coping, but of course we've doubled our replicas traffic [15:36:46] <_joe_> marostegui: we're not really coping atm it seems [15:37:03] _joe_: On the DBs I still don't see anything going super flat [15:37:13] But yeah, the traffic keeps increasing [15:37:31] php-fpm workers saturation is going down in codfw [15:37:37] But up in eqiad [15:37:54] <_joe_> yeah I have repooled eqiad for mw-web-ro [15:38:22] 6 months ago we didn't have this issue, did we? I don't recall it [15:38:28] <_joe_> we didn't [15:38:40] <_joe_> I have been saying for a week we were running too hot on mw [15:38:46] yeah [15:38:53] because we've not reduced our DB capacity [15:38:58] <_joe_> and I think the reason is thumbs.wikimedia.org [15:39:02] we have the same amount of dbs [15:39:15] RESOLVED: [2x] PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at codfw: 23.4% idle - https://bit.ly/wmf-fpmsat - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:39:32] <_joe_> ok so moving mw-web also resolved the issue for mw-api-ext [15:39:48] <_joe_> which kinda proves it's a shared resource issue [15:39:55] yeah, I am seeing our db traffic now increasing in eqiad [15:39:59] <_joe_> I am not sure it's memcached or databases or what [15:40:03] <_joe_> yeah that is expected [15:40:14] yep [15:40:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: codfw mw-web releases routed via main (k8s) 801.6ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [15:40:16] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations, 06Release-Engineering-Team: Test the Docker Registry's GC with Releng/MediaWiki images - https://phabricator.wikimedia.org/T437297#12350968 (10elukey) @dancy hi! I am close to complete the first cleanup of the old MediaWiki manifests, and now I'd l... [15:40:21] <_joe_> so, we need to work hard between now and tomorrow to determine what happened here [15:40:23] Are y'all completely rolling back the switchover? Just making sure as Cirrussearch is unhappy ATM too [15:40:26] <_joe_> or we can't do the switchover [15:40:37] <_joe_> inflatador: not completely, no [15:40:38] !log oblivian@cumin1004 END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance [15:40:38] do we know why the number of connection is so different? https://grafana.wikimedia.org/goto/sb5cxd?orgId=default [15:40:58] <_joe_> I am at the moment only switching back mw-web-ro and mw-api-ext-ro [15:41:02] <_joe_> in a few [15:41:08] s7 rows read are crazy in codfw and have been crazy for a while [15:41:32] so s7 has centralauth [15:41:44] and some other big wikis, but centralauth lives there [15:41:45] _joe_ got it. I will discuss with Search Platform then, we have some knobs we can turn in mwconfig [15:42:08] <_joe_> nto sure why not all the reads are back in eqiad atm for mw-web [15:42:20] (03CR) 10BCornwall: [C:03+1] varnish: Fix thumb.wm.o normalization to match upload.wm.o [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [15:42:52] (03CR) 10BCornwall: [V:03+1 C:03+1] "Dunno what happened but it's working now. Not gonna question it, no brainpower to spare." [puppet] - 10https://gerrit.wikimedia.org/r/1341936 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [15:43:11] This reminds me a bit of when we had one DC cold and replicas weren't too warmed up, but ofc this is not the case now as they are as warm as eqiad's [15:43:34] marostegui: _joe_ wtf is this https://grafana-rw.wikimedia.org/goto/stwf8v?orgId=default [15:44:20] :-/ [15:44:25] what could be reading ONLY from codfw??? [15:45:00] <_joe_> marostegui: I think none of the DCs is able to sustain our current level of uncached traffic to mw-web [15:45:28] (03CR) 10Dpogorzelski: [C:03+1] profiles/amd_gpu: Add GPUs ettings verification script [puppet] - 10https://gerrit.wikimedia.org/r/1341714 (https://phabricator.wikimedia.org/T431553) (owner: 10Klausman) [15:45:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at codfw: 4.194% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=codfw%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:46:01] (03CR) 10Klausman: [C:03+2] profiles/amd_gpu: Add GPUs ettings verification script [puppet] - 10https://gerrit.wikimedia.org/r/1341714 (https://phabricator.wikimedia.org/T431553) (owner: 10Klausman) [15:46:26] cdanis: any chance you can split that graph by host? [15:46:47] I want to make sure we are not graphing non production hosts there which could mess up with the graph, eg: backup host taking backups == rows read [15:47:03] (03PS3) 10Sbisson: Keep Article Guidance on where it is on today [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342285 (https://phabricator.wikimedia.org/T433293) [15:47:21] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 22 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342285 (https://phabricator.wikimedia.org/T433293) (owner: 10Sbisson) [15:48:40] 10ops-codfw, 10ops-eqiad, 06DC-Ops: Audit Serial Console Connections (2026) - https://phabricator.wikimedia.org/T438876 (10cmooney) 03NEW p:05Triage→03Low [15:48:54] marostegui: ah okay so that load is all db2218 [15:48:58] https://grafana-rw.wikimedia.org/goto/sccgbh?orgId=default [15:49:05] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet [15:49:08] cdanis: right, so that is production [15:49:17] so it is real [15:50:33] !log installing libhtml-parser-perl security updates [15:50:37] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:50:41] the db2218 load is mostly coming from mw-web and mw-api-ext in codfw [15:50:45] weird it's hitting mostly that replica [15:50:52] cdanis: this is interesting ,compare this two replicas in s7: https://grafana.wikimedia.org/d/000000273/mysql?var-server=db2218&var-port=9104&from=now-12h&to=now&timezone=utc&var-job=$__all&refresh=1m&viewPanel=panel-3 https://grafana.wikimedia.org/d/000000273/mysql?var-server=db2222&var-port=9104&from=now-12h&to=now&timezone=utc&var-job=$__all&refresh=1m&viewPanel=panel-3 [15:51:14] maybe the optimizer going crazy on that host and doing silly things? let me check if I can find some queries there [15:51:34] PROBLEM - Host cirrussearch1120 is DOWN: PING CRITICAL - Packet loss = 100% [15:51:34] (03CR) 10JHathaway: [C:03+1] profile::pki::client: disable client auth via mTLS [puppet] - 10https://gerrit.wikimedia.org/r/1344005 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [15:51:40] (03CR) 10JHathaway: [C:03+1] profile::pki::multirootca: disable mTLS requirements [puppet] - 10https://gerrit.wikimedia.org/r/1344004 (https://phabricator.wikimedia.org/T436809) (owner: 10Elukey) [15:52:48] (03PS1) 10Muehlenhoff: Record LDAP access for nbattista [puppet] - 10https://gerrit.wikimedia.org/r/1344013 [15:54:15] (03CR) 10Muehlenhoff: [C:03+2] Record LDAP access for nbattista [puppet] - 10https://gerrit.wikimedia.org/r/1344013 (owner: 10Muehlenhoff) [15:54:22] cdanis: I think I am seeing something with the optimizer [15:54:32] marostegui: should we just depool that host? [15:54:35] or lower its weight maybe [15:54:38] give me a sec [15:55:50] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet [15:55:51] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet [15:55:57] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet [15:56:43] yep, the optimizer is being silly with globaluser table on that host [15:56:53] let me depool refresh the stats and check again [15:57:06] !log marostegui@cumin1004 START - Cookbook sre.mysql.depool depool db2218: optimizer issues [15:58:12] !log marostegui@cumin1004 END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues [15:58:49] (03PS1) 10Ebernhardson: cirrus: Send more_like traffic to eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 [15:59:08] 10ops-codfw, 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Enable Performance Governor on Cirrussearch hosts - https://phabricator.wikimedia.org/T438877#12351042 (10bking) 05Open→03In progress p:05Medium→03High [15:59:12] cdanis: host depooled, I am refreshing table stats and will manually run an explain to see what the optimizer wants to do [15:59:21] (03PS2) 10Ebernhardson: cirrus: Send more_like traffic to eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 [16:00:05] jhathaway and rzl: #bothumor My software never has bugs. It just develops random features. Rise for Puppet request window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1600). [16:00:05] zabe: A patch you scheduled for Puppet request window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [16:00:46] (03PS1) 10Slyngshede: mw-web: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344016 (https://phabricator.wikimedia.org/T435462) [16:02:04] 10ops-codfw, 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Enable Performance Governor on Cirrussearch hosts - https://phabricator.wikimedia.org/T438877#12351053 (10bking) Hello DC Ops, We have 110 Cirrussearch hosts, with 55 each in EQIAD/CODFW. [[ https://fault-tolerance.toolfor... [16:02:08] (03PS1) 10Muehlenhoff: Record LDAP access for rmota [puppet] - 10https://gerrit.wikimedia.org/r/1344017 [16:02:19] 10ops-codfw, 10ops-eqiad, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Enable Performance Governor on Cirrussearch hosts - https://phabricator.wikimedia.org/T438877#12351055 (10bking) a:05bking→03None [16:03:46] (03CR) 10Muehlenhoff: [C:03+2] Record LDAP access for rmota [puppet] - 10https://gerrit.wikimedia.org/r/1344017 (owner: 10Muehlenhoff) [16:04:20] !log sukhe@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [16:04:45] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [16:04:59] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12351075 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by sukhe@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed ho... [16:05:02] (03PS1) 10Ayounsi: move-vlan: handle 'systemctl restart networking' better [cookbooks] - 10https://gerrit.wikimedia.org/r/1344018 [16:06:50] (03CR) 10Jasmine: [C:03+1] mw-web: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344016 (https://phabricator.wikimedia.org/T435462) (owner: 10Slyngshede) [16:09:13] (03PS1) 10Dreamy Jazz: Enable AbuseReview on jawiki for likely PII [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344020 (https://phabricator.wikimedia.org/T438867) [16:12:13] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:13:38] FIRING: CirrusSearchNodeIndexingNotIncreasing: OpenSearch instance cirrussearch1120-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [16:13:45] !log oblivian@cumin1004 START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance [16:15:12] (03PS1) 10CDobbins: cp2059: add this host as a text node [puppet] - 10https://gerrit.wikimedia.org/r/1344022 (https://phabricator.wikimedia.org/T436691) [16:15:37] (03CR) 10Ssingh: [C:03+1] cp2059: add this host as a text node [puppet] - 10https://gerrit.wikimedia.org/r/1344022 (https://phabricator.wikimedia.org/T436691) (owner: 10CDobbins) [16:17:13] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:18:01] (03CR) 10CDobbins: [C:03+2] cp2059: add this host as a text node [puppet] - 10https://gerrit.wikimedia.org/r/1344022 (https://phabricator.wikimedia.org/T436691) (owner: 10CDobbins) [16:18:22] Is the puppet window being used? [16:19:02] !log oblivian@cumin1004 END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance [16:19:06] I see a patch in it, but nothing seems to have happened yet in the window [16:19:09] (03CR) 10BBlack: [C:04-1] "-1 for now, but this is more about the intent than the technical merits of the patch. We should continue discussion in T425216 ." [puppet] - 10https://gerrit.wikimedia.org/r/1341937 (https://phabricator.wikimedia.org/T425216) (owner: 10Krinkle) [16:20:05] RESOLVED: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [16:21:35] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12351163 (10MoritzMuehlenhoff) [16:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [16:25:59] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet [16:26:15] !log cdobbins@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [16:26:41] !log cdobbins@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [16:28:18] 10ops-codfw, 06SRE, 06DC-Ops, 13Patch-For-Review: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12351230 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by cdobbins@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WA... [16:29:38] Dreamy_Jazz: zabe had a patch, but we were holding off, due to dc switch over woes [16:30:02] Do these woes prevent me from using scap? [16:30:22] (Yet to read backscroll) [16:30:27] 06SRE, 06Commons, 10MediaWiki-File-management, 06Traffic, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12351266 (10BBlack) >>! In T425216#12278778, @Krinkle wrote: >>>! In T435283#12248179, @BBlack wrote: >>... [16:32:02] good question Dreamy_Jazz [16:32:04] Dreamy_Jazz: yes, hold off for now pelase [16:32:06] *please [16:32:14] thanks rzl [16:32:43] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 22 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344020 (https://phabricator.wikimedia.org/T438867) (owner: 10Dreamy Jazz) [16:33:07] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet [16:33:08] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet [16:33:13] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet [16:33:15] I'll wait for the backport window then :D [16:33:20] Thanks [16:34:07] Dreamy_Jazz: we'll hopefully be in a good place by then but please check again before deploying :) [16:34:12] Sure [16:36:10] 10ops-codfw, 10ops-eqiad, 06SRE, 06DC-Ops: Audit Serial Console Connections (2026) - https://phabricator.wikimedia.org/T438876#12351323 (10cmooney) @Jclark-ctr @VRiley-WMF, going through the ports on scs-a8-eqiad these were the issues I found. Port 8 - should be connected to ps1-a8-eqiad, but there is not... [16:41:03] 10ops-codfw, 10ops-eqiad, 06SRE, 06DC-Ops: Audit Serial Console Connections (2026) - https://phabricator.wikimedia.org/T438876#12351371 (10cmooney) [16:41:42] !log cdobbins@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [16:42:06] !log cdobbins@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [16:42:12] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12351376 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by cdobbins@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed... [16:42:43] !log marostegui@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing [16:43:51] Dreamy_Jazz: I stand corrected, go ahead [16:44:00] Thanks, proceeding [16:44:30] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344020 (https://phabricator.wikimedia.org/T438867) (owner: 10Dreamy Jazz) [16:45:26] (03Merged) 10jenkins-bot: Enable AbuseReview on jawiki for likely PII [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344020 (https://phabricator.wikimedia.org/T438867) (owner: 10Dreamy Jazz) [16:45:55] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1344020|Enable AbuseReview on jawiki for likely PII (T438867)]] [16:46:00] T438867: AbuseReview: Enable personal info content policy on jawiki - https://phabricator.wikimedia.org/T438867 [16:46:29] !log marostegui@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing [16:48:51] !log oblivian@cumin1004 START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance [16:50:20] !log dreamyjazz@deploy1003 dreamyjazz: Backport for [[gerrit:1344020|Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [16:50:27] Checking.... [16:51:50] !log dreamyjazz@deploy1003 dreamyjazz: Continuing with deployment [16:54:08] !log oblivian@cumin1004 END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance [16:54:41] FIRING: ConfdResourceFailed: confd resource _var_lib_gdnsd_discovery-rest-gateway.state.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [16:59:08] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1344020|Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s) [16:59:14] T438867: AbuseReview: Enable personal info content policy on jawiki - https://phabricator.wikimedia.org/T438867 [16:59:22] RESOLVED: GnmiInterfaceCountersDrop: ... [16:59:22] asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=asw1-b3-magru:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [16:59:41] FIRING: [16x] ConfdResourceFailed: confd resource _var_lib_gdnsd_discovery-rest-gateway.state.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [17:00:04] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1700) [17:00:54] !log sukhe@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [17:01:19] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [17:01:25] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12351494 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by sukhe@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed ho... [17:01:41] !log sukhe@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet [17:02:05] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet [17:02:11] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12351503 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by sukhe@cumin1004 from sretest2013 to cp2059.codfw.wmnet completed: - sretest2013 (**WARN**) - ✔️... [17:02:20] FIRING: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [17:03:18] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet [17:04:45] After deploying scap I'm seeing a lot of 404 errors for calling out to an internal service [17:06:05] 10ops-codfw, 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Enable Performance Governor on Cirrussearch hosts - https://phabricator.wikimedia.org/T438877#12351518 (10Jclark-ctr) There would be a good chance we could have issues in Rows E and F for power, and we would probabl... [17:06:06] !log marostegui@cumin1004 START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed [17:07:47] The logs seem to have started before my deploy, to appears related to the switchover [17:09:33] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet [17:09:34] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet [17:09:39] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet [17:10:14] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet [17:13:17] (03PS1) 10Btullis: dse-k8s: allow unauthenticated reads of the OIDC discovery endpoints [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344034 (https://phabricator.wikimedia.org/T435591) [17:15:22] !log oblivian@puppetserver1001 conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw [17:16:57] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet [17:16:58] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet [17:17:03] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet [17:17:40] (03PS2) 10Btullis: dse-k8s: allow unauthenticated reads of the OIDC discovery endpoints [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344034 (https://phabricator.wikimedia.org/T435591) [17:17:41] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet [17:19:20] 10ops-codfw, 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Enable Performance Governor on Cirrussearch hosts - https://phabricator.wikimedia.org/T438877#12351593 (10Jhancock.wm) We've already got power alerts in A2, B2, C4, and C5 for codfw since the start of the switchover... [17:19:41] RESOLVED: [16x] ConfdResourceFailed: confd resource _var_lib_gdnsd_discovery-rest-gateway.state.toml has errors - https://wikitech.wikimedia.org/wiki/Confd#Monitoring - https://grafana.wikimedia.org/d/OUJF1VI4k/confd - https://alerts.wikimedia.org/?q=alertname%3DConfdResourceFailed [17:22:20] RESOLVED: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [17:22:37] !log dzahn@dns1004 START - running authdns-update [17:24:10] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet [17:24:11] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet [17:24:17] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet [17:24:36] 10SRE-SLO, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th): page_change SLO windows - https://phabricator.wikimedia.org/T438054#12351615 (10APizzata-WMF) Hey @elukey, thanks for the reply. Unfortunately we can compute these metrics only once a month, but I get your point. Please let me know if anyth... [17:24:59] !log dzahn@dns1004 END - running authdns-update [17:25:00] !log sukhe@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [17:25:25] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [17:25:30] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12351618 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by sukhe@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed ho... [17:26:20] !log btullis@cumin1004 END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet [17:26:20] FIRING: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [17:29:58] 10ops-codfw, 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Enable Performance Governor on Cirrussearch hosts - https://phabricator.wikimedia.org/T438877#12351647 (10Jhancock.wm) safe racks in codfw: A4 B4 B5 C1 C2 D1 D3 D4 maybe racks, might cause a power alert, but try... [17:31:28] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host wdqs1029 [17:31:46] !log bking@cumin2003 START - Cookbook sre.dns.netbox [17:32:00] (03PS1) 10Dzahn: planet: add a feed [puppet] - 10https://gerrit.wikimedia.org/r/1344039 [17:32:51] !log sukhe@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [17:33:15] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [17:33:22] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12351671 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by sukhe@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed ho... [17:34:42] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-b2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438869#12351675 (10phaultfinder) [17:36:50] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003" [17:36:54] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003" [17:36:54] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:36:54] !log bking@cumin2003 START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [17:36:57] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [17:36:58] (03PS1) 10Ilias Sarantopoulos: fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344040 [17:36:58] !log bking@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029 [17:37:26] !log bking@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029 [17:37:26] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029 [17:38:47] (03CR) 10Bking: [C:03+1] "Confirmed working via `test-cookbook -c 1344018 sre.hosts.move-vlan inplace wdqs1029`" [cookbooks] - 10https://gerrit.wikimedia.org/r/1344018 (owner: 10Ayounsi) [17:39:07] 06SRE, 10Wikimedia-Mailing-lists: Change owner of Wikiversity-l mailing list - https://phabricator.wikimedia.org/T438820#12351702 (10Dzahn) If you are currently the owner and have access; you should be able to just add a new admin and once they confirm they can login remove yourself, no? [17:39:32] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet [17:39:33] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet [17:39:39] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet [17:41:40] (03CR) 10Kosta Harlan: [C:03+1] fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344040 (owner: 10Ilias Sarantopoulos) [17:42:34] (03PS2) 10Ilias Sarantopoulos: fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344040 [17:42:42] jouncebot: nowandnext [17:42:42] For the next 0 hour(s) and 17 minute(s): MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1700) [17:42:42] In 0 hour(s) and 17 minute(s): MediaWiki train - Utc-7 Version (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1800) [17:43:27] Going to use scap to fix those errors I mentioned about above [17:43:51] !log bking@cumin2003 START - Cookbook sre.hosts.move-vlan for host wdqs1030 [17:44:13] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344040 (owner: 10Ilias Sarantopoulos) [17:45:12] (03Merged) 10jenkins-bot: fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344040 (owner: 10Ilias Sarantopoulos) [17:45:24] !log bking@cumin2003 START - Cookbook sre.dns.netbox [17:45:34] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1344040|fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] [17:46:58] !log sukhe@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [17:47:22] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [17:47:32] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12351732 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by sukhe@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed ho... [17:49:39] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-b2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438869#12351747 (10phaultfinder) [17:50:01] !log dreamyjazz@deploy1003 dreamyjazz, isaranto: Backport for [[gerrit:1344040|fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [17:50:12] (03PS1) 10Bking: Cirrussearch: Enable performance governor on Row D hosts [puppet] - 10https://gerrit.wikimedia.org/r/1344041 (https://phabricator.wikimedia.org/T438877) [17:50:35] !log dreamyjazz@deploy1003 dreamyjazz, isaranto: Continuing with deployment [17:50:46] bking@cumin2003 move-vlan (PID 33359) is awaiting input [17:51:18] !log marostegui@cumin1004 END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed [17:53:25] !log bking@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003" [17:53:29] !log bking@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003" [17:53:29] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [17:53:30] !log bking@cumin2003 START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [17:53:33] !log bking@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors [17:53:34] !log bking@cumin2003 START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030 [17:53:51] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344041 (https://phabricator.wikimedia.org/T438877) (owner: 10Bking) [17:54:29] !log bking@cumin2003 END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030 [17:54:29] !log bking@cumin2003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030 [17:55:43] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1344040|fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s) [17:55:56] (03PS1) 10Ebernhardson: opensearch semantic: Update to opensearch 3.8.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344042 (https://phabricator.wikimedia.org/T438058) [17:56:33] (03PS2) 10Bking: Cirrussearch: Enable performance governor on Row D hosts [puppet] - 10https://gerrit.wikimedia.org/r/1344041 (https://phabricator.wikimedia.org/T438877) [18:00:05] thcipriani and thcipriani: #bothumor I � Unicode. All rise for MediaWiki train - Utc-7 Version deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T1800). [18:00:26] darnit. /me updates calendar. [18:00:34] (03CR) 10Ebernhardson: [C:03+2] opensearch semantic: Update to opensearch 3.8.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344042 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [18:03:14] (03Merged) 10jenkins-bot: opensearch semantic: Update to opensearch 3.8.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344042 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [18:03:34] (03CR) 10Ebernhardson: [C:03+1] "seems reasonable from the application side. I haven't verified the host->rack mapping." [puppet] - 10https://gerrit.wikimedia.org/r/1344041 (https://phabricator.wikimedia.org/T438877) (owner: 10Bking) [18:04:38] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-b2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438869#12351836 (10phaultfinder) [18:06:14] (03CR) 10Btullis: [C:03+1] Cirrussearch: Enable performance governor on Row D hosts [puppet] - 10https://gerrit.wikimedia.org/r/1344041 (https://phabricator.wikimedia.org/T438877) (owner: 10Bking) [18:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:07:06] !log Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0 [18:07:08] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:09:13] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for dpislaru - https://phabricator.wikimedia.org/T438827#12351849 (10Dzahn) a:03thcipriani [18:09:44] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet [18:14:56] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for dpislaru - https://phabricator.wikimedia.org/T438827#12351873 (10Seddon) Approved from myside! [18:16:04] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet [18:16:05] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet [18:16:11] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet [18:17:10] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12351878 (10Jclark-ctr) @VRiley-WMF You did decom a server with matching 450 config A might be able to swap failed parts... [18:18:14] !log btullis@cumin1004 END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet [18:24:45] (03PS1) 10Andrea Denisse: kubernetes: Register the aux Kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1343838 (https://phabricator.wikimedia.org/T438801) [18:25:46] (03PS1) 10TrainBranchBot: group0 to 1.47.0-wmf.21 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344046 (https://phabricator.wikimedia.org/T438217) [18:25:49] (03CR) 10TrainBranchBot: [C:03+2] "Initiated by jhuneidi@deploy1003" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344046 (https://phabricator.wikimedia.org/T438217) (owner: 10TrainBranchBot) [18:26:44] (03Merged) 10jenkins-bot: group0 to 1.47.0-wmf.21 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344046 (https://phabricator.wikimedia.org/T438217) (owner: 10TrainBranchBot) [18:34:52] (03PS1) 10Eric Gardner: ReaderExperiments: Enable on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344048 [18:34:52] (03PS1) 10Eric Gardner: ReaderExperiments: Set the preferred-sources debug flag on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344049 (https://phabricator.wikimedia.org/T436692) [18:34:54] (03PS1) 10Eric Gardner: ReaderExperiments: Drop the stale ShareHighlight config var [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344050 (https://phabricator.wikimedia.org/T424764) [18:35:30] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet [18:35:31] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet [18:35:37] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet [18:36:02] (03CR) 10Bking: [C:03+2] Cirrussearch: Enable performance governor on Row D hosts [puppet] - 10https://gerrit.wikimedia.org/r/1344041 (https://phabricator.wikimedia.org/T438877) (owner: 10Bking) [18:36:18] 06SRE, 10SRE-Access-Requests: Requesting access to deployment for EAlbizzati-WMF - https://phabricator.wikimedia.org/T437741#12351927 (10Dzahn) 05Open→03Stalled [18:36:21] !log jhuneidi@deploy1003 rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs T438217 [18:36:25] T438217: 1.47.0-wmf.21 deployment blockers - https://phabricator.wikimedia.org/T438217 [18:38:47] (03PS2) 10CDanis: kubernetes: Register the aux Kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1343838 (https://phabricator.wikimedia.org/T438801) (owner: 10Andrea Denisse) [18:38:48] (03CR) 10CDanis: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343838 (https://phabricator.wikimedia.org/T438801) (owner: 10Andrea Denisse) [18:40:43] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet [18:41:00] !log dancy@deploy1003 Installing scap version "4.291.0" for 3 host(s) [18:41:48] !log dancy@deploy1003 install-world aborted: (no justification provided) (duration: 00m 48s) [18:43:12] !log dancy@deploy1003 Installing scap version "4.291.0" for 3 host(s) [18:44:38] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-b2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438869#12351959 (10phaultfinder) [18:44:59] !log dancy@deploy1003 Installing scap version "4.291.0" for 3 host(s) [18:45:09] (03CR) 10CDanis: [C:03+1] "thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1343838 (https://phabricator.wikimedia.org/T438801) (owner: 10Andrea Denisse) [18:46:24] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:47:10] !log dancy@deploy1003 Installing scap version "4.291.0" for 3 host(s) [18:48:23] mutante: Are you around? I need some assistance recovering scap on the deploy server [18:49:04] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet [18:49:05] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet [18:49:10] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet [18:49:45] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet [18:50:41] 06SRE, 06Infrastructure-Foundations: Create nodejs 26 production images - https://phabricator.wikimedia.org/T437509#12351993 (10Jdforrester-WMF) >>! In T437509#12344663, @MoritzMuehlenhoff wrote: > @Jdforrester-WMF This is now available in our Docker registry: https://docker-registry.wikimedia.org/nodejs26-sli... [18:51:31] mutante: Nevermind! I figured out I have sudo access to the right user. [18:51:44] dancy: what do you ... heh. ok :) [18:51:54] !log dancy@deploy1003 Installing scap version "4.291.0" for 3 host(s) [18:51:54] (03CR) 10Btullis: [C:03+1] use_linux612_on_bookworm: Bump kernel to 6.12.107 [puppet] - 10https://gerrit.wikimedia.org/r/1343856 (owner: 10Muehlenhoff) [18:52:20] mutante: Thanks! [18:52:26] (03CR) 10Andrea Denisse: "I'm not sure why the experimental build is failing, it says error but the output doens't show any kind of error..." [puppet] - 10https://gerrit.wikimedia.org/r/1343838 (https://phabricator.wikimedia.org/T438801) (owner: 10Andrea Denisse) [18:53:44] !log dancy@deploy1003 Installation of scap version "4.291.0" completed for 3 hosts [18:53:58] !log dancy@deploy1003 Installing scap version "4.291.0" for 2 host(s) [18:55:02] !log dancy@deploy1003 Installation of scap version "4.291.0" completed for 2 hosts [18:55:29] (03CR) 10Andrea Denisse: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343838 (https://phabricator.wikimedia.org/T438801) (owner: 10Andrea Denisse) [18:56:36] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet [18:56:37] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet [18:56:40] !log jclark@cumin1004 START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie [18:56:43] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet [18:56:52] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12352037 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jclark@cumin1004 for host ml-serve1016.eqiad.wmnet with OS trixie [18:58:46] !log btullis@cumin1004 END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet [19:01:37] !log Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0 [19:01:39] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [19:02:33] hmm, can't tell is the cpu governor helped much :( [19:02:55] (03CR) 10Andrea Denisse: [C:03+2] kubernetes: Register the aux Kubernetes service [puppet] - 10https://gerrit.wikimedia.org/r/1343838 (https://phabricator.wikimedia.org/T438801) (owner: 10Andrea Denisse) [19:04:54] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-a2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438865#12352069 (10phaultfinder) [19:11:32] (03CR) 10LWatson: [C:03+1] ReaderExperiments: Drop the stale ShareHighlight config var [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344050 (https://phabricator.wikimedia.org/T424764) (owner: 10Eric Gardner) [19:14:02] (03CR) 10LWatson: [C:03+1] ReaderExperiments: Enable on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344048 (owner: 10Eric Gardner) [19:14:54] 10ops-codfw, 10ops-eqiad, 06SRE, 06DC-Ops, and 2 others: Enable Performance Governor on Cirrussearch hosts - https://phabricator.wikimedia.org/T438877#12352198 (10bking) We've enabled the performance governor on row D hosts (starting at about 18:45 UTC, about 25 minutes before I wrote this). We haven't see... [19:15:34] (03CR) 10LWatson: [C:03+1] ReaderExperiments: Set the preferred-sources debug flag on testwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344049 (https://phabricator.wikimedia.org/T436692) (owner: 10Eric Gardner) [19:15:48] !log jclark@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage [19:16:47] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12352203 (10Jclark-ctr) After talking to @elukey and @cmooney, it looks like the only bootable ports are the RJ45 ports on the rear. I temporarily ran a 25ft Cat6 cable to... [19:17:08] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet [19:17:09] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet [19:17:15] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet [19:18:35] (03PS1) 10Ebernhardson: admin_ng: add the opensearch-semantic-search-ssd namespace [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344062 (https://phabricator.wikimedia.org/T438058) [19:18:37] (03PS1) 10Ebernhardson: opensearch-semantic-search-ssd: new release on local SSDs [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344063 (https://phabricator.wikimedia.org/T438058) [19:19:03] !log jclark@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage [19:19:21] !log cdobbins@cumin1004 START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie [19:19:38] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-b2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438869#12352235 (10phaultfinder) [19:21:41] !log sukhe@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [19:22:02] (03PS2) 10Dzahn: add attribution static site to helmfile, create values file [deployment-charts] - 10https://gerrit.wikimedia.org/r/1341335 (https://phabricator.wikimedia.org/T437635) [19:22:05] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [19:22:12] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12352241 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by sukhe@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed ho... [19:23:19] !log btullis@cumin1004 END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet [19:24:56] !log sukhe@cumin1004 START - Cookbook sre.hosts.rename from sretest2013 to cp2059 [19:25:15] (03PS1) 10Ayounsi: Add hiera facts for the "pool" action of the rack depool cookbook [puppet] - 10https://gerrit.wikimedia.org/r/1344065 (https://phabricator.wikimedia.org/T327300) [19:25:21] !log sukhe@cumin1004 END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059 [19:25:28] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12352256 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.rename started by sukhe@cumin1004 from sretest2013 to cp2059 completed: - sretest2013 (**WARN**) - ✔️ Downtimed ho... [19:33:47] !log jclark@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004" [19:34:02] !log jclark@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004" [19:34:04] !log jclark@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie [19:34:13] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12352303 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jclark@cumin1004 for host ml-serve1016.eqiad.wmnet with OS trixie completed: - ml-serve1016... [19:37:44] (03CR) 10Bking: [C:03+1] "I'm assuming the existing opensearch-semantic-search-test will be used to compare the performance, and that's why we need the new namespac" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344062 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [19:38:09] (03CR) 10Ryan Kemper: [C:03+1] cirrus: Send more_like traffic to eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 (owner: 10Ebernhardson) [19:38:19] (03CR) 10Bking: [C:03+1] cirrus: Send more_like traffic to eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 (owner: 10Ebernhardson) [19:38:51] !log cdobbins@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage [19:39:12] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12352335 (10Jclark-ctr) Server is up right now with 10g. if 25g is needed we will need 25g sfp's ordered @klausman ` jclark@ml-serve1016:~$ ip addr 1: lo: 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12352337 (10Jclark-ctr) [19:42:11] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet [19:42:12] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet [19:42:18] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet [19:42:20] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage [19:42:34] (03PS3) 10Ebernhardson: cirrus: Send more_like traffic to eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 [19:42:47] (03PS1) 10AOkoth: devtools: setup new deploy host [puppet] - 10https://gerrit.wikimedia.org/r/1344075 (https://phabricator.wikimedia.org/T429494) [19:44:21] !log btullis@cumin1004 END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet [19:44:28] I have a patch to check and see whether to deploy in 16 mins [19:46:40] (03CR) 10Bking: [C:03+2] admin_ng: add the opensearch-semantic-search-ssd namespace [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344062 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [19:48:21] (03PS1) 10Ssingh: Revert "sretest2013: remove manual references to this host" [puppet] - 10https://gerrit.wikimedia.org/r/1344077 [19:48:39] (03PS1) 10Ssingh: Revert "cp2059: add this host as a text node" [puppet] - 10https://gerrit.wikimedia.org/r/1344078 [19:49:17] (03PS3) 10Andrea Denisse: oncall: Add Klaxon to the aux clusters [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344072 (https://phabricator.wikimedia.org/T438897) [19:49:17] (03CR) 10Andrea Denisse: "Please don't merge it untill we publish the v0.1.0 image on GitLab." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344072 (https://phabricator.wikimedia.org/T438897) (owner: 10Andrea Denisse) [19:49:42] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-b2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438869#12352366 (10phaultfinder) [19:50:35] (03CR) 10Ssingh: [C:03+2] Revert "cp2059: add this host as a text node" [puppet] - 10https://gerrit.wikimedia.org/r/1344078 (owner: 10Ssingh) [19:50:49] (03CR) 10Ssingh: [C:03+2] Revert "sretest2013: remove manual references to this host" [puppet] - 10https://gerrit.wikimedia.org/r/1344077 (owner: 10Ssingh) [19:54:26] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 22 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 (owner: 10Ebernhardson) [19:54:30] FIRING: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker1021 has a BGP session which is not in the 'established' state. [19:55:01] (03PS1) 10Bking: dse-k8s: Add opensearch-semantic-search-ssd namespace [puppet] - 10https://gerrit.wikimedia.org/r/1344081 (https://phabricator.wikimedia.org/T438058) [19:55:34] (03PS4) 10Ebernhardson: cirrus: Send more_like traffic to eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 [19:57:26] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344081 (https://phabricator.wikimedia.org/T438058) (owner: 10Bking) [19:58:34] 10ops-eqiad, 06SRE, 06DC-Ops: krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12352395 (10VRiley-WMF) @MoritzMuehlenhoff Hey, I was checking with Brian on this and he mentioned to double check with you. Should we put this server as a standby unit until we may need it? [19:58:48] (03CR) 10Ryan Kemper: [C:03+1] dse-k8s: Add opensearch-semantic-search-ssd namespace [puppet] - 10https://gerrit.wikimedia.org/r/1344081 (https://phabricator.wikimedia.org/T438058) (owner: 10Bking) [19:59:10] (03CR) 10Bking: [C:03+2] dse-k8s: Add opensearch-semantic-search-ssd namespace [puppet] - 10https://gerrit.wikimedia.org/r/1344081 (https://phabricator.wikimedia.org/T438058) (owner: 10Bking) [19:59:21] PROBLEM - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [19:59:35] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: #bothumor I � Unicode. All rise for UTC late backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T2000). [20:00:05] codenamenoreste, stephanebisson, Dreamy_Jazz, and ebernhardson: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:00:10] \o [20:00:26] My patch was already deployed, forgot to remove [20:02:05] !log sukhe@cumin1004 START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie [20:02:14] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12352409 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by sukhe@cumin1004 for host sretest2013.codfw.wmnet with OS trixie [20:02:29] i can ship these, looking through them it seems plausible to ship them together if people are arround [20:02:54] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet [20:02:55] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet [20:03:01] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet [20:03:18] (03CR) 10Dzahn: [C:03+1] devtools: setup new deploy host [puppet] - 10https://gerrit.wikimedia.org/r/1344075 (https://phabricator.wikimedia.org/T429494) (owner: 10AOkoth) [20:03:20] RESOLVED: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [20:04:30] FIRING: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker1022 has a BGP session which is not in the 'established' state. [20:05:19] haven't heard from anyone, proceeding with my config patch [20:05:36] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ebernhardson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 (owner: 10Ebernhardson) [20:06:20] FIRING: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [20:06:32] (03Merged) 10jenkins-bot: cirrus: Send more_like traffic to eqiad [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1344014 (owner: 10Ebernhardson) [20:06:52] !log ebernhardson@deploy1003 Started scap sync-world: Backport for [[gerrit:1344014|cirrus: Send more_like traffic to eqiad]] [20:06:55] FYI see https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1342825 [20:09:29] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499#12352427 (10Dzahn) Hi @elukey today I was missing some images from the docker-registry related to zuul. I setup a new mach... [20:10:21] codenamenoreste: i can ship yours once this finishes [20:11:16] !log ebernhardson@deploy1003 ebernhardson: Backport for [[gerrit:1344014|cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:12:19] !log ebernhardson@deploy1003 ebernhardson: Continuing with deployment [20:12:40] !log bking@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [20:13:06] !log cdobbins@cumin1004 conftool action : set/pooled=yes; selector: name=ncredir2002.* [20:13:43] (03CR) 10Bking: [C:03+1] dse-k8s: allow unauthenticated reads of the OIDC discovery endpoints [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344034 (https://phabricator.wikimedia.org/T435591) (owner: 10Btullis) [20:13:59] FIRING: CirrusSearchNodeIndexingNotIncreasing: OpenSearch instance cirrussearch1120-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [20:14:30] RESOLVED: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker1022 has a BGP session which is not in the 'established' state. [20:14:45] FIRING: [4x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker1022 has a BGP session which is not in the 'established' state. [20:15:08] !log bking@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [20:15:30] RESOLVED: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker1022 has a BGP session which is not in the 'established' state. [20:17:13] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:17:21] !log ebernhardson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1344014|cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s) [20:17:45] codenamenoreste: still around? I can start shipping yours [20:18:20] 👍🏻 [20:18:39] (03CR) 10TrainBranchBot: [C:03+2] "Approved by ebernhardson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342825 (https://phabricator.wikimedia.org/T436652) (owner: 10Codename Noreste) [20:19:35] I'll deploy my change after [20:19:39] (03Merged) 10jenkins-bot: eswiki: Add abusefilter-access-protected-vars to abusefilter user group [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342825 (https://phabricator.wikimedia.org/T436652) (owner: 10Codename Noreste) [20:19:59] !log ebernhardson@deploy1003 Started scap sync-world: Backport for [[gerrit:1342825|eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] [20:20:03] T436652: Add abusefilter-access-protected-vars to the abusefilter group on Spanish Wikipedia - https://phabricator.wikimedia.org/T436652 [20:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [20:24:20] !log ebernhardson@deploy1003 ebernhardson, codenamenoreste: Backport for [[gerrit:1342825|eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:24:52] perfect timing :) Will wait a sec, assuming codename comes back [20:25:14] codenamenoreste: it's up on test servers, can you test? [20:26:08] sure, let me bring out the laptop [20:27:50] RESOLVED: CirrusSearchMoreLikeLatencyTooHigh: CirrusSearch more_like 95th percentiles latency is too high (mw@codfw to dnsdisc) - https://wikitech.wikimedia.org/wiki/Search#Health/Activity_Monitoring - https://grafana.wikimedia.org/d/dc04b9f2-b8d5-4ab6-9482-5d9a75728951/elasticsearch-percentiles?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchMoreLikeLatencyTooHigh [20:28:13] Turning on k8-debug the right is definitely included [20:28:32] awesome, deploying [20:28:32] Go ahead [20:28:36] !log ebernhardson@deploy1003 ebernhardson, codenamenoreste: Continuing with deployment [20:33:04] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet [20:33:33] !log ebernhardson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342825|eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s) [20:33:37] T436652: Add abusefilter-access-protected-vars to the abusefilter group on Spanish Wikipedia - https://phabricator.wikimedia.org/T436652 [20:34:29] (03CR) 10TrainBranchBot: [C:03+2] "Approved by sbisson@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342285 (https://phabricator.wikimedia.org/T433293) (owner: 10Sbisson) [20:35:29] (03Merged) 10jenkins-bot: Keep Article Guidance on where it is on today [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342285 (https://phabricator.wikimedia.org/T433293) (owner: 10Sbisson) [20:35:35] RECOVERY - Host cirrussearch1120 is UP: PING OK - Packet loss = 0%, RTA = 0.36 ms [20:35:49] !log sbisson@deploy1003 Started scap sync-world: Backport for [[gerrit:1342285|Keep Article Guidance on where it is on today (T433293)]] [20:35:53] T433293: Introduce a global on/off switch for the extension and make it available via CommunityConfiguration - https://phabricator.wikimedia.org/T433293 [20:36:09] PROBLEM - OpenSearch health check for shards on 9400 on cirrussearch1120 is CRITICAL: CRITICAL - elasticsearch http://localhost:9400/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9400): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [20:40:01] !log sbisson@deploy1003 sbisson: Backport for [[gerrit:1342285|Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:40:08] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet [20:40:09] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet [20:40:15] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet [20:40:45] !log sbisson@deploy1003 sbisson: Continuing with deployment [20:43:28] !log bking@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [20:44:21] !log bking@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [20:44:39] !log bking@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [20:44:50] (03CR) 10Andrea Denisse: "The image is published, this can be merged: https://gitlab.wikimedia.org/repos/sre/klaxon/-/pipelines/230080" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344072 (https://phabricator.wikimedia.org/T438897) (owner: 10Andrea Denisse) [20:45:09] PROBLEM - OpenSearch health check for shards on 9200 on cirrussearch1120 is CRITICAL: CRITICAL - elasticsearch http://localhost:9200/_cluster/health error while fetching: HTTPConnectionPool(host=localhost, port=9200): Read timed out. (read timeout=4) https://wikitech.wikimedia.org/wiki/Search%23Administration [20:45:32] !log aqu@deploy1003 Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] [20:45:43] !log sbisson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1342285|Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s) [20:45:47] T433293: Introduce a global on/off switch for the extension and make it available via CommunityConfiguration - https://phabricator.wikimedia.org/T433293 [20:45:57] !log bking@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [20:46:13] !log aqu@deploy1003 Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s) [20:48:04] !log aqu@deploy1003 Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] [20:49:17] FIRING: JobQueueLowTrafficConsumerWidespreadHighLatency: ... [20:49:17] Processing delay times for low-traffic consumer rules are unusually high - https://wikitech.wikimedia.org/wiki/MediaWiki_JobQueue/Operations#JobQueueLowTrafficConsumerWidespreadHighLatency - https://grafana.wikimedia.org/d/fe130675-0c2d-4991-9dec-f54cf6a9c4d8/jobqueue-low-traffic-jobs?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DJobQueueLowTrafficConsumerWidespreadHighLatency [20:50:01] RECOVERY - OpenSearch health check for shards on 9400 on cirrussearch1120 is OK: OK - elasticsearch status production-search-omega-eqiad: cluster_name: production-search-omega-eqiad, status: green, timed_out: False, number_of_nodes: 25, number_of_data_nodes: 25, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1732, active_shards: 5195, relocating_shards: 2, initializing_shards: 0, unassigned_shards: 0, de [20:50:01] assigned_shards: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:50:52] (03PS1) 10RLazarus: alertmanager: Repoint serviceops task to the correct project [puppet] - 10https://gerrit.wikimedia.org/r/1344095 [20:51:01] RECOVERY - OpenSearch health check for shards on 9200 on cirrussearch1120 is OK: OK - elasticsearch status production-search-eqiad: cluster_name: production-search-eqiad, status: green, timed_out: False, number_of_nodes: 55, number_of_data_nodes: 55, discovered_master: True, discovered_cluster_manager: True, active_primary_shards: 1286, active_shards: 3843, relocating_shards: 5, initializing_shards: 0, unassigned_shards: 0, delayed_unassi [20:51:01] rds: 0, number_of_pending_tasks: 0, number_of_in_flight_fetch: 0, task_max_waiting_in_queue_millis: 0, active_shards_percent_as_number: 100.0 https://wikitech.wikimedia.org/wiki/Search%23Administration [20:51:18] (03PS2) 10RLazarus: alertmanager: Repoint serviceops task to the correct project [puppet] - 10https://gerrit.wikimedia.org/r/1344095 [20:54:56] 06SRE, 10Beta-Cluster-Infrastructure, 06Traffic: Beta cluster haproxy does not support `warn-blocked-traffic-after` keyword - https://phabricator.wikimedia.org/T428052#12352640 (10Southparkfan) 05Open→03Resolved a:03BCornwall Thanks to Brett's work on T436468, this hack is no longer needed. I have... [20:55:04] !log aqu@deploy1003 Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s) [20:57:00] (03PS1) 10Dzahn: Revert "site: add zuul1005 as a zuul executor" [puppet] - 10https://gerrit.wikimedia.org/r/1344096 [20:57:35] (03PS1) 10PipelineBot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344097 [20:58:04] (03CR) 10Dzahn: [C:03+2] Revert "site: add zuul1005 as a zuul executor" [puppet] - 10https://gerrit.wikimedia.org/r/1344096 (owner: 10Dzahn) [20:58:39] FIRING: [2x] CirrusSearchNodeIndexingNotIncreasing: OpenSearch instance cirrussearch1120-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [20:59:10] !log aqu@deploy1003 Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] [20:59:21] RECOVERY - Check unit status of httpbb_kubernetes_mw-web_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-web_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [20:59:40] !log aqu@deploy1003 Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s) [20:59:49] !log aqu@deploy1003 Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] [21:00:05] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T2100) [21:04:45] PROBLEM - PyBal backends health check on lvs2013 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs2021.codfw.wmnet, wdqs2022.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [21:05:10] !log sukhe@cumin1004 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie [21:05:22] 10ops-codfw, 06SRE, 06DC-Ops: Q1:rack/setup/install sretest2013 (configF test host) - https://phabricator.wikimedia.org/T436691#12352665 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by sukhe@cumin1004 for host sretest2013.codfw.wmnet with OS trixie executed with errors: - sretest20... [21:07:39] RECOVERY - PyBal backends health check on lvs2013 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [21:09:37] (03Abandoned) 10JHathaway: WIP: Move DB backups from cumin1003 to cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1341326 (https://phabricator.wikimedia.org/T427897) (owner: 10JHathaway) [21:09:42] (03Abandoned) 10JHathaway: WIP - stdlib [puppet] - 10https://gerrit.wikimedia.org/r/1337995 (owner: 10JHathaway) [21:09:50] (03Abandoned) 10JHathaway: WIP: stdlib [puppet] - 10https://gerrit.wikimedia.org/r/1342382 (owner: 10JHathaway) [21:10:20] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet [21:15:23] ^ will look at wdqs pybal [21:15:47] almost certainly codfw is getting slammed by something [21:18:06] jouncebot: nowandnext [21:18:06] For the next 0 hour(s) and 41 minute(s): Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260922T2100) [21:18:06] In 4 hour(s) and 41 minute(s): Automatic deployment of MediaWiki to pretrain wikis - see mw:Pretrain (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260923T0200) [21:18:20] might sneak out a MW Apache config change, if nobody yells stop [21:18:57] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet [21:18:58] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet [21:19:04] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet [21:20:41] 06SRE, 06Infrastructure-Foundations, 10Puppet-Infrastructure: Upgrade stdlib to remove legacy facts - https://phabricator.wikimedia.org/T438912 (10jhathaway) 03NEW [21:21:02] (03PS1) 10JHathaway: stdlib upgrade: remove is_integer() calls [puppet] - 10https://gerrit.wikimedia.org/r/1344099 (https://phabricator.wikimedia.org/T438912) [21:23:39] FIRING: [2x] CirrusSearchNodeIndexingNotIncreasing: OpenSearch instance cirrussearch1120-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [21:23:48] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344099 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [21:24:10] !log aqu@deploy1003 Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s) [21:24:13] !log aqu@deploy1003 Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] [21:24:20] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1344099 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [21:25:22] !log aqu@deploy1003 Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s) [21:26:50] (03CR) 10JHathaway: [C:03+2] stdlib upgrade: remove is_integer() calls [puppet] - 10https://gerrit.wikimedia.org/r/1344099 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [21:27:31] (03CR) 10RLazarus: [V:03+1 C:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9472/co" [puppet] - 10https://gerrit.wikimedia.org/r/1339694 (https://phabricator.wikimedia.org/T437403) (owner: 10Zabe) [21:28:23] (03CR) 10RLazarus: [V:03+1 C:03+2] Add Apache configuration for wikipedia-ar-arbcom.wikimedia.org [puppet] - 10https://gerrit.wikimedia.org/r/1339694 (https://phabricator.wikimedia.org/T437403) (owner: 10Zabe) [21:28:38] RESOLVED: CirrusSearchNodeIndexingNotIncreasing: OpenSearch instance cirrussearch1120-production-search-eqiad is not indexing - https://wikitech.wikimedia.org/wiki/Search/OpenSearch/Administration#Alerts/Dashboards - https://grafana.wikimedia.org/d/JLK3I_siz/elasticsearch-indexing?orgId=1&from=now-3d&to=now&viewPanel=57 - https://alerts.wikimedia.org/?q=alertname%3DCirrusSearchNodeIndexingNotIncreasing [21:30:23] 07sre-alert-triage, 06Data-Platform-SRE: Alert in need of triage: PuppetFailure (instance apifeatureusage1001:9100) - https://phabricator.wikimedia.org/T438701#12352734 (10bking) 05Open→03Resolved a:03bking [[ https://wikimedia.slack.com/archives/C055QGPTC69/p1789990811096739 | Per this Slack message... [21:30:37] 07sre-alert-triage, 06Data-Platform-SRE (2026-09-18 - 2026-10-09): Alert in need of triage: PuppetFailure (instance apifeatureusage1001:9100) - https://phabricator.wikimedia.org/T438701#12352737 (10bking) p:05Triage→03Low [21:41:26] (03PS1) 10JHathaway: stdlib upgrade: Remove use of compat types [puppet] - 10https://gerrit.wikimedia.org/r/1344101 (https://phabricator.wikimedia.org/T438912) [21:42:48] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12352808 (10wiki_willy) Hi @klausman - to follow up, do you need 25g on this one because it's a MI350X instead of a MI300X? Just trying to see what changed, and if we nee... [21:44:02] (03CR) 10JHathaway: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1344101 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [21:47:44] !log rzl@deploy1003 Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 T437403 [21:47:48] T437403: Create arbcom_arwiki - https://phabricator.wikimedia.org/T437403 [21:49:06] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet [21:50:50] (03CR) 10Btullis: opensearch-semantic-search-ssd: new release on local SSDs (035 comments) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344063 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [21:50:59] !log rzl@deploy1003 rzl: https://gerrit.wikimedia.org/r/1339694 T437403 synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [21:51:13] (03CR) 10JHathaway: "ready for review!" [puppet] - 10https://gerrit.wikimedia.org/r/1344101 (https://phabricator.wikimedia.org/T438912) (owner: 10JHathaway) [21:53:32] !log rzl@deploy1003 rzl: Continuing with deployment [21:57:16] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet [21:57:17] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet [21:57:23] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet [21:57:58] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet [21:58:26] !log rzl@deploy1003 Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 T437403 (duration: 13m 43s) [21:58:30] T437403: Create arbcom_arwiki - https://phabricator.wikimedia.org/T437403 [22:06:25] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet [22:06:26] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet [22:06:32] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet [22:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:10:26] (03PS1) 10Lerickson: Enable wdqs-proxy metrics on port 9100. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344104 (https://phabricator.wikimedia.org/T438789) [22:12:20] (03CR) 10Lerickson: "Note: I haven't tested this yet because I need a build that enables management and the metrics path. Once I have one (see the linked wdqs-" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1344104 (https://phabricator.wikimedia.org/T438789) (owner: 10Lerickson) [22:16:22] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-a2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438865#12352930 (10Jhancock.wm) something might need to come out of the rack. the switchover put it over more than just a brief surge [22:17:20] (03CR) 10Zabe: "No worries:)" [puppet] - 10https://gerrit.wikimedia.org/r/1339694 (https://phabricator.wikimedia.org/T437403) (owner: 10Zabe) [22:21:24] (03PS1) 10Ryan Kemper: wdqs: Extend codfw main restart cooldown [puppet] - 10https://gerrit.wikimedia.org/r/1344106 (https://phabricator.wikimedia.org/T242453) [22:23:47] (03PS2) 10Ryan Kemper: wdqs: Extend codfw main restart cooldown [puppet] - 10https://gerrit.wikimedia.org/r/1344106 (https://phabricator.wikimedia.org/T242453) [22:24:01] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1344106 (https://phabricator.wikimedia.org/T242453) (owner: 10Ryan Kemper) [22:26:40] (03PS3) 10Ryan Kemper: wdqs: Extend codfw main restart cooldown [puppet] - 10https://gerrit.wikimedia.org/r/1344106 (https://phabricator.wikimedia.org/T242453) [22:27:52] (03CR) 10Ryan Kemper: [C:03+2] wdqs: Extend codfw main restart cooldown [puppet] - 10https://gerrit.wikimedia.org/r/1344106 (https://phabricator.wikimedia.org/T242453) (owner: 10Ryan Kemper) [22:30:21] !log [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts [22:30:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [22:36:35] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet [22:45:19] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet [22:45:20] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet [22:45:25] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet [22:46:39] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:09:46] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-a2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438865#12353014 (10phaultfinder) [23:11:59] 10ops-eqiad, 06SRE, 06DC-Ops: Q4: eqiad: (12) PDUs for ML expansion - https://phabricator.wikimedia.org/T400778#12353016 (10wiki_willy) [23:15:28] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet [23:23:36] !log btullis@cumin1004 START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet [23:23:37] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet [23:23:37] !log btullis@cumin1004 END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{dse-k8s-worker10[02-28].eqiad.wmnet} and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad) [23:39:49] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1344112 [23:39:49] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1344112 (owner: 10TrainBranchBot) [23:52:47] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1344112 (owner: 10TrainBranchBot)