[00:05:24] FIRING: [4x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:10:24] FIRING: [5x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:10:25] FIRING: [5x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:01:29] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12326132 (10SLyngshede-WMF) [08:10:39] FIRING: [5x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:12:30] 06Traffic, 10Liberica, 10Prod-Kubernetes, 06ServiceOps, and 2 others: Migrate Wikikube k8s apiserver and services to IPIP - https://phabricator.wikimedia.org/T420436#12326354 (10Jelto) [08:43:51] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12326457 (10SLyngshede-WMF) Traffic is moving to ULSFO: https://grafana.wikimedia.org/d/000000093/cdn-frontend-network?from=now-30m&to=now&timezone... [09:22:52] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12326599 (10ayounsi) [09:43:21] 10netops, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-28 - 2026-09-18), 07Essential-Work, 07Kubernetes: Calico IPv4 block exhaustion on dse-k8s cluster, blocking new node provisioning - https://phabricator.wikimedia.org/T429773#12326742 (10Gehel) [09:47:38] FIRING: [6x] LVSRealserverMSS: Unexpected MSS value on 103.102.166.224:443 @ cp5025 - https://wikitech.wikimedia.org/wiki/LVS#LVSRealserverMSS_alert - https://grafana.wikimedia.org/d/Y9-MQxNSk/ipip-encapsulated-services?orgId=1&viewPanel=2&var-site=eqsin&var-cluster=cache_upload - https://alerts.wikimedia.org/?q=alertname%3DLVSRealserverMSS [09:49:07] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12326850 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=1f3524ce-e01c-441f-8517-6b989c2f4db7) set by slyngshede@cumin1003 for 4:00:00 on 1 host(s) and their services with rea... [09:49:39] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12326871 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=f93ce7f9-6cc1-4e08-afd7-61f4166200d0) set by slyngshede@cumin1003 for 2 days, 0:00:00 on 1 host(s) and their services... [10:07:38] RESOLVED: [6x] LVSRealserverMSS: Unexpected MSS value on 103.102.166.224:443 @ cp5025 - https://wikitech.wikimedia.org/wiki/LVS#LVSRealserverMSS_alert - https://grafana.wikimedia.org/d/Y9-MQxNSk/ipip-encapsulated-services?orgId=1&viewPanel=2&var-site=eqsin&var-cluster=cache_upload - https://alerts.wikimedia.org/?q=alertname%3DLVSRealserverMSS [10:08:38] FIRING: [6x] LVSRealserverMSS: Unexpected MSS value on 103.102.166.224:443 @ cp5025 - https://wikitech.wikimedia.org/wiki/LVS#LVSRealserverMSS_alert - https://grafana.wikimedia.org/d/Y9-MQxNSk/ipip-encapsulated-services?orgId=1&viewPanel=2&var-site=eqsin&var-cluster=cache_text - https://alerts.wikimedia.org/?q=alertname%3DLVSRealserverMSS [10:13:38] RESOLVED: [6x] LVSRealserverMSS: Unexpected MSS value on 103.102.166.224:443 @ cp5025 - https://wikitech.wikimedia.org/wiki/LVS#LVSRealserverMSS_alert - https://grafana.wikimedia.org/d/Y9-MQxNSk/ipip-encapsulated-services?orgId=1&viewPanel=2&var-site=eqsin&var-cluster=cache_text - https://alerts.wikimedia.org/?q=alertname%3DLVSRealserverMSS [11:01:55] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12327258 (10Fabfur) [11:04:31] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12327268 (10cmooney) [11:10:42] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12327289 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by slyngshede@cumin1003 for host cp5025.eqsin.wmnet with OS trixie [12:10:40] FIRING: [5x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:13:28] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12327653 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by slyngshede@cumin1003 for host cp5025.eqsin.wmnet with OS trixie completed: - cp5025 (**WARN**) - Downtimed on Ici... [12:32:16] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12327707 (10ssingh) [12:45:23] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12327750 (10cmooney) @Jclark-ctr thanks for the help with the switches in rack A1 and the cabling. Both //lsw1-a1-eqiad// and //ssw1-a1-eqiad// are reachabl... [12:53:32] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052#12327789 (10dkertesz) 05Open→03Resolved [13:02:25] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12327834 (10dkertesz) [13:15:25] FIRING: [5x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:25:25] FIRING: [5x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:30:25] RESOLVED: [3x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:31:40] FIRING: VarnishPrometheusExporterDown: Varnish Exporter on instance cp5026:9331 is unreachable - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/000000304/varnish-dc-stats?viewPanel=17 - https://alerts.wikimedia.org/?q=alertname%3DVarnishPrometheusExporterDown [13:32:43] FIRING: HaproxyKafkaExporterDown: HaproxyKafka on cp5026 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://grafana.wikimedia.org/d/d3e4e37c-c1d9-47af-9aad-a08dae2b3fd5/haproxykafka?orgId=1&var-site=eqsin&var-instance=cp5026 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [13:34:35] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12328013 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by slyngshede@cumin1003 for host cp5026.eqsin.wmnet with OS trixie [13:36:40] FIRING: VarnishPrometheusExporterDown: Varnish Exporter on instance cp6002:9331 is unreachable - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/000000304/varnish-dc-stats?viewPanel=17 - https://alerts.wikimedia.org/?q=alertname%3DVarnishPrometheusExporterDown [13:37:43] FIRING: HaproxyKafkaExporterDown: HaproxyKafka on cp6002 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://grafana.wikimedia.org/d/d3e4e37c-c1d9-47af-9aad-a08dae2b3fd5/haproxykafka?orgId=1&var-site=drmrs&var-instance=cp6002 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [13:38:21] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12328069 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by sukhe@cumin1004 for host cp6002.drmrs.wmnet with OS trixie [14:19:26] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhautions - https://phabricator.wikimedia.org/T437984#12328397 (10ayounsi) [14:30:06] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12328478 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by sukhe@cumin1004 for host cp6002.drmrs.wmnet with OS trixie completed: - cp6002 (**WARN**) - Downtimed on Icinga/A... [14:35:01] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12328545 (10ssingh) [14:42:15] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12328628 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1004 for host ncredir7003.magru.wmnet with OS trixie [14:46:06] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12328671 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by slyngshede@cumin1003 for host cp5026.eqsin.wmnet with OS trixie completed: - cp5026 (**WARN**) - Downtimed on Ici... [14:50:52] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12328723 (10SLyngshede-WMF) [14:50:59] 06Traffic, 13Patch-For-Review: Reimage cp-upload nodes as cp-text nodes - https://phabricator.wikimedia.org/T436363#12328724 (10SLyngshede-WMF) 05Open→03Resolved [15:42:19] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12329150 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1004 for host ncredir7003.magru.wmnet with OS trixie completed: - ncredir7003 (**WARN**) - Downtimed on Icin... [16:55:14] 06Traffic, 10Citoid, 06MediaWiki-API-Platform-Team, 06ServiceOps, and 3 others: Consider adding a php interface over the citoid api in the Citoid extension - https://phabricator.wikimedia.org/T435916#12329788 (10ppelberg) **Meta** @Mvolz, I'm boldly assuming this is future work that you're exploring. Accor... [16:55:22] 06Traffic, 10Citoid, 06MediaWiki-API-Platform-Team, 06ServiceOps, and 2 others: Consider adding a php interface over the citoid api in the Citoid extension - https://phabricator.wikimedia.org/T435916#12329789 (10ppelberg) [18:25:13] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12330336 (10CDobbins) [18:55:51] 06Traffic, 06Commons: Error: 503, Backend fetch failed (when uploading a 37MB SVG file) - https://phabricator.wikimedia.org/T410201#12330460 (10Aklapper) [20:14:12] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12330829 (10Jclark-ctr) @cmooney Unplugged and reseated B7, and the link came back up. I also connected the links between the D8 and D1 spines with the A1 sp...