[03:22:40] FIRING: SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:22:40] FIRING: SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:47:34] FIRING: DiskSpace: Disk space build2001:9100:/ 2.897% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=build2001 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [11:22:40] FIRING: SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:37:25] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:47:49] FIRING: DiskSpace: Disk space build2001:9100:/ 2.89% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=build2001 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [13:29:22] 10netops, 06Infrastructure-Foundations, 06SRE Observability: Prometheus rule evaluation failures (instance titan1001) - https://phabricator.wikimedia.org/T435494#12237036 (10tappof) [14:37:25] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:47:47] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237508 (10ssingh) [14:48:25] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237509 (10ssingh) Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the maint window of 08:00 UTC. Can you please co... [14:52:52] 10netops, 06Infrastructure-Foundations, 06SRE: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506 (10cmooney) 03NEW p:05Triage→03Medium [14:57:15] 10netops, 06Infrastructure-Foundations, 06SRE: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506#12237562 (10cmooney) Hmm.... I removed the config for both FPCs, and requested they go to "offline", however the system alarms have not cleared: ` cmooney@re0.cr2-eqiad> show ch... [15:04:18] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237593 (10RobH) >>! In T435406#12237508, @ssingh wrote: > Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the main... [15:06:38] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237604 (10RobH) [15:07:16] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237605 (10RobH) a:05ssingh→03RobH [16:47:50] FIRING: DiskSpace: Disk space build2001:9100:/ 2.89% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=build2001 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [17:00:11] 10netops, 06Infrastructure-Foundations, 06SRE: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506#12238341 (10cmooney) @VRiley was able to unseat the cards, which means the FPC slots now show as 'empty' rather than 'offline' ` cmooney@re0.cr2-eqiad> show chassis fpc... [17:40:03] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238568 (10ssingh) @RobH, @cmooney: @SLyngshede-WMF will be handling this from Traffic, just as an FYI. [18:10:53] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238743 (10cmooney) >>! In T435406#12237508, @ssingh wrote: > Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the m... [18:37:40] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:41:34] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238851 (10ssingh) >>! In T435406#12238743, @cmooney wrote: >>>! In T435406#12237508, @ssingh wrote: >> Thanks, @RobH. We will then aim to de... [19:27:35] FIRING: [2x] DiskSpace: Disk space build2001:9100:/ 2.886% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [19:35:26] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543 (10cmooney) 03NEW p:05Triage→03High [20:17:25] FIRING: [4x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:19:54] 10netops, 06Infrastructure-Foundations, 06SRE Observability: Prometheus rule evaluation failures (instance titan1001) - https://phabricator.wikimedia.org/T435494#12239145 (10Scott_French) This would likely be T435543. Impact should have resolved around 15:50 when the new circuits were depooled. [20:28:44] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239173 (10cmooney) [20:29:27] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239180 (10cmooney) [20:34:43] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239204 (10cmooney) I've sent a mail to the HE noc (cc'd noc@wikimedia) to raise a ticket about this. Let's see what they say. [20:47:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:52:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:52:34] FIRING: [2x] DiskSpace: Disk space build2001:9100:/ 2.886% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [20:57:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:02:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:07:25] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:22:25] FIRING: [4x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:57:15] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239515 (10cmooney) Ticket ID HE#7216121 [23:26:46] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12239713 (10RobH)