[07:35:15] FIRING: [2x] ProbeDown: Service idp2005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:40:15] RESOLVED: [2x] ProbeDown: Service idp2005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [08:13:07] 10netops, 10Cloud-VPS, 06Data-Platform-SRE, 10Data-Services, and 4 others: Plan to make clouddumps more resilient and easier to operate - https://phabricator.wikimedia.org/T411248#12260584 (10fgiunchedi) [16:56:25] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:11:25] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:21:25] RESOLVED: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:22:55] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:57:03] ^^ checking out the krb alerts here [18:01:06] OK, we should be good, I just needed to restart the service after putting the keytab in place. [18:01:14] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1330562 is ready for review if anyone has the cycles [18:04:55] FIRING: SystemdUnitFailed: replicate-krb-database.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:13:10] FIRING: [2x] SystemdUnitFailed: krb5-admin-server.service on krb1004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:55:15] ^^ looking again at the failures [19:10:55] RESOLVED: SystemdUnitFailed: replicate-krb-database.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:11:35] ^^ something is wrong with the new krb VM, the disk is going read-only and crashing the service. I'm gonna try migrating to another ganeti host [20:56:27] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12263461 (10cmooney) [20:57:42] 10netops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE, 06Traffic: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12263463 (10cmooney) 05Open→03Resolved I'm gonna close this one, everything went smoothly on the day. We did not get time to factory...