[02:39:01] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:39:01] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:03:17] 10CAS-SSO, 06Infrastructure-Foundations: Upgrade to CAS 8.0.X - https://phabricator.wikimedia.org/T439509 (10SLyngshede-WMF) 03NEW [07:03:26] 10CAS-SSO, 06Infrastructure-Foundations: Upgrade to CAS 8.0.X - https://phabricator.wikimedia.org/T439509#12372634 (10SLyngshede-WMF) p:05Triage→03Medium [07:19:01] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:25:47] ^ docker-reporter-kubernetes-aux_eqiad-images.service should recover in a bit, this was failing due to a stale etherpad image, which Jelto now re-deployed using a fresh trixie image with the fix from https://phabricator.wikimedia.org/T438866 f [08:25:01] 10netops, 06Infrastructure-Foundations: Junos: ae speed not reported corectly through gNMI - https://phabricator.wikimedia.org/T439202#12372876 (10cmooney) No objection to us working with Juniper to fix this, or even the workround with the description / rewrite. But I think for our monitoring we should be ign... [08:39:01] FIRING: [2x] SystemdUnitFailed: wmf_auto_restart_krb5-admin-server.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:55:49] 10netops, 06Infrastructure-Foundations, 13Patch-For-Review: Juniper hostkey-algorithm deprecation - https://phabricator.wikimedia.org/T439136#12372989 (10ayounsi) 05Open→03Stalled p:05High→03Low Marking it as stalled until we have upgraded all the devices. [08:56:20] 10netops, 06Infrastructure-Foundations, 13Patch-For-Review: Juniper hostkey-algorithm deprecation - https://phabricator.wikimedia.org/T439136#12372992 (10ayounsi) [08:56:21] 10netops, 06Infrastructure-Foundations, 06Traffic: 2026 Junos upgrade - https://phabricator.wikimedia.org/T416444#12372993 (10ayounsi) [09:07:59] 10netops, 06Infrastructure-Foundations, 13Patch-For-Review: Junos: ae speed not reported corectly through gNMI - https://phabricator.wikimedia.org/T439202#12373029 (10ayounsi) a:03ayounsi Great point, I sent https://gerrit.wikimedia.org/r/c/operations/alerts/+/1345934 for that. Transit/Peering interfaces... [09:26:53] 10netops, 06Infrastructure-Foundations, 13Patch-For-Review: Junos: ae speed not reported corectly through gNMI - https://phabricator.wikimedia.org/T439202#12373119 (10ayounsi) 05Open→03Resolved Deployed! [10:40:17] 10netops, 06Infrastructure-Foundations, 06SRE: eqiad row A/B switch refresh configuration - https://phabricator.wikimedia.org/T439535 (10cmooney) 03NEW p:05Triage→03Medium [10:41:26] 10netops, 06Infrastructure-Foundations, 06SRE: eqiad row A/B switch refresh configuration - https://phabricator.wikimedia.org/T439535#12373591 (10cmooney) [10:41:30] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: eqiad row A/B switch refresh prep - https://phabricator.wikimedia.org/T418012#12373592 (10cmooney) [10:46:29] 10netops, 06Infrastructure-Foundations, 06SRE: Support channelized / breakout interfaces on SR-Linux - https://phabricator.wikimedia.org/T439537 (10cmooney) 03NEW p:05Triage→03Medium [10:47:20] 10netops, 06Infrastructure-Foundations, 06SRE: eqiad row A/B switch refresh configuration - https://phabricator.wikimedia.org/T439535#12373630 (10cmooney) [10:47:21] 10netops, 06Infrastructure-Foundations, 06SRE: Support channelized / breakout interfaces on SR-Linux - https://phabricator.wikimedia.org/T439537#12373629 (10cmooney) [12:38:48] FIRING: PuppetFailure: Puppet has failed on testvm2001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [12:39:01] FIRING: [3x] SystemdUnitFailed: wmf_auto_restart_krb5-admin-server.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:48:48] RESOLVED: PuppetFailure: Puppet has failed on testvm2001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:49:01] FIRING: [3x] SystemdUnitFailed: wmf_auto_restart_krb5-admin-server.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:14:42] cjd91, topranks, https://gerrit.wikimedia.org/r/c/operations/homer/public/+/1346028 [14:15:33] that's probably blocking the VMs re-image in eqsin [14:18:41] cjd91: I manually pushed the change in eqsin, can you try the re-image agin? [14:18:42] again [14:19:07] XioNoX: yep, doing it now [14:28:16] XioNoX: +1 on the change, though I'm wondering should it have affected packets routing through et-0/0/1.100? [14:29:34] topranks: lets see how it goes, I remember the dhcp_relay trying to be too "secure" and dropping some unrelated dhcp packets [14:30:00] so as the packets transit the router to go from one rack to the other, my bet is that it's happening here [15:23:52] cjd91: how is i going? [15:28:33] this error appeared a while back, but apparently it's also starting the first puppet run? https://www.irccloud.com/pastebin/25kVPsvU/ [15:29:58] the rest of it, which I forgot to copy https://www.irccloud.com/pastebin/BXIcC3x9/ [17:49:01] FIRING: SystemdUnitFailed: wmf_auto_restart_krb5-admin-server.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:55:40] slurk: it worked, right? looks like I can ssh to the host and it's running trixie [18:55:44] er, cjd91 ^ [18:57:14] XioNoX: yes, it worked. thank you (and t.opranks) for your help! [18:57:16] * slurk falls back asleep [21:49:01] FIRING: SystemdUnitFailed: wmf_auto_restart_krb5-admin-server.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:22:01] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Standardize management routers interfaces - https://phabricator.wikimedia.org/T421674#12377400 (10Papaul)