[07:04:13] 10netops, 06Infrastructure-Foundations, 06SRE: Missing series for BGP session_state from eqiad CRs since upgrade to 23.4R2-S8.7 - https://phabricator.wikimedia.org/T435909#12278845 (10cmooney) >>! In T435909#12275570, @ayounsi wrote: > Adding a 3rd netflow host in eqiad shuffled the targets around, it caused... [07:39:28] FIRING: KeyholderUnarmed: 2 unarmed Keyholder key(s) on cumin1004:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [07:44:28] RESOLVED: KeyholderUnarmed: 2 unarmed Keyholder key(s) on cumin1004:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [08:12:25] FIRING: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:17:00] XioNoX: Sorry, moving my question regarding cp5022 so our new team member can follow along :-) [08:21:51] slyngs: looks like an issue on the host, `bast5005:~$ ssh cp5022 -vvv` shows `debug1: Connection established.` but then hangs [08:21:59] any relevant logs on the server side? [08:23:24] Ah, I just saw the console and though it was fine. It's not actually responding to anything [08:24:25] it's responding to pings, but that's an expensive server to just respond to pings [08:25:02] I'm going to power cycle it and see if it comes back, we've depooled and downtimed it for now [08:26:01] Responding to ping is more than it did previously though :-) [08:29:18] Okay, it's back now. I'll do some investigating [08:39:15] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12279180 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host netflow7002.magru.wmnet with OS trixie [08:46:37] piger: https://grafana.wikimedia.org/d/000000377/host-overview?from=now-3h&to=now&timezone=utc&var-server=cp5022&var-datasource=000000026&var-cluster=cache_text&refresh=5m [08:47:17] It's clearly having an issue at around 7:52 / 53 [08:48:14] slyngs: this is that host that was on the wrong vlan last week is it? [08:48:56] Yes, and XioNoX found an IPv6 issue yesterday and then it was reimaged and looked good, so it got repooled [08:49:07] hmm ok [08:49:37] can't imagine what might cause issues now if it's been freshly reimaged on the right IP [08:50:30] It hasn't been health, that's why it didn't get moved to the right VLAN I believe. It was down at the time [08:50:47] yeah I got the IP wrong last week when I manually changed it (https://netbox.wikimedia.org/extras/changelog/293402/) [08:51:16] yeah you said that alright. I would not be surprised if what you're seeing is not the result of the same poor hw health [08:53:58] The CPU temps are fun, they diverge when the host is repooled. Half the cores are like 10 - 15 degrees hotter than the rest [08:54:52] Ah, the other ones do that to [09:12:25] RESOLVED: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:47:01] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12280451 (10ayounsi) [12:47:56] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12280452 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1004 for host netflow6001.drmrs.wmnet with OS trixie [13:36:57] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12280737 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1004 for host netflow6001.drmrs.wmnet with OS trixie completed: - netflow6001 (**PASS**) - Downtimed... [14:51:48] 10CFSSL-PKI, 06Infrastructure-Foundations: Add a load balanced service in front of PKI - https://phabricator.wikimedia.org/T436809 (10elukey) 03NEW p:05Triage→03Medium [15:15:42] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814 (10cmooney) 03NEW p:05Triage→03Medium [15:36:16] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814#12281602 (10Jclark-ctr) @cmooney can these be pulled at anytime or do you want to be online when pulled? [16:38:25] FIRING: SystemdUnitFailed: check_netbox_uncommitted_dns_changes.service on netbox1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:43:25] RESOLVED: SystemdUnitFailed: check_netbox_uncommitted_dns_changes.service on netbox1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:51:46] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814#12282002 (10cmooney) >>! In T436814#12281602, @Jclark-ctr wrote: > @cmooney can these be pulled at anytime or do you want to be online when pulled? T... [21:23:25] FIRING: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:28:25] RESOLVED: SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed