[00:04:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:04:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:04:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:49:57] FIRING: [3x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:15:00] ^ can be ignored, WIP to move that to trixie [10:19:57] FIRING: [3x] SystemdUnitFailed: debian-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:46:17] 10netops, 06Infrastructure-Foundations, 10Observability-Alerting, 06SRE: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801 (10cmooney) 03NEW p:05Triage→03Low [14:19:57] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:38:39] 10netops, 06Infrastructure-Foundations, 10Observability-Alerting, 06SRE: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801#12247067 (10tappof) The alerts come from Pint, a linter/validator for Prometheus. It just says that the expression defined in the TransitBGPDown rul... [14:39:31] 10netops, 06Infrastructure-Foundations, 10Observability-Alerting, 06SRE: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801#12247073 (10tappof) [14:51:20] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack: wait_for_optimal() should ignore acked alerts - https://phabricator.wikimedia.org/T319277#12247196 (10LSobanski) 05In progress→03Resolved a:03SLyngshede-WMF [14:51:47] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack: wait_for_optimal() should ignore acked alerts - https://phabricator.wikimedia.org/T319277#12247202 (10LSobanski) Also related: https://gerrit.wikimedia.org/r/c/operations/software/spicerack/+/1140208 [14:55:08] 10CAS-SSO, 06Data-Platform-SRE, 06Infrastructure-Foundations: Enable the OAuth device authorization grant in CAS and register a wmf-s3-login client - https://phabricator.wikimedia.org/T435602#12247242 (10LSobanski) [14:56:03] 10CAS-SSO, 06Data-Platform-SRE, 06Infrastructure-Foundations: Design and create the LDAP groups for the Data Lake access tiers - https://phabricator.wikimedia.org/T435601#12247243 (10LSobanski) [14:56:46] 10CAS-SSO, 06Data-Platform-SRE, 06Infrastructure-Foundations: Establish LDAP access groups and login tooling for human access to the Data Lake on S3 - https://phabricator.wikimedia.org/T435473#12247247 (10LSobanski) [14:57:32] 10netops, 06SRE Observability: Prometheus rule evaluation failures (instance titan1001) - https://phabricator.wikimedia.org/T435494#12247266 (10LSobanski) [14:58:40] 10netops, 06SRE Observability: Prometheus rule evaluation failures (instance titan1001) - https://phabricator.wikimedia.org/T435494#12247271 (10cmooney) Are there still any problems here folks? I'm hoping it cleared up when we shifted the traffic away from the problematic link on Thursday? [15:17:17] Heads-up that I am reimaging `krb1003` now, ref https://netbox.wikimedia.org/dcim/devices/3166/ for hardware info [15:28:56] see my comment on https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328634, the role needs fixing [15:30:53] ACK, will fix. Do you know if that would cause reimages to fail? [15:31:14] just curious as they are failing, and I was gonna try firmware updates too if the role wasn't the problem [15:31:24] 10CAS-SSO, 06Collaboration-Services, 06Infrastructure-Foundations, 10GitLab (Auth & Access), 06Release-Engineering-Team (Radar): Add GitLab to offboarding workflow - https://phabricator.wikimedia.org/T339843#12247583 (10LSobanski) GitLab login is based on the IDP account so disabling the latter will stop... [15:33:38] OK, https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328645 is up. I'll self-merge once it passes CI unless anyone has any objections [15:34:07] +1d [15:34:21] the initial reimage would work fine, but Puppet would fail when the KDC role gets applied [15:35:36] OK, will try fw updates then [15:36:30] sounds good [16:35:01] moritzm & inflatador I wanted to do some investigation on the hardware failure on krb1002, is it okay if I reboot the host? [16:53:28] jhathaway sure, it's down now [16:53:40] thanks [16:56:20] Have y'all ever played with the on-board RAID controller on the Supermicros? I imagine it's junk but just curious [17:05:47] I have not [17:06:36] Emperor: may have? [17:08:37] I doubt it would help us here, but I noticed it when I was messing around Friday [17:11:06] the system seems to run at normal speed when booted into a rescue image [17:11:14] https://www.supermicro.com/en/products/motherboard/x12spw-tf It says it's an Intel® C621A controller....that sounds better than I would've expected [17:11:38] but I honestly have no idea [17:27:51] yeah, based on https://www.intel.com/content/www/us/en/support/articles/000056660/server-products/sasraid.html looks like it's embedded junk [17:33:58] in other news, the replacement host just reimaged, so I'll start looking at docs for deploying a new krb replica, should have patches fairly soon [17:45:37] 10netops, 06Infrastructure-Foundations, 10Observability-Metrics, 06SRE: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855 (10cmooney) 03NEW p:05Triage→03High [17:46:45] 10netops, 06Infrastructure-Foundations, 10Observability-Metrics, 06SRE: Expand blackbox icmp probes to ping specific router interfaces/circuits - https://phabricator.wikimedia.org/T435855#12248368 (10cmooney) [18:08:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:32:48] CR up for adding the replacement host as krb replica if anyone has time to look: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328688 [18:48:43] looking [18:59:35] Thanks, running puppet-merge now [19:03:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:04:51] OK merged, following the directions at https://wikitech.wikimedia.org/wiki/Data_Platform/Systems/Kerberos/Administration#Handling_failures_and_failover to stand up the replica [19:08:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:13:53] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:23:53] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:24:57] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:26:15] OK, so I've followed the directions to stand up a replica but `/usr/sbin/kprop -d -f /srv/backup/kdc_database_krepl_20260824192439 krb1003.eqiad.wmnet` is still failing due to a connection error [19:26:39] checking basic stuff now but LMK if y'all have suggestions. Docs are at https://wikitech.wikimedia.org/wiki/Data_Platform/Systems/Kerberos/Administration#Replica_node [19:29:09] well, the `krb5-kpropd.service` isn't even running on `krb1003`, let's try that first ;) [19:30:02] FIRING: [5x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:30:39] OK, that's fixed. Will still need to remove `krb1002` from some hieradata before `replicate-krb-database.service` can run without crashing [19:53:51] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328700 Another CR for bringing on the new replica. I also noticed that the replication script at https://w.wiki/ToiE bails out when a single replica fails. I'll try and get a patch up for that too [19:58:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:05:17] inflatador: +1d [20:05:37] jhathaway ACK, thanks! [20:08:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:17:53] jhathaway one more for ya, no rush https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328709 [20:19:51] +1d, bashing out erb is one of my pet peeve's, but I closed my eyes on the templating piece :P [20:39:33] Thanks, 100% agree re: templating out bash [20:48:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:53:53] FIRING: [3x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:40:26] 10netops, 06Infrastructure-Foundations, 10Observability-Alerting, 06SRE: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801#12249177 (10cmooney) Ok thanks that's a major issue with our stats pipeline then. The cr2-eqord we can ignore, if we have no stats for BGP that's a... [21:55:45] 10netops, 06Infrastructure-Foundations, 10Observability-Alerting, 06SRE: AlertLintProblem for TransitBGPDown check - https://phabricator.wikimedia.org/T435801#12249291 (10cmooney) 05Open→03Resolved a:03cmooney Actually I worked it out. We stopped getting stats on July 15 when we upgraded the rou...