[08:26:27] the unit/functional/integ tests for zarcillo are green in gitlab ci [10:34:24] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db1218:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:49:13] RESOLVED: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db1218:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:03:45] hello data persistence friends - this week, I've seen a number replication lag alerts for backup source hosts popping up in -operations (presumably while snapshots are in progress?). this feels like it's happening more frequently than I recall, but I'm also not sure ... my question: is this expected, and / or did we change something about their notification settings? [13:05:57] swfrench-wmf: I think what you saw was the failure of a backup source that crashed maybe? https://phabricator.wikimedia.org/T440283 [13:06:10] They are not supposed to alert no, because they have notifications disabled [13:06:13] for when the backups run [13:07:19] thanks, marostegui! hmmm ... so, I've seen them for, e.g., db1265 [13:07:38] swfrench-wmf: then that isn't supposed to be the case, I will double check [13:09:12] sounds good, and thank you - it's not paging or anything, but I'll occasionally see them go by in -operations (and it's a bit concerning at first until you start to recognize the host names) [13:09:21] Yeah, I am seeing that they all have notifications enabled [13:09:34] And I think this could be coming from the codfw issue where icinga was mass silenced etc [13:09:47] so maybe notifications disabled were re-enabled when we did the massive clean up again [13:09:59] swfrench-wmf: they were for replication lag, right? [13:11:04] PROBLEM - MariaDB Replica Lag: s4 on db1265 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 629.05 seconds https://wikitech.wikimedia.org/wiki/MariaDB/Troubleshooting%23Incident_Response [13:11:06] yes, exactly - e.g., [13:11:06] > <+icinga-wm> PROBLEM - MariaDB Replica Lag: s3 on db2239 is CRITICAL: CRITICAL slave_sql_lag Replication lag: 641.07 seconds [...] [13:11:10] yep [13:11:33] Yeah those are supposed to be with notifications disabled, I think it was maybe reenabled on the mass icinga operation [13:11:36] I will downtime them [13:11:44] Sorry for the noise [13:11:46] amazing, thank you! :) [13:12:21] no worries at all - it's just a bit jarring to see them pop up, until you realize what hosts they are, heh [13:12:26] thanks again [13:12:31] thank you! [15:22:28] I seems that db2201 rebooted, not sure if this was planned? [15:22:34] since there was no downtime, likely not? [15:23:05] marostegui, cezmunsta ^ [15:26:07] nothing in SEL [15:26:28] That host had issues before [15:26:34] It wasn't planned [15:26:35] it's a backup souce, not live replica [15:26:39] I'll get to it [15:26:40] Thanks [15:26:55] I also poked at syslog, no sign of issues (like oom-killer or so) [15:27:07] probably best to depool it and have DC ops upgrade all firmware [15:28:27] Yeah, it crashed some days ago and I asked them to do that but giving me a heads up before [15:29:09] moritzm: https://phabricator.wikimedia.org/T440283#12402714 [15:29:49] ah, so that might have been that troubleshooting indeed [15:29:54] yep [15:30:00] I asked in -sre so confirmed [15:30:29] ah, right. thanks [16:07:11] federico3: [18:06:54] FIRING: ZarcilloDataIngestionMetricAbsent: Zarcillo data ingestion metric is missing - https://wikitech.wikimedia.org/wiki/MariaDB/Zarcillo - https://grafana.wikimedia.org/d/celrzpf6av8qob/zarcillo - https://alerts.wikimedia.org/?q=alertname%3DZarcilloDataIngestionMetricAbsent [16:07:26] heh [16:08:20] trying to update the conf... [16:10:19] mesides, should we alarm here instead? [16:12:00] also the alert does not describe where it's missing [16:12:18] yeah, we probably need o11y to review that [16:12:33] Do you happen to know when tappof is back? maybe check their calendar [16:32:49] he's going to be back on Monday [16:33:32] maybe we can merge the CR in the meantime to see if it improves