[00:04:25] FIRING: SystemdUnitFailed: uwsgi-graphite-web.service on graphite2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:09:25] FIRING: [2x] SystemdUnitFailed: uwsgi-graphite-web.service on graphite1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:10:20] ^ I think they're related to the Graphite decom, I'll silence them. [01:09:25] FIRING: SystemdUnitFailed: wmf_auto_restart_uwsgi-graphite-web.service on graphite2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:54:25] FIRING: [2x] SystemdUnitFailed: wmf_auto_restart_uwsgi-graphite-web.service on graphite1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:54:48] Ah, I've silenced them already, let me look... [03:57:04] Okay, so I silenced uwsgi-graphite-web.service but this one is wmf_auto_restart_uwsgi-graphite-web lol, silenced both on both hosts. [09:20:20] hello! any idea what's wrong with my alert in https://gerrit.wikimedia.org/r/c/operations/alerts/+/1321906 ? and how to fix it? thanks [09:27:52] XioNoX: I think it's because gnmi_interfaces_interface_state_counters_in_octets you added is the same as the earlier instance which assumes an interval of 2m. so you can halve the link or double the traffic [09:30:15] hnowlan: where do you see that? I thought the error was about the labels "AssertionError: Alert TransportLinksInUsageNoRedundancy is not going to fire: external label reference in non-global alert (team-netops/interfaces.yaml)" ? [09:32:17] oh, I just saw the last error which is that the alert didn't fire at all in the test. one sec [09:34:56] thinking out loud it looks for : `EXT_LABELS_RE = re.compile(r"{.*(prometheus|site)\s*[=!~]")` [09:35:34] yeah if you change the values for the metrics the tests pass [09:35:43] I think it's looking for a label that isn't there, although it's quite confusing [09:35:54] ah! [09:36:42] what did you change? [09:39:10] I changed gnmi_interfaces_interface_state_high_speed to 10000x35 - but maybe doubling the traffic would be a more representative test? [09:40:31] I wonder if there's a way to make the pint errors a bit more human-friendly, they're fairly obtuse [09:44:21] hnowlan: thanks, pushed a new PS [09:44:45] yeah that would be great, but no idea how doable it is [09:45:29] at the least there's some easy documentation work to make it more comprehensible, filed https://phabricator.wikimedia.org/T436621 [09:48:07] hnowlan: hmm, CI is still failing [09:54:22] XioNoX: oh hmm - the other fail is real, my bad. Looking [10:02:25] so this is a bit of a weird one that kinda underscores the need for better UX here :P [10:02:46] because the deploy-tag is set to ops, the `site` label isn't available [10:04:02] in tests at least [10:04:08] it will be added in real life by the instance [10:04:22] quick fix is to just glob for the DC in the instance label [10:04:50] you could change the alert to be global but tbh I am not qualified to say what that will do - maybe that makes more sense practically given the kind of alert [10:04:55] tappof: any thoughts on that? [10:13:43] I'll take a look [10:31:00] hnowlan: XioNoX It's a bit of an edge case. Normally, the site label is added as an external label, so it is only visible to Thanos components. In this specific case (and a few others), the site label is also added by the scrape job itself. [10:31:04] I'll check whether we can add exceptions in CI or whether we'll need to move the rule to Thanos Ruler. [10:31:14] I'll let you know. [10:33:02] tappof: thanks, hnowlan's workaround idea seems fine to me too [10:35:14] yeah XioNoX, I think so. You can add a file named interfaces_global.yaml under team-netops, with # deploy-tag: global at the top of the file.. [10:36:21] then you must also move the test to a new file named `interfaces_global_test.yaml`. [10:37:06] ok [10:39:53] FWIW I can confirm: it looks like you want/need to alert on metrics across sites (?) that's the job for global alerts [10:40:13] not across sites, no, only for each POP [10:40:19] (excluding core sites) [10:42:11] tests pass locally, pushing to gerrit [10:42:14] I see ok, yes another way to achieve the same without global would be another alert file with # deploy-site: list of all pops [10:42:54] IIRC there was a task open about supporting exclusion for deploy tags, so you could say !eqiad !codfw for example [10:43:06] not sure, as the metric exist in all the POPs, it's just a "filter" on the alert that excludes the core sites [10:43:29] iirc exclusions are only if the metric doesn't exist in the specific sites [10:44:10] yes slightly different context in the sense that I'm thinking about the deploy tags, i.e. "deploy this alert file only in these sites" [10:44:22] > another way to achieve the same without global would be another alert file with # deploy-site: list of all pops [10:44:34] I believe this won't pass the CI because the rule relies on a label that is typically external and therefore usually missing on leaf Prometheus instances.. [10:44:50] alright, thanks all, CI is happy ! https://gerrit.wikimedia.org/r/c/operations/alerts/+/1321906 [10:45:09] true, in the case I mentioned the site!=eqiad,codfw label matching wouldn't be needed [10:45:27] to be clear: global is 100% good IMO [10:46:33] Yeah, yeah... any input from godog here is always sincerely appreciated :) [10:47:06] <3 <3 thank you tappof, appreciate it [20:05:34] FIRING: DiskSpace: Disk space prometheus1008:9100:/srv/prometheus/k8s-dse 3.953% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=prometheus1008 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [21:05:34] RESOLVED: DiskSpace: Disk space prometheus1008:9100:/srv/prometheus/k8s-dse 3.718% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=prometheus1008 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [23:06:35] FIRING: DiskSpace: Disk space prometheus1008:9100:/srv/prometheus/k8s-dse 3.54% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=prometheus1008 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [23:11:34] RESOLVED: DiskSpace: Disk space prometheus1008:9100:/srv/prometheus/k8s-dse 3.54% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&viewPanel=12&var-server=prometheus1008 - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace