[01:21:40] FIRING: [5x] SystemdUnitFailed: uwsgi-graphite-web.service on graphite1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:52:40] FIRING: [2x] LogstashKafkaConsumerLag: Too many messages in logging-eqiad for group logstash7-codfw - https://wikitech.wikimedia.org/wiki/Logstash#Kafka_consumer_lag - https://alerts.wikimedia.org/?q=alertname%3DLogstashKafkaConsumerLag [03:25:07] FIRING: SlothSliErrorRatioRateAbsent: SLI error ratio rate is absent (wdqs / scholarly-update-lag) - TODO - https://grafana.wikimedia.org/d/slot-pilot-slo-detail/sloth-slo-detail?var-sli_window=5m&var-service=wdqs&var-slo=scholarly-update-lag - https://alerts.wikimedia.org/?q=alertname%3DSlothSliErrorRatioRateAbsent [03:42:40] RESOLVED: [2x] LogstashKafkaConsumerLag: Too many messages in logging-eqiad for group logstash7-codfw - https://wikitech.wikimedia.org/wiki/Logstash#Kafka_consumer_lag - https://alerts.wikimedia.org/?q=alertname%3DLogstashKafkaConsumerLag [05:20:07] FIRING: [2x] SlothSliErrorRatioRateAbsent: SLI error ratio rate is absent (wdqs / scholarly-update-lag) - TODO - https://alerts.wikimedia.org/?q=alertname%3DSlothSliErrorRatioRateAbsent [05:24:30] FIRING: [5x] SystemdUnitFailed: uwsgi-graphite-web.service on graphite1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:36:26] hello o11y, let me know what you think of https://phabricator.wikimedia.org/T439099 [09:20:07] FIRING: [2x] SlothSliErrorRatioRateAbsent: SLI error ratio rate is absent (wdqs / scholarly-update-lag) - TODO - https://alerts.wikimedia.org/?q=alertname%3DSlothSliErrorRatioRateAbsent [09:24:30] FIRING: [5x] SystemdUnitFailed: uwsgi-graphite-web.service on graphite1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:42:19] XioNoX: should be fairly doable, thanks [09:43:09] <3 [13:20:07] FIRING: [2x] SlothSliErrorRatioRateAbsent: SLI error ratio rate is absent (wdqs / scholarly-update-lag) - TODO - https://alerts.wikimedia.org/?q=alertname%3DSlothSliErrorRatioRateAbsent [13:24:30] FIRING: [5x] SystemdUnitFailed: uwsgi-graphite-web.service on graphite1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:26:40] RESOLVED: SystemdUnitFailed: ifup@eno12399np0.service on logging-sd2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:55:07] RESOLVED: [2x] SlothSliErrorRatioRateAbsent: SLI error ratio rate is absent (wdqs / scholarly-update-lag) - TODO - https://alerts.wikimedia.org/?q=alertname%3DSlothSliErrorRatioRateAbsent [16:49:30] FIRING: SystemdUnitFailed: prometheus-redis-exporter@6379.service on arclamp2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:51:40] RESOLVED: SystemdUnitFailed: prometheus-redis-exporter@6379.service on arclamp2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:13:44] hi observability friends -- is thanos looking healthy in codfw? should it be repooled? (if so, please feel free to do so!) [18:18:01] thanks cdanis, ok giving it another look over and then will do so [18:24:01] confirming thanos codfw is pooled [18:24:48] i see the thanos-(query|swift|web) discovery services still depooled in codfw..? [18:25:17] where do you see that? [18:25:57] sudo cookbook sre.discovery.datacenter status all [18:26:14] iiinteresting [18:54:27] my bad, was not looking at dnsdisc, those are pooled now