[00:54:24] FIRING: SystemdUnitFailed: fstrim.service on sessionstore1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:54:39] FIRING: SystemdUnitFailed: fstrim.service on sessionstore1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:16:14] those errors are related to sda failing [08:54:25] FIRING: [2x] SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:52:58] federico3: where's the threashold for the depooled hosts defined? [10:53:01] threshold [10:53:46] you mean in terms of number of hosts? [10:54:20] No, like, where do we define how many days need to past for a host to start alerting [10:56:52] in the "for:" key in mysql-depooled.yaml - right now it's instant as long as the silence expired or is removed, but if we find it too noisy we can put some days [10:57:51] I see, thank you [11:10:52] I've updated the patch for updating the CNAME for the DV masters to codfw. S5 moved: https://gerrit.wikimedia.org/r/c/operations/dns/+/1319815 [11:11:07] Are we all set on the database side? [11:11:40] slyngs: no, we have to enable replication from codfw -> eqiad https://phabricator.wikimedia.org/T436500 [11:11:43] federico3 cezmunsta ^ [11:12:19] I'll just subscribe to that task then. Thanks [11:12:26] yes, scheduled for tomorrow [11:12:28] slyngs: Btw I think federico rebased the patch a few days ago [11:12:34] indeed [11:12:37] the dns one I mean [11:13:16] Nice, thank you. Still S5 was wrong :-) [11:13:42] Down from me originally being wrong on like 8 hosts [11:56:47] @slyngs did you use a tool to generate the PR? [11:57:22] ... no and yes, if https://noc.wikimedia.org/dbconfig/codfw.json is a tool... Otherwise the only tool is me :-) [12:54:40] FIRING: [2x] SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:24:25] FIRING: [2x] SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:26:40] ^-- I did just reset-failed the rclone service [15:24:25] RESOLVED: SystemdUnitFailed: fstrim.service on sessionstore1005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:25:11] ^ that host came back? [15:25:12] Emperor: ^ [15:26:16] [17:24:59] <+icinga-wm> PROBLEM - Host sessionstore1005 is DOWN: PING CRITICAL - Packet loss = 100% [15:26:19] and back down again XD [15:31:47] marostegui: dc-ops are working on it [15:31:55] ah ok thanks! [16:17:09] hello! We have a low priority change for orchestrator that we'd like review on. We'd like to migrate its check from icinga/tcp to prometheus/http and that requires changing the check type a little bit. https://gerrit.wikimedia.org/r/c/operations/puppet/+/1343590 [16:18:15] hnowlan: can we delay that for Thursday, once the DC switch is done? [16:20:13] absolutely, zero rush [16:20:26] it can be next week or whenever [16:30:08] hnowlan: Thanks, i've added myself and c.ezmunsta to the reviewers list, so we have it "pending" and get to it after the switchover [16:30:23] thanks! <3