[11:55:25] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@x1.service on db2201:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:58:34] ^ this is the host with HW issues: https://phabricator.wikimedia.org/T434532 [11:58:44] I am going to silence that for a month [11:58:55] it is a backup source [12:15:25] FIRING: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:21:02] @marostegui we should plan to restart dborch for https://phabricator.wikimedia.org/T431658 - I can do it now [12:21:16] federico3: please go ahead! thank you [12:23:55] done [12:25:25] RESOLVED: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:37:14] thank you federico3 [14:00:25] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db2201:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:04:20] Hi All, I am sure DP team is independent and can handle things during Kwaku's absence. Just in case if you need any help during his absence please feel free to reach out to me. [15:06:43] Thank you kavitha_ [17:32:04] hi dp! traffic & serviceops are getting started on dc switchover prep [0] and wanted to reach out to ask if dp has a dc switchover contact we could coordinate with on pre-dc switchover tasks (e.g [1]) [17:32:04] [0] - https://phabricator.wikimedia.org/T433363 [17:32:04] [1] - https://phabricator.wikimedia.org/T403966 [17:34:12] [non-urgent] [cc slyngs] [17:41:43] & from more recently, organized as pre-work > circular replication > post-work [17:41:43] [2] https://phabricator.wikimedia.org/T416705 [17:41:43] [3]https://phabricator.wikimedia.org/T416706 [17:41:43] [4] https://phabricator.wikimedia.org/T416708 [18:00:40] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db2201:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:00:40] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db2201:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed