[01:20:25] FIRING: [8x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2192:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:19:36] I've brought up misc masters in codfw and all proxies reloaded [05:20:25] FIRING: [8x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2192:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:25:25] FIRING: [8x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2192:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:30:25] FIRING: [8x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2192:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:26:42] If you need another pair of hands for anything, just ping me [07:27:34] thanks Emperor , so far I think it is all fine now, I am on _security working out the alerts/notification status with volan.s and tappo.f [07:27:55] cool [07:30:38] the apus cluster in codfw lost power to 2/4 backends and all the frontends, and managed to put itself back to a functioning state without intervention, which I'm quite happy/relieved about (I had to sort out an unhappy monitor, but I think that was probably correct quorum-related refusal to start) [09:14:07] Emperor: please could you check what you are seeing from ms-backup200[34] against Swift? [09:15:23] I see the queue building from the commonswiki API, but the processing just seems to be stuck [09:21:28] I have stopped the sync processes for the moment to try to find out why the log seems to just stall [09:30:25] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2202:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:57:36] cezmunsta: I picked one frontend (ms-fe2009) and one of those servers (ms-backup2003) and I see some requests that look to be going OK e.g. 2026-09-24T07:29:59.400395+00:00 ms-fe2009 proxy-server: 10.192.8.7 10.192.8.7 24/Sep/2026/07/29/59 GET /v1/AUTH_mw/wikipedia-commons-local-public.ee/e/ee/1062362_I_Woolsthorpe_Manor_House%252C_interior_Colsterworth_20260605_0063.jpg HTTP/1.0 200 - python-swiftclient-4.7.0 AUTH_tkde73b7afa... - [09:57:36] 32231059 - txd8a5bf02779d4cc58d663-006ab4d177 - 0.1839 - - 1790234999.215436697 1790234999.399361372 0 [09:58:35] cezmunsta: but the last download from that fe/backup pair was at 08:00 UTC this morning, after which there were auth requests (successful) every half hour or so most recently at 09:11:40 [09:59:45] [but maybe because that's when you stopped them running?] [10:00:47] cezmunsta: I'm inclined to suggest starting at least one up again and seeing what it logs about its interactions with swift... [10:01:21] Emperor: thanks, yes that would match. I was seeing lots of duplicates and so I just wanted to make sure that nothing odd was happening in terms of where the requests were going. It looks like the duplicates are ones that have been processed in eqiad, so I presume that they are ones that have copied across [10:01:57] codfw swift looks to be receiving almost 0 requests since the incident though [10:02:00] which is odd [10:03:22] I will let you know once started again, I am just trying to see why it looks to stall given that there is a queue to process :) [10:19:05] Emperor: OK, I have found what was going on, an unhandled exception in the code, so checking the targets now [10:22:49] cezmunsta: ah, progress :) [I have a meeting at half past] [10:30:07] It looks like these issues started when the database was accessible again: Sep 23 19:24:17 ms-backup2003 backup-wiki[5709]: mediabackups.MySQLMetadata.MySQLConnectionError -> Sep 23 19:54:24 ms-backup2003 backup-wiki[7745]: boto3.exceptions.S3UploadFailedError: Failed to upload , which is the interval between auto-restarts [10:45:24] looks like the hosts themselves were back online around 18:24ish yesterday. [10:46:11] Anyhow, have things started back up OK? [10:58:27] For now, as I disabled Puppet again the other day, I have dropped the restart time to 30s. I still see exceptions, but then there are also log entries saying "complete successfully". For the moment, I am just seeing how it goes - but I don't see the backedup count increasing, but I am wondering if that is the code vs a backup. I cleared the log earlier so that it will be easier t look back and reset things if need be [10:59:21] Meanwhile, back to checking the code to see if that looks to be a valid explanation [11:51:02] OK, adding in handling for this is getting the backedup counts rising, so I will apply the same to the other instance and start that up again [13:30:25] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2202:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:35:13] So the codfw processes are catching up on the pending queue, I will check what to do with the ones that had issues tomorrow once the backlog has gone [19:35:25] RESOLVED: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2202:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed