[06:50:12] we have a PSU issue on db2226, I'm opening a task for dc-ops [07:12:39] federico3: how did you find that? [07:13:15] cezmunsta: alertmanager - https://alerts.wikimedia.org/?q=%40cluster%3Dwikimedia.org&q=instance%3D~%5E(db%7Cpc%7Ces)%5B12%5D.* [07:17:40] Sorry, I was meaning what drew your attention to it, I don't seem to see a notification anywhere, nor email [07:21:35] I just scan the alertmanager page [07:22:01] cezmunsta: have you also seen the crashloop alerts for backup200{3,4}? [07:22:48] you can filter by team on alerts.wm.org (a thing I try and remember to do) [07:22:57] Emperor: yes [07:26:29] I think that check needs adjusting - a crash lopp != exit code 0 [07:26:43] s/pp/op/ [07:26:55] Main PID: 315774 (code=exited, status=0/SUCCESS) [07:31:23] They had the 1800s dropped to 30s to catch-up, which they did and so then started stopping more frequently. I have stopped them completely for the moment so that I can try to figure out which need to go back to the queue from the errors yesterday [07:32:07] ack. [07:32:21] [this is my last day for 2 weeks, so do shout if you need anything from me] [07:33:38] Thanks, will do ... enjoy your 2w off! [09:51:53] marostegui: OK, test-s4 just did the full circuit of prepare->finalize->prepare->finalize - given that the prod rerun is now a few weeks away, happy for the merge? [09:52:20] Yeah go for it! I 1+ed already right? [10:14:14] Yep, just rebasing then I will merge it [21:10:25] FIRING: [2x] SystemdUnitFailed: mariadb.service on db2160:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed