[10:35:43] marostegui: do we really need to see the queries that are running when waiting for depool? i.e. use desired columns instead of SELECT * [10:36:10] i don't normally check, unless a host is taking a bit to get depooled and then I do check [10:36:20] that's my use case, not sure about what others do [10:38:06] I have a screen full of characters that I cannot read as part of the rolling restarts :) [10:38:13] XD [10:38:24] I usually check only if it's taking a long time and sometimes eyeball it if it prints out something unexpected [10:38:24] I guess we can just delete it [10:38:27] federico3: thoughts ^ [10:38:28] ah [10:38:47] should we maybe just print it after a few mins or attempts instead of all the time? [10:39:09] yes, works for me [10:39:31] also I configure byobu with a long backlog just in case [10:39:45] cezmunsta: want to give that patch a try? [10:40:50] To avoid the noise, sure :) [10:41:07] haha [10:42:30] "START - Cookbook sre.mysql.pool pool db1197: Maintenance" this is showing up as notifications disabled, yet seems to have been marked by the rolling restart for a repool - does that sounds correct? [10:44:26] https://phabricator.wikimedia.org/P95505 [10:48:32] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1279290 [10:48:33] wow [10:48:37] Fixing [10:49:15] +1 [10:49:45] merging [10:50:00] how did this host go through in the last rolling restarts then? [10:50:35] I am using the cookbook for repool, which allows checking of Icinga. The restarts previously only checked replication lag [10:50:42] ah right [10:51:43] This is the only one that I have seen so far and most runs have used the cookbook. I will keep an eye out for any others though, should they appear [10:52:37] k thanks [10:53:16] I will wait to see how it is handled by the code, hopefully self-healing :) [11:16:51] It recovered itself and started repooling [12:54:58] cezmunsta: so this can be closed right? https://phabricator.wikimedia.org/T433493 [12:55:57] Yep, missed that one getting generated [12:56:03] no worries! [12:56:22] closing it [12:56:24] done [12:56:28] ah thanks! [13:35:22] federico3: how are you avoiding the warnings coming from ".tox/py311-lint_unit/lib/python3.11/site-packages/phabricator/__init__.py" in cookbook tests? [13:37:14] ... or do you know if anything is being done to resolve them? [14:44:06] for the dbas: db1245 may resurect today: https://phabricator.wikimedia.org/T431115 [14:44:34] It was marked as notifications disabled, but as it never ran puppet, I cannot discard a race condition or something if it comes back [14:44:42] I think it is marked as failed on netbox anyway [14:44:55] just FYI [14:58:58] FIRING: SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:00:32] that's x4, weird unit failure [15:02:15] the host seems fine, though [17:37:48] FIRING: PuppetFailure: Puppet has failed on ms-be2098:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [18:59:28] FIRING: SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:07:48] FIRING: [2x] PuppetFailure: Puppet has failed on ms-be2097:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [22:13:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:53:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:54:28] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:08:03] FIRING: [2x] PuppetFailure: Puppet has failed on ms-be2097:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure