[02:54:28] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:08:03] FIRING: [2x] PuppetFailure: Puppet has failed on ms-be2097:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [03:53:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:58:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:23:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:28:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:41:20] I am OOO but the depool cookboko doesn't work [05:42:57] this needs fixing asap https://phabricator.wikimedia.org/T433564 [07:08:04] FIRING: [2x] PuppetFailure: Puppet has failed on ms-be2097:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [07:19:25] marostegui: "2026-07-29 14:02:25 - 2026-07-29 14:49:11" was when db1218 was repooling yesterday after its restar for the kernel [07:27:56] cezmunsta: Interesting...could be a coincidence. We've had many hosts already crashing like this (rebooting, nothing on logs) in the last few weeks [07:28:32] I going to send the list to papaul and willy to see if we can get dell to troubleshoot this at least having all these issues logged [07:29:55] +1 looking at the depool cookbook issue shortly - whilst you are around, there is a hanging wikiadmin connection on commonswiki, I haven't spotted an actual query, but it remains connected and blocks the restart... who is best to ask about the remote end? [07:30:20] what's the host? [07:30:33] db1252 [07:30:59] It has done it a few times, there were 2 connections at one point and then just the one [07:31:13] there are also wikiuser ones [07:31:26] Yes, when that script fails it repools [07:31:34] I wanted to go ahead and kill the script but I ran into: https://phabricator.wikimedia.org/T428573 [07:33:09] I will exclude it for now and maybe they drop off of their own later [07:33:15] it is depooled? [07:34:21] Now, no - timeout on depool = repool and exit [07:34:45] I will depool with dbctl [07:35:51] depooled, I'd say, let's stop mariadb and reboot (manually if you want) [07:36:26] do you want me to do it? [07:36:48] No, I will give it a short while and then start the run again [07:37:00] cool, I've NOT downtimed it or anything [07:37:03] just depooled via dbctl [07:37:10] Is there anything left to do re the crashed host? [07:37:47] I can get it replicating again if you weren't going to do anything beforehand [07:38:50] cezmunsta: I'd leave it stopped, papaul would get someone to update bios and firmware, so they'll have to reboot it a couple of times [07:38:57] +1 [07:39:00] I'd wait for DCOPs to confirm on the task they are done, and then we can start it again [07:39:15] I am going to go ooo again but I will get the list of hosts sent later today [07:39:29] enjoy o/ [07:39:33] It will be too hot outside anyways to do anything XD [07:39:34] o/ [07:41:09] thanks for the help! [09:28:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:33:53] I lost db1245 again :-( [10:33:58] FIRING: [3x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db1245:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:38:58] FIRING: [5x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db1245:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:50:21] Sigh, past me is an idiot [10:51:52] cezmunsta: just wondering if the depool cookbook is broken how come rolling restarts scripts are working? [10:57:39] Can I get a +1 to https://gerrit.wikimedia.org/r/c/operations/puppet/+/1319442 so I can unblock the trixie migration by (I hope!) now being able to reimage ms-be1080, please? [10:59:59] * cezmunsta looks [11:02:15] marostegui: it is not broken, so long as you can connect to the host [11:02:45] As for rolling restarts, they don't depool using the cookbook... [11:06:29] I used the depool cookbook with db1252 to make sure that it does work: 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance [11:06:58] cezmunsta: thanks :) [11:08:04] FIRING: [2x] PuppetFailure: Puppet has failed on ms-be2097:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [11:09:38] ^-- that's the pair of new nodes that didn't image properly, I will try and get to them a bit later, been working on trying to unbreak ms-be1080 so the trixie upgrade can restart [11:10:42] [I'll silence them for a day in the mean time] [11:17:47] cezmunsta: I think I reported that bug some weeks ago and I think it got fixed but looks like it is not [11:17:52] Let me see if I can find the task [11:17:59] I'm on my phone so we'll see xdd [11:19:50] cezmunsta: https://phabricator.wikimedia.org/T427381 [11:20:57] ack - based upon timestamps I would guess that it was lost in the refactor for depool.py [11:21:09] I will check what was done anyway [11:22:27] Thank you so much for working on [11:22:34] *on this [11:23:23] np - confirmed, test_runner_s_depool no longer exists in the unit test [11:25:36] Added it to the triage notes for Monday [11:49:04] FIRING: MysqlPredictiveFreeDiskSpace: Host db1265:9100 predictive low disk space on /srv - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [11:56:01] jynus: ^ I see you on the host, there seems to be plenty of space there, did you do something, or a false +ve? [11:57:30] it is being setup, I am against having such alerts preciselly for that [11:57:40] but still, someone set it up [11:57:49] setup == provisioning [11:58:55] ack ty [11:59:20] https://phabricator.wikimedia.org/T407942#12170792 [12:01:10] predictive alerts are like precog crime, presumption that something will go bad; they should be dashboards, not alerts [12:03:58] FIRING: [6x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db1245:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:09:04] RESOLVED: MysqlPredictiveFreeDiskSpace: Host db1265:9100 predictive low disk space on /srv - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [12:13:58] FIRING: [3x] SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db1245:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:15:25] I think all mine are acked or resolved now, or at least I cannot find them on alertmanager [12:57:04] FIRING: MysqlPredictiveFreeDiskSpace: Host db1265:9100 predictive low disk space on /srv - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [13:02:04] FIRING: [2x] MysqlPredictiveFreeDiskSpace: Host db1265:9100 predictive low disk space on /srv - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [13:17:04] FIRING: [2x] MysqlPredictiveFreeDiskSpace: Host db1265:9100 predictive low disk space on /srv - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [13:23:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:28:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:50:32] db1265 and db1285 instances are recovering from lag and when done, I will add them to zarcillo [13:51:07] please note that if you ran any schema change between 20 hours yesterday and today, they will need reaplication (s3, s4, s7 or s8) [13:53:09] * cezmunsta nods [13:53:49] jynus: is the ifup one something that you have seen before on other hosts? [13:53:56] s/one/alert [13:53:57] no, I was about to ask [13:54:10] my first intention would be to 1) depool just in case [13:54:14] It was a failed IPv6 as already assigned [13:54:21] manually restart the service or the host [13:54:26] mmmm [13:54:29] https://phabricator.wikimedia.org/T433610 [13:54:34] interesting, I would depool and contact netops [13:55:01] I noticed that there is a timer for the VMs to reset these alerts [13:55:06] like, discard it is a one time fluke [13:55:21] but otherwise if there is an ip conflict I woudl consider a huge issue [13:55:24] host-wise [13:55:59] enough to depool if there are enough resources [13:56:12] It ha replicas [13:56:20] *has [13:56:35] is it a master or what it is its role? [13:56:55] Due to move to x4 [13:57:52] I don't know how that's handled, but a warning even if no current ongoing issue would make me suspicious and I would make it depoolable asap just in case [13:58:29] For the VMs, which get the reset-failed, this is the note: "# A bandaid for ifupdown race on boot" [13:58:31] if you are on your own, it can wait [13:58:52] yeah, but for the vms it would make sense [13:58:59] I've never seen it on a real host [13:59:11] I would consult netops, maybe they say: don't worry [13:59:28] Yep, will do [13:59:33] ^ topranks any insight? know issue? [13:59:42] or XioNoX ^ [13:59:57] what? [14:00:13] https://phabricator.wikimedia.org/T433610 re the SystemdUnitFailed for ifup [14:00:37] In particular the "Error: ipv6: address already assigned." [14:00:54] fluke or real issue in your opinion [14:01:07] did it happen more than once? [14:01:23] do you know, cezmunsta? [14:01:52] is it safe to reload without causing network interruption? [14:02:07] host seems fine at the moment [14:02:08] it looks like just poor interaction between the "up" command, ifupdown and systemd [14:02:11] I don't think so, was marked failed on 28/07 and seems to have been making noise since [14:02:25] all of it probably needs a little tlc [14:02:42] I believe what's happening is if the service restarts it tries to run this command as per /etc/network/interfaces [14:02:43] so safe to reload, we were wondering if to depool just in case [14:02:50] that is a question ^ [14:02:51] up ip addr add 2620:0:860:12d:10:192:58:6/64 dev eno12399np0 [14:03:17] because that's been done already iproute2 returns an error code to the script [14:03:21] cmooney@db2248:~$ sudo ip addr add 2620:0:860:12d:10:192:58:6/64 dev eno12399np0 [14:03:21] Error: ipv6: address already assigned. [14:04:04] there is nothing wrong here the system is fine. the problem is the framework around restarting the systemd service / ifup which ends up in a "failed" state becuase of that [14:04:20] jynus: yeah it is safe to reload [14:04:21] I see, so a bug in the logic, not on the host is what I am understanding [14:04:29] to actually make systemd happy we probably need to reboot the host [14:04:33] minor bug [14:04:36] gotcha [14:04:57] so we can ignore it/ack it and schedule for a reboot, cezmunsta [14:05:04] +1 [14:05:08] but not in an emergency way [14:05:14] thanks for the context, topranks [14:05:17] and XioNoX [14:05:29] I think that it appeared because of a reboot :) [14:05:52] cezmunsta: can you at least document a summart of that on the ticket ? [14:06:08] cezmunsta: hmm maybe there is something deeper going on [14:06:12] oh [14:06:32] "reboot system boot 6.12.95+deb13-am Tue Jul 28 14:59 - still running" [14:06:36] topranks: the reasoning to ping you was if you could have some extra debuging tools [14:06:41] the failure from the task is because of the service restarting [14:06:56] it failed trying to re-add the same IP to the int that was already on it [14:07:04] RESOLVED: MysqlPredictiveFreeDiskSpace: Host db1265:9100 predictive low disk space on /srv - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [14:07:17] at boot the IP isn't there, so that command should run ok, the service should start [14:07:28] race condition? [14:07:33] the question is therefore what made the service try to restart after the first time [14:07:45] perhaps there is some other problem causing systemd to try that [14:07:48] jynus: yeah [14:08:08] it is bad if there is a low chance of it failing on every reboot [14:08:47] seems odd though, we have a vanilla /etc/network/interfaces here, and we've thousands of hosts like that and not seen this before [14:08:54] what if depooled, then manual ifdown, the restart? [14:08:56] is the host pooled? [14:09:03] yeah, it is active right now [14:09:15] that is why we seeked to see how urgent it was [14:10:10] so back to my original plan of non-emergency depooling and then debug after depool [14:10:55] jynus: yeah I think that seems best [14:11:21] looking at the logs it does seem to have failed first time, so perhaps there is some race condition we've not seen before [14:11:49] that's good, finding bugs even if extremelly rare should be celebrated [14:11:59] cezmunsta: if you are on your own today, it can wait [14:12:01] IMHO [14:12:28] but let's document a summary of the discussion here on the ticket and proposed depooling and debug [14:12:56] as the host seemed healthy last time I looked at it [14:13:16] and if I can count with you topranks for when it is depooled for further debugging [14:13:33] jynus: yep happy to thanks [14:13:50] cezmunsta: ok to proceed like that? [14:14:56] also a good moment to review how to depool manually a host, as I saw another ticket from M. where the cookbook was broken [14:15:23] jynus: not broken, it was due to the host having crashed [14:15:46] ah, yeah, dependency on the host being up, right [14:15:53] that's good [14:16:43] but yeah, the idea was "prepare for the worst", I don't think it will be the case, but in the past e.g. a disk broken, a host showed signs of degradation before a hard down [14:17:05] The ticket for x4 creation is "Setting up a fake x4 section " Amir1 any input on the codfw section of x4? [14:17:36] yeah, I don't know the progress of that, maybe it is not in real production, Amir will know [14:18:14] but if it is, it will be normal section, so no ring or shard, normal sX section-like [14:18:36] but I get confused with the many xY sections now [14:18:38] x4 is for now should be just a couple of s4 hosts pretending to be x4. On the db-side everything is s4 [14:18:55] We are mostly ready for the split though. zabe would know the progress of that [14:18:57] Amir1: but does it receive already real mw traffic? [14:18:58] FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:19:09] Amir1: ^ the reason for asking :) [14:19:26] the host should be pooled under s4 already. Isn't it? [14:19:36] so it's pooled under s4 and x4 [14:19:54] with the difference being x4 is not getting any [14:20:00] I see, so it would need depool from both, and I guess a master switch too [14:20:08] yeah [14:20:38] they connect to the same port, same mariadb instance. It's just pretending to be two different things (until we split them apart) [14:20:46] gotcha [14:21:32] cezmunsta: so review the procedures for a secondary master switchover in case there was an emergency, but let's wait when there is more people around to proceed [14:22:30] imho [14:24:25] I am going to see what the switchmaster UI makes of x4 [14:24:50] Yep, that broke :D [14:28:06] It seems happy enough - I have added a silence for it for the time being [14:36:57] ty for input all [14:47:58] going away for the day [14:48:10] o/ enjoy [14:49:52] cezmunsta: need help? [14:51:23] marostegui: maybe take a quick look at https://phabricator.wikimedia.org/T433610 and then let me know what you think? [14:53:36] cezmunsta: first time I see that error. For now you can just leave it, I'll switch it on Monday. That section is sort of WIP so our tooling won't work on it as you mention on the ticket [14:53:54] I'll switch it to a different host on Monday and we can reboot/investigate [14:54:21] kk thanks for the eyes [14:54:43] Busy day today I see! [14:55:14] Mostly that noisy alert :D