[06:47:45] greetings [07:19:10] I'll take a look at tools-k8s-worker-nfs-74 [07:31:17] morning! [07:31:22] godog: thanks! [07:36:48] is there a task for the latest network issues? T432426 was closed long ago [07:36:48] T432426: Network unavailable on a few VMs - https://phabricator.wikimedia.org/T432426 [07:38:22] so the -74 is basically unresponsive on console, not sure how else to debug further [07:38:32] I'll reboot unless someone has ideas [07:40:08] that's new, anything on horizon logs? [07:41:08] good point, checking [07:43:15] bunch of oom messages, and indeed it looks like the vm is trashing https://phabricator.wikimedia.org/P96485 [07:44:41] ok I'll reboot [07:46:17] thanks [07:48:33] sure np, re: latest network issues I'm not sure tbh, there was yesterday the neutron agent restart and that seemed isolated, anything else ? [07:49:30] it has happened several times in the last few months that the nodes just lost their network (rebooted at least a couple workers lately, besides yesterday) [07:49:54] I was just curious if there was any known reason [07:50:49] ah, not as far as I'm aware other than T432426 [07:50:50] T432426: Network unavailable on a few VMs - https://phabricator.wikimedia.org/T432426 [07:51:13] I have this old patch https://gerrit.wikimedia.org/r/c/cloud/wmcs-cookbooks/+/1195629 to allow skipping the drain for rebooting k8s workers, any reviewers? Otherwise I'll abandon to clear my in-progress things [07:52:44] this also https://gitlab.wikimedia.org/toolforge-repos/admin-web/-/merge_requests/1, it's been there for some time [07:52:46] (the stack) [07:53:23] I pasted the wrong log earlier, this is the correct one though nothing changes of substance https://phabricator.wikimedia.org/P96486 [07:54:31] that sounds like user workload (webservices) being killed by OOM yep [07:57:17] infra tracing has been failing to connect on that node as long as there's logs (16th Sept), with ConnectTimeout [07:57:29] Sep 16 17:36:15 tools-k8s-worker-nfs-74 infra-tracing-nfs[3826149]: requests.exceptions.ConnectTimeout: HTTPSConnectionPool(host='localhost', port=30004): Max retries exceeded with url: /loki/api/v1/push (Caused by ConnectTimeoutError(, 'Connection to localhost timed out. (connect> [07:58:55] hmm... this log has always been around when I've seen network issues like this one (and OOM) [07:58:56] Sep 22 06:04:01 tools-k8s-worker-nfs-74 sssd[593]: Child [456473] ('wikimedia.org':'%BE_wikimedia.org') was terminated by own WATCHDOG. Consult corresponding logs to figure out the reason. [07:59:18] kinda makes sense though, it can't connect to ldap, but the error is not very clear [07:59:44] Sep 22 05:57:03 tools-k8s-worker-nfs-74 systemd[1]: Starting restart_networkd_on_network_failure.service - Restart systemd-networkd if no routes are found... [07:59:49] ^that sounds interesting [08:01:57] heh, dcaro if you have time to dig deep today and want to please go ahead, I currently have too much stuff on my plate for today [08:02:05] that happens all the time though xd [08:03:36] having too much stuff going on and/or restart networkd ? [08:07:31] both to be honest xd [08:07:57] lol [08:10:56] hmm `plugin type=\"calico\" failed (delete): netplugin failed with no error message: signal: killed"` at around the time when the network seems to start breaking [08:11:20] this same line `Sep 22 05:54:48 tools-k8s-worker-nfs-74 containerd[594]: time="2026-09-22T05:54:48.661055234Z" level=error msg="RemovePodSandbox for \"91d6f7ce2824b7a252b48f7306cd3bf60c35a5dfeb1e0905ca4bdab397eee35a\" failed" error="failed to forcibly stop sandbox \"91d6f7ce2824b7a252b48f7306cd3bf60c35a5dfeb1e0905ca4bdab397eee35a\": failed to destroy netwo>` [08:13:04] dmesg does not have logs after 05:50 it seems :/ [08:13:26] that's weird [08:13:42] anyhow, too much other stuff to get through, I'll leave it there [08:52:43] morning [12:40:32] taavi: for gerrit 1343962 doesn't this PCC looks a bit weird to you? https://puppet-compiler.wmflabs.org/output/1343962/9466/conf1007.eqiad.wmnet/index.html [12:42:38] volans: how so? it moves the ACL definitions from the manifest to hiera so that we can use the profile on cloudinfra-etcd without that ACL (while still keeping it on the conf* nodes) [12:44:04] doh, I'm an idiot, for me there was only one of the two, the line ended perfectly at the right browser border so I didn't horizontally scroll (and there is no evident scrollbar) :D [12:44:26] so only wikireplica-db-analytics was added and not the other [12:44:28] lol [12:47:44] taavi: one small note, as the DC switchover is around the corner, consider if postponing merging it until after, maybe check with sly.ngs [16:58:50] * dhinus off [17:26:45] * dcaro off