[06:17:38] greetings [06:57:52] morning [07:36:42] godog: there's another toolforge worker (in tools) that lost the network :/ [07:38:21] tools-k8s-worker-nfs-39, I'm checking, I think the journal is old enough [07:38:26] do we have a task for those? [07:44:52] mhh I wonder if that's the same problem as T432426 [07:44:52] T432426: Network unavailable on a few VMs - https://phabricator.wikimedia.org/T432426 [07:46:59] cloudvirt1077 is me btw [07:58:56] ack, [08:00:48] just found on the tools bastion, it has been having segmentation failures since the 15th [08:17:48] morning [08:20:24] godog: volans the infra tracing scripts are failing on the bastion, do you want me to open a task? [08:20:29] Sep 21 08:19:49 tools-bastion-15 infra-tracing-nfs[3949]: Unable to find kernel headers. Try rebuilding kernel with CONFIG_IKHEADERS=m (module) or installing the kernel development package for your running kernel version. [08:20:56] dcaro: interesting, did they got a new kernel and rebooted recently? [08:21:15] I just rebooted yep, did not upgrade manually, but might have happened with autoupgrades I think [08:21:40] checking if the deps are up to date too [08:22:06] in particular the linux-headers- [08:24:12] dcaro: I see it working fine now [08:24:13] Sep 21 08:23:32 tools-bastion-15 infra-tracing-nfs[4104]: 2026-09-21 08:23:32,682 INFO: Sent 13 log lines across 3 streams to Loki [08:24:30] oh, okok, maybe puppet upgraded it [08:24:59] yes at 08:19:31 [08:25:01] the machine seemed to be stuck trying to read the last puppet run file (processes got stuck in D state when trying, blocking some logings that tried to update the MOTD) [08:25:12] so maybe puppet was stuck too :/ [08:25:30] that file is not on NFS, so it's a bit troubling [08:27:26] so my understanding is: the kernel was updated, the headers were not installed, the host was rebooted, nfs-tracing couldn't start, the first puppet run installed the headers and refreshed infra-tracing, so everything seems to work fine after the first puppet run [08:28:03] ideally the unattended upgrades should install the headers too if were installed with the previous kernel when upgrading it, not sure if there is an option for it [08:28:51] that sounds correct [08:29:06] morning. something's up with metricsinfra-puppetserver-1. it fell off the network earlier today, I tried hard-rebooting it via horizon and now it's failing to bring up its interfaces again. seems to becaused by a live migration, based on the horizon action log? [08:29:20] with the suspicion that puppet got stuck in between and did not upgrade the packages automatically until the reboot [08:29:41] taavi: there's two toolforge workers without network today too [08:29:51] (not sure it's related, but just for visibility) [08:30:56] taavi: I did drain 1077 earlier today in T435537 and 1063 maybe half an hour ago [08:30:57] T435537: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537 [08:31:44] volans: P:wmcs::kubeadm::infra_tracing::nfs seems to only install the headers package for the current kernel version, I think if it would install the 'linux-headers-cloud-amd64' metapackage as well it would automatically get updated like the kernel itself [08:31:56] dcaro: which ones? [08:32:23] nfs-39 and 70 [08:32:44] they have IPs, but can't reach the dns [08:33:35] nfs-39 was also live migrated this morning, -70 was not [08:35:17] sorry, none of them had ips, I did a `systemctl restart network` on -39, and it got an ip, but is unable to reach the dns anyhow [08:36:25] taavi: indeed, I don't recall exactly why I did it that way, I vaguely recall I was pointed at some existing code to get it this way, but we can totally revisit and use linux-headers-cloud-amd64 instead [08:38:29] the tracing process starting timing out to send logs on `Sep 19 02:17:27` for nfs-70 [08:39:02] *started [08:39:44] sssd crashed also by that time [08:41:29] a bit earlier `Sep 19 00:51:45 tools-k8s-worker-nfs-70 sssd[594]: Child [3874542] ('wikimedia.org':'%BE_wikimedia.org') was terminated by own WATCHDOG` [08:41:48] also the previous day though, so maybe unrelated [09:04:51] I'll take a look at metricsinfra puppetserver-1 and mx-out05 unreachable unless someone is looking already [09:32:52] ok what I got so far: I'm not seeing dhcp requests even making it to cloudnet hosts, whereas other hosts on 1073 like pontoon-demo-cloudgw-01.testlabs seem to be fine, I'll open a task [09:37:00] T438697 [09:37:01] T438697: Instances migrated to cloudvirt1073 lost their network following a live migration - https://phabricator.wikimedia.org/T438697 [09:40:16] mmhh those instances are on VLAN/legacy and 1073 was part of the rack move/rebalance, I'm wondering if we're missing sth at the network level [09:41:21] taavi: maybe you know or ^ rings a bell? I'm not familiar with VLAN/legacy is supposed to work / be transported from cloudvirt to cloudnet [09:42:07] let me check [09:43:03] thank you [09:44:04] the VLAN is there on the switch interface at least https://phabricator.wikimedia.org/P96479 [09:45:45] indeed [09:48:22] for reference cloudvirt1073 moved from F4 to D5 in T435921 [09:48:22] T435921: Rebalance cloudvirts out of F4 and into D5 - https://phabricator.wikimedia.org/T435921 [09:50:10] I'm going to kick the neutron agent running on it to see if that does anything [09:50:26] ack [09:52:17] wow that did it didn't it? I see both instances got network back [09:52:18] seems like it did, metricsinfra-puppetserver-1 is now reachable again [09:52:49] "My disappointment is immeasurable, and my day is ruined" [09:53:35] thank you taavi [09:54:38] looking at the neutron-openvswitch-agent logs at the time the migration happened I see it doing things it's supposed to be doing, so I have no clue what when wrong there [09:56:38] same here [10:06:20] dcaro: nfs-39 is back, haven't looked into -70 though yet [10:06:27] \o/ [10:07:07] is that because it had an ip? (I restarted the network before from the console and got one) [10:07:28] I did not touch 70 so had no ip [10:07:56] heh not sure, taavi restarted neutron agent on the cloudvirt and vlan/legacy came back [10:12:15] I'll take a look at -70 too [10:15:12] I'll just reboot it, load average 3k [10:21:40] ack [10:21:47] that sounds like stuff piling up on io? [10:25:47] indeed [13:05:18] hmh. conftool still uses the v2 etcd api, and of course the v2 keyspace backup tool only works on the host itself :/ [13:05:36] so I can't just come up with a Single Thing to back up both the toolforge k8s etcd clusters and the new cloudinfra etcd cluster [13:11:40] sighsob... we're then either pending the work in prod for conftool v3 apis or have to do two separate things [13:12:09] but wait, isn't possible to enable both APIs on etcd IIRC? [13:12:21] maybe I misremember [13:13:28] yes but AIUI you can't use the v3 api to query keys set with the v2 api or vice versa, or at least that's how I read the second paragraph on https://etcd.io/docs/v3.5/op-guide/recovery/ [13:14:39] mmmh, looks like, I would probably just try and see [13:16:28] yeah, I'll have to try [13:16:50] if not I guess backing up on the local storage nodes is good enough as long as we do it on all nodes and we can migrate it to the v3 snapshot once conftool migrates to it [13:16:58] quick reviews on clinic scripts https://gitlab.wikimedia.org/repos/cloud/wmcs/utils/-/merge_requests/11 and the stacked https://gitlab.wikimedia.org/repos/cloud/wmcs/utils/-/merge_requests/12 [13:23:41] taavi: +1 [13:39:18] godog: was cloudvirt1077 not already set to virtualization workloads? I think I've lost track of what's happening on T435537 [13:39:19] T435537: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537 [13:41:02] andrewbogott: heh I have updated the task description with the current TODO, i.e. bring cloudvirt to the provision cookbook state [13:41:33] i.e. the settings in the task description [13:52:53] But the point of that task originally was to set the powersave mode differently from what the provision script does... has the provision cookbook been patched out of band? [13:53:45] not afaik [13:57:40] my understanding is that we want intel_pstate and then nudge the cpu via EPB/EPP to do the right thing [14:26:05] did we get rid of the 'needs review' gitlab label? [14:35:57] yep, we are not monitoring that anymore [14:36:18] (still pending on defining exactly what replaces it) [14:40:54] remind me, exit code 99 from tofu provisioning CI is "all good but there are changes?" [14:41:37] volans: yes iirc [14:41:46] k thx [14:44:19] can I get a +1 on T438363 (moving from toolsdb to trove) [14:44:20] T438363: Request creation of Trove Database for Broad Interwiki Enhanced Retrieval VPS project - https://phabricator.wikimedia.org/T438363 [14:54:49] this one need two +1 (2x quota) https://phabricator.wikimedia.org/T438516 [15:06:25] dcaro: added a second +1 [15:06:36] thanks! [15:07:58] T438363 still needs a +1 [15:07:59] T438363: Request creation of Trove Database for Broad Interwiki Enhanced Retrieval VPS project - https://phabricator.wikimedia.org/T438363 [15:14:00] got it, thanks andrewbogott! [15:24:16] hmpf... openstack provider is failing to install [15:24:20] https://www.irccloud.com/pastebin/XjEvEU1c/ [15:24:38] the 3.0.0 seems like a major version bump maybe? [15:25:34] oh no, two years old, from the changelog `## 3.0.0 ( 25 September, 2024 )` [15:25:43] dcaro: where are you seeing that? [15:25:53] https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/jobs/968905 [15:25:58] an MR to add a trove project [15:26:14] * taavi tries to retry [15:26:32] my first guess would be a github transient failure, those are not uncommon these days unfortunately [15:26:49] yep, I'll retry [15:26:53] https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/jobs/968927 [15:27:02] well, it's now a different error that supports my theory [15:27:27] ack xd [15:27:31] I'll wait a bit then [15:27:37] (does the CI job not use the vendor_modules we include with that git repo?) [15:28:39] no idea, how can I check? (if it pulls it again I guess it does not?) [15:30:00] I don't see the config file that would need, so probably past me forgot that option [15:30:26] retried once more and it finally passed [15:32:36] thanks \o/ [15:33:00] do you want me to create a task for the vendoring? [15:35:56] sure [15:39:14] T438754 [15:39:14] T438754: [tofu-infra,ci] Make the ci use the vendored providers directory in the git repo - https://phabricator.wikimedia.org/T438754 [15:56:08] I'm considering sending an email to the cloud last asking if anyone will speak up in defense of instance suspend and pause. Is asking just wp:beans? [15:56:56] Terminology: suspend and pause both retain the runtime/ram state of a VM, as opposed to a shutdown which does not. Pause keeps the state in local ram on the cloudvirt, suspend dumps it to disk. [15:57:35] It's my bofh opinion that if you care about the runtime ram state of your VM then... you shouldn't. But maybe that's draconian. [15:57:48] what user flows does it enable? (as in, when will users want to suspend/pause) [15:58:21] resilience wise I would lean towards being restart-ready, but maybe there's things like migration that would be easier [15:58:31] I can't think of any good ones. I mean... if something very interesting was happening on a server and you want to stop time until you have time to investigate... [15:58:43] but as soon as you restart time there would be clock drift and other weird things [15:59:10] yep, I remember having those issues in the past (when using a vm to test stuff, then suspend when not in use) [15:59:11] (paused and suspended VMs can't be live migrated, which is the inciting reason for us wanting to remove the feature) [15:59:58] so with the info I have so far, I would agree to remove the feature [16:01:29] I think it's ok to ask for "anyone using this?" kind of thing, can we check when was last used or something? [16:02:02] I thing filippo ran into a suspended instance this morning, that's why it came up. [16:03:01] then yep, +1 for asking why that was (see if there's a use case for it) [16:05:40] https://etherpad.wikimedia.org/p/suspendsuspending [16:05:50] Also, I think that's a cloud list thing not a cloud-announce thing, does that sound right? [16:19:07] agree, +1 for cloud list [16:33:44] * dcaro off [16:33:47] cya tomorrow! [17:14:48] andrewbogott: the clouddb memory alert has come back, on a different host (clouddb1022) [17:14:52] I reopened T438200 [17:14:52] T438200: clouddb memory alert - https://phabricator.wikimedia.org/T438200 [17:23:31] I saw it but still don't know what to do about it. Guess I'll restart things if it goes oom [17:29:14] I added some more info to the task, hopefully they can help manuel in debugging [17:29:31] it doesn't seem to be load-driven, as in there are currently no queries running [17:29:54] I can try killing all the "Sleeping" threads to see if it makes a difference [17:31:58] that did not make any difference, maybe it was some earlier query that is no longer running [17:32:10] manuel said that it never frees memory, only ever consumes it? [17:32:29] I would hope that the cap would be set to something less than 100% [17:33:05] yes I think that's right, but it should be capped... we also have 2 running instances on that host, each one should have its cap [17:36:09] huh [17:37:53] I wouldn't worry for now, let's restart it again, if this keeps happening manuel might want to collect some tracedumps [17:38:06] innodb_buffer_pool_size are configured to 100 and 105, it's already well over that [17:38:40] I think that's just one of the components using memory, finding the "total" is complicated IIRC [17:38:59] if I just restart mariadb will I need to restart replication after? [17:39:39] yes, and I would probably stop replication manually before, even if systemctl stop should also do it cleanly [17:40:37] ok, let's see if I can get this right... [17:41:47] commands are "stop replica;" and then "start replica;" from the "mysql.s3" console [17:42:18] https://wikitech.wikimedia.org/wiki/MariaDB/Rebooting_a_host uses 'slave' and not 'replica' are they synonyms? [17:42:29] yes exactly the same [17:42:53] ok, I stopped replicas, restarted both mariadbs, restarted replicas. [17:42:56] Look right to you? [17:43:19] yes [17:43:32] memory is down [17:43:49] and replicas up? [17:44:00] yes, just checked [17:44:02] thank you! [17:44:07] thanks for checking my work! [17:44:19] Now -> lunch [17:48:24] * dhinus off [18:28:07] * andrewbogott back