[08:16:23] Oh wow, hughops to all involved [08:40:09] * volans errand to run, bbiab [08:45:02] dhinus: when you are around, can you double check that you are ok with the naming in https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/192 ? [08:45:44] dcaro: I'm here but was reading the backscroll from the incident :) [08:46:01] no rush, just fyi [08:47:08] looks good, approved! [08:47:34] ack [09:40:47] dcaro: I noticed that components/gen/toolforge_models.py still has the old syntax (enabled instead of notify) [09:41:55] * volans back [09:43:04] dhinus: oh, it should not, let me check [09:43:29] oh yes, that might be true, it generates the models from the public api endpoint openapi.json, not the code itself [09:43:40] so until it's deployed, it does not fetch the "new" openapi.json [09:44:13] it does not use those models though, as it has the actual original ones instead [09:44:21] (for the components-api side of things) [09:44:40] I see! [09:44:54] I guess we could try to filter the models that are generated somehow (did not look into it), and avoid generating the ones we don't need [09:45:34] not urgent [09:46:23] today I created T439108 also [09:46:23] T439108: [components-api] create a ci scheduled pipeline to update the toolforge models - https://phabricator.wikimedia.org/T439108 [13:47:53] dhinus: are you going to rejoin the call? [13:49:22] dcaro: I opened a thread in slack, the mic is no longer working, I even tried the discord app instead of in-browser :) [15:33:58] andrewbogott: fwiw I just fixed T379550 for the seanleong-wmde account, in case you feel like poking into the keystone logs in case there's something useful there [15:33:59] T379550: openstack: keystone may be failing to add users to the bastion project in Keystone and/or LDAP - https://phabricator.wikimedia.org/T379550 [15:34:21] also, note to self: the openstack CLI doesn't give any useful warnings if you try to remove a typo'ed role from an user [15:34:28] * andrewbogott scowls at that bug [15:34:50] not sure if the initial addition that failed was before or after the logstash fix, though [15:35:47] happen to know when the... [15:35:52] ok, that answers that [15:36:30] first mention I see is from 14 minutes ago [15:36:53] that's probably me trying to fix it? [15:36:57] yeah [15:37:28] this is from the support query in -cloud from yesterday that I missed during the codfw saga [15:40:45] they were already in there on 2026-05-07 [15:40:58] so the actual failure is probably lost to time [15:41:04] but thank you for fixing! [15:49:33] can someone take a look at this one? T438255 (looks to be potentially tool-specific, but would be good to make sure) [15:49:34] T438255: Can't deploy anymore on Toolforge (pod stays pending) - https://phabricator.wikimedia.org/T438255 [15:50:34] aputhin: on a quick look that goes in the giant 'we should have larger workers so that we can better schedule large things' pile [15:52:34] taavi: is there anything stopping us from having bigger workers other than someone getting around to running the cookbook? [15:53:12] andrewbogott: I think we want someone to look at the data (that g.odog collected?) to figure out what kind of a CPU:RAM ratio we want for those [15:55:28] ok [15:58:29] bliviero: ^ this is one of the things I was forgetting on Tuesday on what we should work in Q2 [16:06:16] taavi: noted [16:15:59] for the immediate issue that was filed (T438255) what should our response be? [16:16:00] T438255: Can't deploy anymore on Toolforge (pod stays pending) - https://phabricator.wikimedia.org/T438255 [16:18:50] the cluster seems pretty busy (compared to my memory of some time ago at least) [16:23:19] are there servers that are depooled that could be brought back in for capacity? [16:24:35] no, but we could create more workers (of the existing size) pretty easily. dcaro, taavi, any objection to me adding another 6 workers? Or X workers? [16:25:07] nope [16:25:20] not really no +1 [16:26:36] dcaro: you were looking at the capacity section on https://grafana.wmcloud.org/d/3jhWxB8Vk/toolforge-general-overview?from=now-12h&to=now&timezone=utc&var-cluster_datasource=P8433460076D33992&var-cluster=tools ? [16:26:44] Looks like 50% usage to me but maybe I'm misreading [16:26:51] actual RAM usage I mean [16:27:04] Oh, I guess for scheduling new jobs it's the other panel that matters [16:27:18] 190% [16:27:21] yep, I was looking at the reservations [16:27:32] does that cap out at 200%? [16:27:33] yep, so we might want to tweak that again [16:28:45] for scheduling, the main thing that matter is the largest free slot on any one node. we used to have a dashboard with that visible but I don't immediately remember which one [16:29:05] yep, not finding it either :/ [16:32:34] hmm.... [16:32:36] 'largest free slot' is still governed by the overprovision number right? [16:32:47] So if we moved from 2x to 3x... [16:33:01] all seem to have >5G memory free [16:33:06] https://www.irccloud.com/pastebin/4b8yoteF/ [16:33:37] (unless I'm reading that wrong) [16:33:47] dcaro: the logs in the task say 'Insufficient cpu', not memory [16:34:03] oh, that's interesting [16:34:20] (oh, too, that's free memory, not unreserved memory I think) [16:34:58] cpu is less reserved than memory even :/ [16:35:10] in that dash, is 'limits' the max allowed reservation [16:35:12] ? [16:35:17] Because if so, we are nowhere close [16:35:37] (and not close for RAM either) [16:36:24] it's the requests [16:36:39] the limits is the "maximum memory/cpu allowed to be used" sort of thing [16:37:07] the request is "minimum that my pod will need allocated to run" [16:37:22] why do 'limits' change over time? [16:38:09] it's the sum of all pods that are currently defined in the cluster [16:38:19] cronjobs starting and stopping, I assume [16:39:57] there was a weird jump this morning (not that it should start making things break though) https://usercontent.irccloud-cdn.com/file/cTrD3A9I/image.png [16:40:16] that was yesterday night [16:40:58] in any case, more workers should be ok, make sure to use the bigger images though (2x the current one iirc is what we said?) [16:41:00] I don't see anything that suggests we have a capacity crunch. So... why can't that user scheule? [16:41:27] it needs to have a node with >5G 'requestable' [16:42:08] and 3 cpu [16:42:17] (at the same time) [16:42:23] not sure if there's one, maybe? [16:43:59] on average there should be quite a few [16:45:09] interesting, this is giving me quite different numbers [16:45:30] dcaro@tools-bastion-15:~$ kubectl-sudo describe nodes | grep -A5 "Allocated resources" [16:45:47] https://www.irccloud.com/pastebin/LEuMrilt/ [16:46:03] cpus are almost saturated [16:46:05] sorry I'm busy with production's repooling, but have you already compared reserved vs allocated resources? [16:47:15] on it yes, I have the suspicion that the graphs we have are not correctly reflecting the actual reserved amounts [16:47:23] allocated is going to be less (maybe much less) than reserved, right? So 'reserved' is a good upper bound [16:47:35] yep [16:47:49] reserved is what the scheduler uses to decide where/if to put the pod in the node [16:47:51] yes and the scheduler looks at reserved IIRC [16:48:30] all but the control nodes show >90% cpu reservation using that command [16:49:13] but we allow up to 180% right? [16:49:17] Or do you mean 90% of allowed? [16:51:12] yep, 90 of allowed, as in, only 10% is left to be considered when putting a new pod in it [16:51:40] wow, very different from what grafana says! [16:51:56] Ok, most workers right now are 8cores/16GB, I'm going to make more workers that are 16/16 [16:52:14] grafana might be averaging all the workers? but even then, there's many >90% (only 6 control nodes) [16:52:25] We still need to do some real number crunching and graph fixing but that that should help for now. [16:52:42] And I assume we're only really talking about nfs workers today? [16:54:25] I think all of them, but let me double check [16:57:13] andrewbogott: you mean 16cores/32G ram? [16:57:27] no, 16/16 [16:57:32] Since we seem to have plenty of RAM [16:57:48] and also because that's a flavor that exists, otherwise I'll have to start with a tofu patch [16:58:32] non-nfs also have capacity issues yep [16:59:06] hm, actually openstack is lying to me about available flavors... [16:59:47] yeah, have to make a tofu patch regardless. [17:00:08] So, do you think better to skip ahead to 16/32? [17:00:56] that should be ok yes, or even more cpu too (given that it seems our assumption that we use more ram than cpu might not be correct) [17:01:54] the VMs having idle memory is not an issue right? (openstack will overallocate?) [17:01:56] At some point it gets hard to schedule giant VMs due to roundoff on the hypervisors... [17:02:05] No, we never overprovision RAM, only CPUs [17:02:24] but we have lots of RAM so it is usually not an issue. [17:02:31] okok, good to know, so then probably more cpu to ram ratio is better (we can double check later too, once we have good data) [17:04:08] https://www.irccloud.com/pastebin/HKaQsoKs/ [17:05:42] dcaro, taavi, https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/364 [17:07:14] LGTM, feel free to make them smaller if you think they will bring problems on the openstack side [17:07:54] Should be OK. I wouldn't want to do 128 though :) [17:09:58] xd ack [17:10:35] I'm going to take 15min before the discord call to stretch my legs, ping me if you need any help, we should follow up on the stats though [17:11:52] sure. It's just the long wait for cookbooks to complete now. [17:25:00] dcaro: oops, I forgot about quotas. https://phabricator.wikimedia.org/T439157 [17:26:24] andrewbogott: got the meeting it 6 min, just +1d, please self-provision xd [17:27:08] thanks, I will [18:37:52] * dcaro off [18:37:55] cya on monday! [19:20:41] I rotated the etherpad and will cover clinic duty tomorrow, then I'll hand it over to r.aymond on monday [19:20:53] * dhinus off