[07:11:15] greetings [07:58:17] morning! [08:58:28] morning [10:09:02] I merged the new tool alerts, and I was NOT expecting them to appear in https://alerts.wikimedia.org/?q=team%3Dwmcs, but they are. looking into it. [10:09:59] something must be adding team=wmcs even if I didn't specify it in the alert definition [10:10:41] dhinus: https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/66/diffs#c9a6313481702b8993d08e5e86a07832c98548cf_0_30 [10:11:04] taavi: ah, that was easier than I thought :D [10:11:35] sorry about that, I'm pushing a fix [10:12:26] also, i thought i asked that i can review your overall alerting system design when you have one :D in particular i'm curious how you plan to turn those into something that will alert the individual maintainers [10:14:49] also also, are you aware that whether a rule in toolforge/alerts.git has the `team: wmcs` tag has no effect on whether it'll be routed to IRC -feed or the feed mailing list? [10:16:30] taavi: sorry if I didn't ping you explicitly, all reviews are welcome both in gitlab and in the task T432863 [10:16:30] T432863: Define alert conditions in Prometheus (if possible) - https://phabricator.wikimedia.org/T432863 [10:17:17] I thought the mailing list was also based on the team label, if it isn't, I'll need to find a way to prevent these alerts from showing there [10:18:02] it is not, the only routing the tooling can currently support is the `project` label which will always be `tools` for any alerts off of that prometheus instance [10:19:29] I guess I'll disable the alerts for now while we find a solution [10:19:37] I'll keep the recording rule [10:30:26] hm, actually https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/production/modules/profile/manifests/toolforge/prometheus.pp#668 is also adding the team: wmcs label regardless of whether an alert has it already [10:31:38] so that's also making all of those show up with the team=wmcs filter on the alert dashboard, but the problem with being routed to irc/mail is a separate one [10:36:42] this disables them to avoid the spam: https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/68 [10:36:57] but CI is failing, fixing it [10:44:44] taavi: do you think the "relabel" in prometheus.pp can be removed, or is it still required for some alerts? [10:46:55] dhinus: it's going to break the PrometheusReloadFailed alert from puppet alerts_default.yml but we can copy that alert with the needed labels to toolforge/alerts.git, otherwise I don't see why we can't remove it [10:47:25] nice, at least that one should be an easy fix [10:47:39] yeah, that's the easy one of the two problems here unfortunately :( [10:48:19] do you think we need a separate prometheus instance? I'll have a look later today, if you have ideas please leave a comment in phab [10:49:07] in the meantime this is now passing CI, can you add a +1? https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/68 [10:49:19] i think all of this depends on how you plan to send the alerts from prometheus to individual tool maintainers, which your phab task don't seem to detail yet at all [10:50:20] that's the separate T432865 [10:50:21] T432865: Set up a Toolforge “alerting service” that emails tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865 [10:51:36] but yeah, you either need completely separate prometheus+AM instances, or you need to extend the metricsinfra tooling for more specific alert routing rules support [10:54:03] thx, I'll check the current metricinfra setup to see how hard it would be to extend the tooling [10:54:13] * dhinus lunch [11:01:22] * dcaro lunch [11:53:57] dhinus: https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/69 and https://gerrit.wikimedia.org/r/c/operations/puppet/+/1324696 [11:56:59] taavi: thank you! +1d gerrit, waiting for CI on gitlab [12:03:51] taavi: do you mind if I push https://gerrit.wikimedia.org/r/c/cloud/wmcs-cookbooks/+/1188832 through? just fixed the test, I'm releasing packages so it would be good to have it in xd [12:04:16] dcaro: please do, thanks [12:36:22] oh, we have a bookworm bastion in toolsbeta [12:36:57] can I nuke it? [12:37:15] * dcaro looks if it's being used [12:40:44] I don't think it's used anywhere [12:40:54] https://gerrit.wikimedia.org/r/c/cloud/wmcs-cookbooks/+/1324703 quick review [12:43:44] huh, thought i removed it already [12:44:32] I'll remove it then :) [13:21:09] hello! Fundraising Tech filed a bug last Friday about their CiviCRM job timing out, that got addressed by andrewbogott which removed some disk setting which caused the slow down on some Ceph/OSD whatever. So that one is resolved and the build is no more timing out [13:22:04] then I got feedback that the build duration varies greatly the fastest ones are at 16 minutes, most are at 24 minutes. That is a sizeable slowdown [13:22:17] so I thought that the hosts running the build in 24 minutes might have a problem [13:22:29] hashar, in this case 'hosts' means cloudvirts? [13:22:35] or VMs? [13:22:41] instances, but yeah cloudvirts under that [13:23:27] in the examples of fast build being given, they seem to have run on `integration-agent-docker-1093` which is backed up by `cloudvirt1058.eqiad.wmnet` [13:23:55] which albeit the host has some business (CPU usage) compared to anothercloudvirt that is idle, the build is still fast [13:24:07] so the issue I have is that the instance on cloudvirt1058 is just dramatically faster [13:24:42] I got two builds happening at 7:00 UTC yesterday (when it is super quiet), one ran in 1450 seonds and the one on cloudvirt1058 ran in just 952 seconds!! [13:25:18] and eventually I looked at the CPU MHz reported in /proc/cpuinfo . All instances have the same 2593 MHz reported [13:25:50] with the exception of `integration-agent-docker-1093` which reports 2893 MHz [13:26:16] rounding: most are 2.6MHz , that other one is 2.9MHz which is 11% faster clock speed [13:26:31] but that don't explain why builds are so much faster on that host ;] [13:26:54] my large brain dump is on https://phabricator.wikimedia.org/T434024#12205386 and following comments [13:27:37] usually I would expect disk access to be more of a factor than clock speed, any chance the VMs have different flavors/IO throttles? [13:27:56] I assume you already eliminated 'running more jobs' as a factor [13:28:10] yeah I eliminated that (hopefully) [13:28:28] I guess the underlying hardware is simply different and performing better for whatever reason [13:28:56] back in 2019 I did notice some dispredancy which I think I ended up blaming the power management system for. The hardware was about to be replaced anyway so it was not further investigated [13:28:58] that's most likely. Cloudvirt numbering is straightforwardedly old to new, so if lower number cloudvirts are slower that's not a shock. [13:29:13] But there are also potential noisy neighbor issues [13:29:46] do we have a list of hardware spec for each of the cloudvirt somewhere? [13:30:40] procurement tickets are in phab, I don't know if you can see them or not. Let's see... [13:31:53] I might be able to grab them from Puppet facts [13:34:15] yeah, the OS should know most things about what hardware it's running on. [13:35:14] Maybe you're already looking at this, but this dash should be helpful for noisy neighbor questions: https://grafana.wikimedia.org/d/000000579/wmcs-openstack-eqiad-summary?orgId=1&from=now-7d&to=now&timezone=utc&var-hypervisor=$__all&refresh=1m [13:44:09] I remember instances can reported "steal" CPU usage, which is the instance not being to have CPU allocated to it because the host is busy with other instances [13:44:35] anyway I imagine a host having a low CPU usage means there is no noisy neighbor issue isn't it? [13:45:10] That's correct, although there could be transient spikes of CPU contention. In this case, most likely if you have a bunch of busy CI nodes on the same cloudvirt. [13:46:14] * hashar nods [14:04:08] taavi: can you tell me about this line in the upgrade docs? "Create a merge request in the lima-kilo repo, updating the Kubernetes version and also the versions of kind, helm, helmfile"? lima-kilo is already running a different version of helm from toolforge, but we want those to be in sync don't we? [14:04:33] ...leading up to the question "do you care about upgrading helm and helmfile as part of this k8s upgrade cycle? And if not, shouldn't we leave l-k where it is?" [14:11:36] dhinus, you might also have an opinion about ^ [14:11:59] i don't have very strong opinions about lima-kilo anymore :P [14:12:14] Upgrading to modern helmfile means we have to update our helm charts as well, there are breaking syntax changes. https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1370 [14:12:40] taavi: OK, that still leaves us with "do you care about upgrading helm and helmfile as part of this k8s upgrade cycle?" [14:14:54] the helmfile version we're running now is comically old but I'm not sure that really matters [14:15:50] if we're not upgrading it in the real toolforge then i don't think we care about upgrading it in lima-kilo [14:16:12] That's what I'm asking though -- do you want us to upgrade it in toolforge? [14:17:05] iirc we use the serviceops maintained package there? in that case i wouldn't [14:17:14] oooh I see. ok [14:17:59] In that case... I will probably push ahead to doing the k8s upgrade in toolsbeta today if you don't object. I'm going to be out tomorrow but also having toolsbeta broken for a day seems fine? [14:21:16] are you just assuming the upgrade will always break things or do you have a more specific reason to suspect breakage? [14:22:12] I do not expect breakage. It's just that my upgrading today is the moral equivalent of deploying on a Friday and I'm trying to keep my conscience clean. [14:22:57] go for it then [14:23:11] ok! [14:35:06] andrewbogott: sorry in a meeting, will reply later [14:36:53] np [14:42:33] I got the list of CPUs on cloudvirt, there are 3 of them being used https://phabricator.wikimedia.org/T434024#12208164 [14:43:15] and the last ones (used on cloudvirt1077-1080) apparently has a smaller clock speed, albeit it has double more threads/cpu whatever, lot more L3 etc [14:54:03] that's the future I suspect... monolithic super-fast CPUs are out of fashion. [14:59:52] then for the typical WMCS workload I guess we want CPUs with a lot of threads, so that make senses [15:00:05] and possibly the fancy new processor end up consuming less power [15:05:53] that's the idea at least [15:35:50] andrewbogott: I think I wrote that line in the k8s upgrade docs. the idea was to use the k8s upgrade as an opportunity to refresh all the versions of the k8s-related tooling to make sure we run a recent version in lima-kilo [15:36:13] but I think having the same (or similar) version to the one we run in prod might make more sense [15:36:19] My first thought is that we don't want to run recent versions, we just want to run the same... [15:36:21] yes :) [15:36:29] I can update the docs. [15:36:53] although I still think we should reformat our helm files because it gets rid of so, so many warnings. [15:39:38] please do. the only reason of updating helm in l-k during the upgrade is probably to have a "realistic" test env to test the new k8s version, so if the helm version is already similar to prod you might keep the current one [15:40:24] +1 for using the prod version one, and +1 for upgrading the syntax to the latest supported by that version [15:40:58] I'm currently working with a customized lima-kilo on an MR, but I can try to rebuild it once I get done with my tests (not today though) [15:42:39] kind is lima-kilo only, I don't have strong opinions on whether to upgrade it during a k8s upgrade, or just keep the old one if it works fine [15:43:45] I think better upgrade (sometimes you need it for the newer k8s images), but yep, can be done after/later [16:39:01] thanks @bd808 for the note on T434684 [16:39:01] T434684: 500 internal error when accessing toolsadmin ssh-keys - https://phabricator.wikimedia.org/T434684 [16:41:42] cya tomorrow! [17:05:57] * dhinus off [17:47:26] andrewbogott: do we know what version of helm/helm charts is being used in prod, is it same? [17:48:03] helmfile yes. helm is similar but not the same so we should update lima-kilo just enough to catch up with toolforge