[06:22:44] greetings [07:57:50] morning! [08:50:33] morning [09:13:03] this is annoying xd, spent 10min wondering how the harbor API changed so much (https://harbor.my/, not our harbor) [12:24:19] I'm taking a look at the puppet failures on cloudvirt [12:27:00] failure acquiring cfssl certs, filing a task [12:49:10] thanks! [13:08:13] T440581 Luca is on it [13:08:14] T440581: i/o timeout from cloudvirt when talking to pki.discovery.wmnet - https://phabricator.wikimedia.org/T440581 [13:08:21] spoiler alert: vlan acls [13:08:52] did we change the network switches? was it an upgrade? (iirc there was one about to happen) [13:09:16] no changes to the network, pki moved from port 80 to 443 tho [13:09:28] or some other port and to 443 [13:12:23] ooohh, ack [13:14:08] actually port and ip, host:80 to lvs:443 [13:33:25] * andrewbogott just started to investigate the puppet failures, should've read IRC first [13:37:24] godog: can you talk me through "max by (instance) (rate(node_schedstat_waiting_seconds_total{cluster="cloudvirt",site="eqiad"}[$__rate_interval]))" ? [13:37:45] node_schedstat_waiting_seconds_total increases monotonically right? So rate() around that tells us how quickly it's increasing... [13:37:54] but I don't understand the max by (instance) part [13:40:49] andrewbogott: yes that's right how quickly it is increasing, the max() is because the metric is exporter per-cpu and that'd be the "worst" cpu [13:41:19] oooh and (instance) is the cpu [13:41:31] ok! So the vertical scale is seconds? [13:42:04] no instance is the host, cpu is in the metric though it gets grouped away [13:42:46] vertical scale is technically seconds per second, per cpu [13:44:13] more of an indication of general busyness as I understand it [13:44:21] ok. I don't think I've used max by before but I think I understand now. [13:44:31] Yeah, seems like it's telling us something useful :) ty [13:45:08] sure np, the more accurate way will be to ask libvirt, good enough for now I think [13:45:50] Yep. I don't think we need accurate numbers so much as "something very bad is happening on this one host" which various metrics should show us [13:46:02] s/accurate/precise/ [13:46:12] indeed [13:47:17] there are half a dozen good, legit toolforge requests this morning! Were there lots of Russians in that discord call by chance? [13:51:44] btw godog, I'm leaving you with one pending clinic request, https://phabricator.wikimedia.org/T440447 -- I think it needs a bit of refining and the user says they'll follow up. Probably we'll have to have them on cloud-vps since they seem to need hundreds of GB of storage. [13:54:05] andrewbogott: ok sounds good, thank you for the heads up [13:54:55] I tried to triage incoming phab tasks but the list is very long and I got distracted fixing one of them :) [14:02:45] andrewbogott: I have no idea of the nationalities/location of people in that call :D [14:03:03] but I think most of them were existing toolforge users [14:51:20] Well somehow we turned up a bunch of new interested users. Would be nice to know where they came from :) [14:52:54] andrewbogott: bliviero this is the task for the toolforge pending pods alerts T414513 [14:52:55] T414513: Add new alerts for Toolforge cluster high load - https://phabricator.wikimedia.org/T414513 [14:53:44] nice [14:55:34] I have the feeling that we should start creating recording rules for k8s stuff, both to keep record of aggregated stats, and to be able to monitor those (as currently we have lots of cardinality in k8s metrics) [15:33:03] thank you dcaro [16:49:57] * dhinus off [17:06:00] dhinus: (when back), how do you make a patch like https://gitlab.wikimedia.org/repos/cloud/toolforge/envvars-admission/-/commit/34db248f0950e4cab92dbc06e01f37b5e33425ed ? Is that automated somehow? [17:10:12] andrewbogott: I'm back here momentarily :) I think I used something like https://github.com/kubernetes/client-go/blob/master/INSTALL.md#using-a-specific-version [17:10:36] it was definitely some "go get" command that generated it, but I don't remember the exact one(s) [17:10:47] fancy! ty [17:10:57] * dhinus off for real :) [17:36:46] * dcaro off [17:36:46] * dcaro