[06:52:39] greetings [07:56:04] morning! [08:43:47] morning [09:33:18] dhinus: I'm having issues deploying on tools, I'm looking [09:34:58] I'm deploying a change right now [09:35:15] and it was successful: "END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component alerts-api" [09:36:16] component1(failed): 400 Client Error: Bad Request for url: https://api.svc.tools.eqiad1.wikimedia.cloud:30003/jobs/v1/tool/automated-toolforge-tests/jobs/ (400): No such image 'tools-harbor.wmcloud.org/tool-automated-toolforge-tests/component1:latest@sha256:9e2729dfb4ddc724bf3929df13eb9a15d2a60872787ff221eb1c07a5c37f1a74' [09:36:29] if it does not try to run the jobs-api tests it might not fail [09:36:48] correct, my deploy did not run the tests [09:36:52] for some reason jobs-api thinks that image does not exist, but it does :/ (I can pull it, there's a build for that image) [09:37:47] is the jobs-api logging anything? [09:39:25] just that the image does not exist :/ [09:39:37] I redeployed and did not fail this time [09:46:55] it failed again now :/, looking [09:47:14] component1(failed): 400 Client Error: Bad Request for url: https://api.svc.tools.eqiad1.wikimedia.cloud:30003/jobs/v1/tool/automated-toolforge-tests/jobs/ (400): No such image 'tools-harbor.wmcloud.org/tool-automated-toolforge-tests/component1:latest@sha256:2b3556fbad1751db3439635c47c576ec7dc56aaf0acab9118acd2d7e1f0dbac6' [09:58:39] hmm.... I suspect there's something borked with one of the jobs-api pods or something similar, now it's passing again [10:04:54] maybe one pod is failing to connect to harbor? [10:07:34] the pod reaches there [10:08:11] just deleted the pod, let's see if there's more issues [11:28:43] from the backlogs: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1343544 https://gerrit.wikimedia.org/r/c/operations/puppet/+/1343545 [11:41:10] if noone is using cloudcumin2001 today I can probably find the time to upgrade it in place to trixie and do 1001 in the next few days [11:45:09] did anyone do anything with paws? It seemed to have gone down and up [11:51:25] dcaro: I restarted some k8s nodes two days ago, didn't touch it after that [11:51:35] 3 nodes were NotReady on monday [11:52:27] ack [11:52:29] volans: I think you can go ahead [11:54:20] thanks [11:56:39] yeah I don't think anyone really uses 2001 [12:38:13] I'm getting this on lima-kilo for loki [12:38:17] https://www.irccloud.com/pastebin/SCFKwDyz/ [12:38:33] anyone found that issue before? [12:40:54] I suspect that the image might be gone (2024) [12:43:02] I just had to go through >5 recaptchas clicking cars and stairs and such... I'm starting to think that I'm half-robot [12:44:03] haha I always get them wrong 70% of the time [12:44:14] minio has been archived :/ [12:44:26] (on github at least) [12:46:26] oops [12:46:27] https://github.com/grafana-community/helm-charts/blob/5572808b898779a9b0abdbf09f5b947229d9a9eb/charts/loki/README.md?plain=1#L170 [12:47:11] minio is deprecated on loki too :/ [12:47:22] anyone has the cached images? xd [12:48:15] LOL I can check, but we need to find an alternative anyway. I also missed that minio was enshittified a few months ago and now has become "AIstor" :( [12:48:16] loki fails to deploy on lima-kilo now for me [12:49:14] this seems to be a drop-in fork (have not tested it): https://github.com/pgsty/silo [12:50:46] nice, we'll have to figure it out yep, loki at least does not offer a replacement (just say to go use s3 instead) [12:52:55] created T439720 [12:52:56] T439720: [lima-kilo,logging,logs-api] find a replacement for minio (s3 for loki) - https://phabricator.wikimedia.org/T439720 [12:57:15] xd, just replacing the minio image for docker.io/pgsty/silo:latest seems to do something [12:57:19] * dcaro feels a bit dirty [12:59:30] yep, that seems to kinda work [13:07:00] dcaro, dhinus: Toolhub is a Tools Platform service, right? could one of you look into it for https://phabricator.wikimedia.org/T439713 ? [13:08:33] moritzm: yes it falls under our team, I'll have a look [13:08:53] cheers [14:13:27] godog: I'm sorry that you're spending so much time with partman :( Just now I tried to reimage cloudcephosd1043.eqiad.wmnet and partman is failing. I will look at the logs if/when my eyes uncross but I thought you might be interested in a regression in that ever-more-complicated recipe. [14:14:48] andrewbogott: ack thank you for letting me know, and yes it is a bit of a mess [14:15:38] andrewbogott: I did remove the special casing here, feel free to put it back if you'd like to test https://gerrit.wikimedia.org/r/c/operations/puppet/+/1345942/2/modules/install_server/files/autoinstall/scripts/partman_early_command.sh [14:16:12] what I'd aiming to have is not yet another hostname matching in that file [14:16:28] Of course on my journey I've rebooted the server a couple of times and the drives are re-labeled each time. I know that this happens and it still makes steam shoot out my ears every time I see it. [14:17:06] indeed [14:17:13] ah, so it isn't calling configure_cephosd_disks at all anymore? Hm... in theory that should make things better :) [14:18:16] I will probably revert that patch and retry if you don't have anything in process for the rest of the day. [14:19:06] andrewbogott: sure no problem, feel free to revert only the shell bits, preseed.yaml changes are a no-op for 1043 [14:19:37] yep [14:19:44] and yes nothing in process for the rest of the day, I'm testing with https://gitlab.wikimedia.org/repos/sre/preseed-test [14:39:09] FYI re: performance governor for cloudvirts https://phabricator.wikimedia.org/T439729 [15:40:40] Can anyone help look at https://phabricator.wikimedia.org/T439317 cloud-vps request ? It’s a bit unclear what the request is about. [15:40:48] As much as I understand resizing cloud vps project vms is something the user has to handle for themselves. [15:41:19] I already left a response, but just to be safe a second pair of eyeballs would be appreciated [15:43:32] Also need a +1 for this Toolforge envvars quota increase request https://phabricator.wikimedia.org/T439237 [15:46:37] Raymond_Ndibe: +1d [15:50:09] On the task, I added a screenshot of the resize button (that I think they can see too) [15:50:57] might be good to add a note in the docs also that the instance can be resized, I only found a side-comment here https://wikitech.wikimedia.org/wiki/Help:Cloud_VPS_instances#Instance_information [15:52:41] Ok thanks David! [17:21:03] * dcaro off [17:21:04] cya! [17:45:04] * dhinus off [20:45:01] received an SOS (request for help) from Silvia Gutierrez as they're trying to stage a training in Mexico, using the beta cluster. Do we have control over what IPs are allowed to access it, and can we open it up for an event? [20:49:40] cross-reference http://phabricator.wikimedia.org/T439757 [20:53:42] bliviero: it seems that's part of the list of blocked_nets in horizon't hiera config for puppet [20:53:52] it lists - 201.141.96.0/21 [20:54:19] I guess we can unblock the whole /21 for now, most likely would not be a big problem [20:56:15] why is this not handled by whoever owns the beta cluster (i.e. DevEx)? [20:57:18] now I need to undertand where to run puppet... [21:01:06] TBH I wasn't sure if we controlled the IP allow/block list, vs DevEx. as bd.808 is out, i chose to ask in our team first before going to Tyler [21:01:35] I can certainly go there (and will) [21:01:41] thank you [21:02:05] {done} puppet run, updated task [21:04:10] oh awesome volans thank you! [21:04:56] it also needed an haproxy reload, sigh :/ [21:05:40] i can find someone in devex to do it, let me find them [21:05:53] I just did it [21:06:05] I hoped puppet would have handled that, but apaprently not in this case, dunno [21:06:22] anyway, I did it because of the SOS but in general yes I don't think a request like this one should land on our team [21:06:38] I assumed the usual owners where not available for some reasons [21:06:45] thank you so much! I appreciate it ! and lesson learned