[07:45:11] !log tools hard reboot tools-k8s-worker-nfs-70 - ran out of memory [07:45:14] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools/SAL [07:51:53] !log tools correction that's tools-k8s-worker-nfs-74 [07:51:56] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools/SAL [14:03:41] !log admin destroying ceph osd.334, osd.335, osd.336 and osd.337 for testing described on T429387 [14:03:48] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [14:03:48] T429387: cloudceph HEALTH_WARN, multiple OSD(s) experiencing slow operations in BlueStore - https://phabricator.wikimedia.org/T429387 [19:20:50] * lucaswerkmeister is about to try out what push-to-deploy does if you haven’t migrated the tool to the components API yet [19:31:06] hm, https://gitlab.wikimedia.org/toolforge-repos/wd-image-positions/-/jobs/971182 isn’t quite the outcome I expected [19:31:12] > fatal: Invalid revision range ed9bacb72c141b8bbbb0e655fb8b6917c7c0bddd..323162c915c7948fcc16d1f34c97a8b536309303 [19:31:30] pretty sure this comes out of what I added in https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/97, but I don’t know why it fails now… [19:31:57] also, in my local clone, that’s a valid revision range. did anything change in how the repo is cloned in CI? does it have an incomplete history? [19:32:52] (it says “Fetching changes with git depth set to 20” further up, but that ought to be enough, one commit is the direct parent of the other…)^U [19:32:54] scratch that [19:33:00] this is my fault for force-pushing to the main branch xD [19:33:34] (I thought X was the parent of Y because `git log ed9bacb72c141b8bbbb0e655fb8b6917c7c0bddd..323162c915c7948fcc16d1f34c97a8b536309303` only showed the latest commit, but that’s not actually true) [19:34:20] okay, so A), I should avoid force-pushing as it’s bound to confuse any “generate a deployment summary based on the git history” logic, but also B), I should make that logic more robust so it doesn’t break the entire build [19:34:45] * taavi screams at lucaswerkmeister for force pushing at a main branch [19:34:56] it’s my repo 😤 [19:34:57] but yeah, the pipeline should probably handle that a bit better [19:36:01] (in the original version of the commit I had failed to notice that the .gitlab-ci.yml already had a `stages` section and just added a new one with test+deploy, but then the old/second stages section apparently overwrote that which in sum resulted in an invalid CI config) [19:44:42] taavi: https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/100 [19:45:00] if you want to quickly review/merge it, I could test it with a rebuild in QuickCategories right away [19:45:20] rebuild in *wd-image-positions [19:45:38] (otherwise testing it would require some tedious dance of switching the repo to use my fork of the CI config, pushing that, then amending the commit and force-pushing to actually trigger the bug… ^^) [19:46:41] is using a variable assignment as an if condition valid bash(?) syntax? [19:46:48] I think so yeah [19:46:53] (and yes this script is very bash) [19:47:10] fair enough, and that's easy enough to revert if it goes wrong [19:47:12] (lol, also I got an email that toolforge-repos/wd-image-positions failed to mirror to lucaswerkmeister/tool-wd-image-positions because force-pushing isn’t allowed *there*) [19:47:14] (point taken) [19:47:31] and I think bash is smart enough to treat `if false` as an exception from “exit on error” [19:47:42] done [19:48:13] if a=$(echo hi; false); then echo what; else declare -p a; fi [19:48:14] yay thanks [19:48:35] I clicked “retry”, let’s hope it refetches the latest version of the CI config [19:48:48] nope [19:49:19] dangit gitlab [19:49:53] it does not [19:50:40] I seem to recall reading in the docs that there are options to retry a whole pipeline (refetches included configs) and only one job (does not), but I can’t find the former :/ [19:50:55] (I can run the pipeline manually but then the old hash will just be 000000) [19:52:23] I guess I can… force-push another new commit :blobfoxthinkgoogly: [19:54:52] done, let’s see… [19:56:08] okay, I think the fix worked, https://gitlab.wikimedia.org/toolforge-repos/wd-image-positions/-/jobs/971227 shows that it used “323162c915..c00bcb6: Migrate to Toolforge Components API with push-to-deploy” as the description [19:56:22] (the different length of the first hash shows that this comes from the fallback code) [19:56:25] lucaswerkmeister: uhh :D it also emails me (and I guess all other admins of /toolforge-repos) of the mirroring failure [19:56:37] oh no! [19:57:52] (also, the push-to-deploy then failed with "Unable to find namespace tool-wd-image-positions or config wd-image-positions-config for wd-image-positions", which is probably the expected result when no previous components service setup was done) [19:58:28] is that (emailing toolforge-repos admins on mirror failure) worth reporting / fixing? (so it only bothers the actual tool maintainers?) [19:58:50] (for the second force-push I disabled the branch protection in the target repo first, so that time the mirroring worked ^^) [20:08:55] !log deployment-prep add deployment-jobrunner06 to 'jobrunner' dsh group - T438656 [20:09:01] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [20:09:02] T438656: Beta cluster jobrunner opcache is not updated during automated deployments - https://phabricator.wikimedia.org/T438656 [20:15:23] !log lucaswerkmeister@tools-bastion-15 tools.wd-image-positions deployed c00bcb6310 (push-to-deploy) and 550f1df480 (upgrade dependencies) [20:15:26] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.wd-image-positions/SAL [20:18:59] !log lucaswerkmeister@tools-bastion-15 tools.wd-image-positions deployed b5a93138aa (Python 3.14) [20:19:00] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.wd-image-positions/SAL [20:20:34] push-to-deploy worked fine for those \o/ [20:26:36] lucaswerkmeister: tbh for now we can solve the mail problem with 'taavi will yell at people who force push in a way that sends those emails' unless somehow everyone sets up mirroring in gitlab to repos that block force pushing and then starts force pushing [20:27:21] hehe, fair [20:53:37] !log deployment-prep remove puppetserver cherry-pick for T428052, running newer version of HAProxy nowadays [20:53:42] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [20:53:43] T428052: Beta cluster haproxy does not support `warn-blocked-traffic-after` keyword - https://phabricator.wikimedia.org/T428052 [20:58:14] hacks-- [21:20:29] !log deployment-prep add deployment-deploy06 to scap::dsh::scap_masters and deployment_hosts - T435393 [21:20:35] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [21:20:36] T435393: Beta Cluster: Switch to PHP 8.5, includes writing Puppet changes to support PHP 8.5 on baremetal - https://phabricator.wikimedia.org/T435393 [21:22:57] !log deployment-prep decommission deployment-cumin-3 - T436470 [21:23:02] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [21:23:03] T436470: Replace deployment-cumin-3 (Bullseye deprecation) - https://phabricator.wikimedia.org/T436470 [22:36:39] tbh for now we can solve the mail problem with 'taavi will yell at people who force push in a way that sends those emails' unless somehow everyone sets up mirroring in gitlab to repos that block force pushing and then starts force pushing [22:36:41] don't give me ideas [22:40:19] !log deployment-prep create empty /etc/helmfile-defaults/mediawiki/release on -deploy04 and -deploy06 to unbreak Scap sync-masters - T435393 [22:40:25] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [22:40:26] T435393: Beta Cluster: Switch to PHP 8.5, includes writing Puppet changes to support PHP 8.5 on baremetal - https://phabricator.wikimedia.org/T435393 [22:41:03] Also, yesterday I finished setting up a brand new app for the first time on the new toolforge infra (I have another one from 2-3 years ago that I haven't migrated to the new infra), if you have anything you want feedback on specifically. Not sure if that'd be useful, but figured I'd ask. It's a web app (and also only stored on github for now) so I didn't do the [22:41:03] gitlab push-to-deploy yet. [22:43:46] !log deployment-prep add wmgRedisLockPassword to PrivateSettings.php - T436480 [22:43:52] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [22:43:53] T436480: Replace deployment-rdb01 (Bullseye deprecation) - https://phabricator.wikimedia.org/T436480