[12:59:13] !log project-proxy drop metricsinfra redis job T429930 [12:59:17] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Project-proxy/SAL [12:59:17] T429930: Convert dynamicproxy from nginx to haproxy - https://phabricator.wikimedia.org/T429930 [13:28:38] Just a heads-up that I'm starting the clouddumps1002 reimage shortly, ref T403154 [13:28:39] T403154: Upgrade clouddumps hosts to bookworm/trixie - https://phabricator.wikimedia.org/T403154 [13:30:28] ack, thanks and good luck [13:50:01] So far so good, it's trying to run Puppet now [14:46:29] !status toolforge kubernetes upgrade in progress [17:01:29] inflatador: clouddumps1002 seems to be pooled already but it doesn't have /srv mounted (and so the fetch jobs have filled the root partition) [17:03:17] taavi apologies, I wasn't given any instructions beyond "reimage the server". Looks like you took care of the depooling? [17:04:07] uh, no, the server is still very much in service and receiving traffic :/ [17:04:23] (fixed) [17:04:41] ah, was looking at https://config-master.wikimedia.org/pybal/eqiad/dumps-nfs [17:05:03] I just rm'd everything in /srv/, will run puppet again [17:05:13] nfs is one of three services running on the host, -rsync and -https were still pooled=yes until a few moments ago [17:05:24] isn't there supposed to be a volume mounted on /srv? [17:06:09] well...shit [17:07:15] `lsblk` shows a correctly-sized volume as sda1, which the installed would have just ignored on an existing machine, I believe you just need to dig the fs uuid (I believe via some `lsblk` option) and add it to fstab? [17:07:50] yeah, will take a closer look [17:08:24] I thought this was gonna be more hands-off, that was probably a stupid assumption on my part. I'll start looking at docs and comparing to dumps1001 [17:14:24] you are right on re:fstab, should be an easy enough fix. These guys don't seem to have their own partman recipe, I'll try and get a patch up for that later [17:18:38] !status ok [17:20:43] taavi OK, puppet ran cleanly and I confirmed that the host can properly mount its 196T volume via PARTUUID entry in fstab...anything else we need to do before repooling? [17:21:24] inflatador: if the data is there, probably not! [17:21:27] should I do it? [17:21:31] Sure [17:21:42] the data is indeed there [17:22:15] and all nfs stuff functional? exportfs, nfs service... [17:22:20] ok, running [17:22:22] taavi@cumin1003 ~ $ sudo confctl select "name=clouddumps1002.wikimedia.org,service=dumps-(rsync|https)" set/pooled=yes [17:22:37] nfs is depooled like it was before the image, i'll leave that to g.odog to handle [17:22:43] ok [17:23:05] I see some failed units related to HDFS, that is probably my team's job to investigate though [17:23:20] so how would one test https access to it ? [17:24:18] `curl --connect-to ::clouddumps1002.wikimedia.org https://dumps.wikimedia.org/other/mediawiki_content_history/labswiki/2026-08-01/xml/bzip2/SHA256SUMS` shows what I would expect [17:26:08] thank you taavi [17:33:42] I added https://wikitech.wikimedia.org/wiki/Portal:Data_Services/Admin/Dumps#Reimaging to the dumps docs. I'm not sure where the formatting weirdness came from, LMK if y'all have ideas [17:34:34] nm, gogt it [17:34:36] fixed (there was a that should have been ) [17:36:14] ACK, fixed [18:03:57] thank you inflatador [18:07:04] !log deployment-prep resolving rebase conflicts in /srv/git/labs/private [18:07:07] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [18:07:50] !log devtools resolving rebase conflicts in /srv/git/labs/private [18:07:50] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Devtools/SAL [18:14:21] bliviero np, let us know when you are ready to do clouddumps1001 [18:18:39] !log civicrm-prototype resolving rebase conflicts in /srv/git/operations/puppet on puppetserver-01.civicrm-prototype.eqiad1.wikimedia.cloud [18:18:41] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Civicrm-prototype/SAL [18:19:42] inflatador: i am guessing we'll wait for green light from godog ! [19:34:10] Cyberpower678: I'm on the verge of fixing (trying to fix) puppet on some hosts in your project. Any objection? It looks like it's been broken for quite a while, at least on cyberbot-exec-01 [23:17:04] Hey, I'm spinning up new trixie text/upload nodes and there are bound to be failures. I'll try to claim/close wmcs phab tickets as they show up [23:27:09] brett: That sounds like a note for folks in #wikimedia-releng about Beta Cluster maybe more than here? [23:28:17] oh, whoops, yes. Thank you :)