[09:47:20] hey folks, as a reminder I am going to get the scap lock in ~10 mins to proceed with the maintenance on the Docker Registry (Dropping old images) [10:03:11] all right cleanup started [10:03:29] this is just for the "Restricted" bucket, the one holding mediawiki images [10:03:33] the rest is not touched [10:05:28] it should take 2 to 3 hours, I'll update once done [10:05:40] Cc: Emperor [10:06:01] Empero.r is out [10:12:18] marostegui: I am sorry, didn't read netbox, I was too focused on my cookbooks [10:13:30] elukey: no worries, if you need anything cumin is always there for you [10:14:52] marostegui: I prefer to consult my spicerack cabinet to be honest, but YMMV [10:16:03] hahaha [10:16:40] ahahah :D [12:06:41] quick update on the mw images drop - I had to restart the mark phase (as in, mark and sweep) because I added a new option that caused it to take a huge amount of time (--delete-untagged FTR). So I am ~1 hour behind schedule, I hope that the second mark will finish in ~15/20 mins so that deletes will start happening [12:33:36] update: mark phase done, now the GC is very very slowly emitting all the blobs to delete (taking ages sigh) [12:33:54] but mark completed in ~1 hour this time [12:34:12] for ~700GB is not super bad, considering the things to check etc.. [12:57:39] ok it is dropping but not in a very fast way [13:01:12] I updated #engineering-all on slack, I need more time and the backport window for MW starts now [13:31:56] completed! [13:32:02] ~400G dropped \o/ [13:32:14] \i/ [13:35:02] now the trick is to do it periodically so we can keep the restricted/mw bucket on apus small [13:57:29] We will do some maintenance on GitLab CI. The cloud runners will be unavailable between 15:00 to 16:00 UTC [13:57:30] heads up for the on-callers - I am reverting a change for pki, I tried to simplify the httpd config but it didn't work as expected [13:57:41] so some puppet failures etcc may happen while I revert [14:06:59] ok! [14:14:47] update: I am going to deploy https://gerrit.wikimedia.org/r/c/operations/puppet/+/1352109 that is a better fix [14:21:07] oook all should work now, lemme know otherwise [15:57:24] GitLab maintenance is done, CI is available again [17:20:01] Anybody with any Swift knowledge here in Empero.r's absence? [17:25:18] cezmunsta: once upon a time I knew how to edit the ring files [17:29:17] before or after it was changed to https://wikitech.wikimedia.org/wiki/Swift/Ring_Management and https://gitlab.wikimedia.org/repos/data_persistence/swift-ring ? :-P [17:30:01] neat! 📸 [17:30:04] cdanis: so, a disk appears to have died and based upon the docs and a previous commit, it seemed as though adding a device as failed under the host should remove it... leaving the timer and service to take care of things. [17:30:04] I am not sure that it has though, a dry run showed "Would add 0 host(s), remove 0, change 0 weights" and I still see the requests coming though -> 507 [17:31:35] This appears with verbose: swift-dispersion-report stderr: ERROR: 10.64.132.24:6014/objects12 is unmounted -- This will cause replicas designated for that device to be considered missing until resolved or the ring is updated. [17:35:05] cezmunsta: I have to go shortly, though by looking at https://gerrit.wikimedia.org/r/c/operations/puppet/+/1062355 the hostname needs to go from a string to a dict too [17:36:28] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1352180/ looks like this one didn't get merged? [17:36:56] oh, but it got live-hacked in, it looks like [17:37:22] ok! then if it isn't that I don't know :( [17:37:24] godog: thanks, yep I spotted that - and cdanis that is the fix .. and yes as the timer beat me to it and I spotted the error [17:37:54] sweet! ok, gotta go [17:42:15] cezmunsta: so it looks like the timer ran successfully but also reported no changes? [17:42:45] yep and the same for the run that I did to make sure after fixing the typo. [17:42:45] [17:47:23] is it just me or does --verbose not seem to work [17:47:54] ... [17:48:05] it has `logging.debug(s)` where it needs logger.debug [17:51:04] Did you use --syslog? [17:52:05] You can use "logging.debug" it just means that you are getting the root logger [17:53:17] so: " sudo /usr/local/bin/swift_ring_manager -o /var/cache/swift_rings --verbose" does show a debug message [17:55:00] "sudo swift-recon --unmounted" showed nothing aside from an initial message and eixited with 0 [17:55:53] The partition is technically still mounted on the remote though, but XFS shutdown and I was going to unmount it once the device was out of the ring [18:32:45] I need to go now, so aside from seeing the errors for the proxy-server and on the remote too, I am presuming that things are all still working. I will leave Puppet disabled on ms-fe1009 pending the MR merge and off for longer on ms-be1074 as it just fails due to the disk issue. [18:35:24] cezmunsta: I figured it out [18:35:55] * cezmunsta returns [18:36:11] cdanis: *\o/* [18:38:03] if you call it `sdn1` then it applies to nothing ... because you have to match the failed disk to the failed prod24_ng object name, which from `mount|grep sdn` on the ms-be host, is objects12 [18:38:10] Would set weight ms-be1074/objects12 in object to 0 [18:38:12] Would add 0 host(s), remove 0, change 1 weights [18:38:47] the script could probably use some more error checking and footgun prevention [18:39:13] OK, let me update that MR [18:47:40] I wrote https://gitlab.wikimedia.org/repos/data_persistence/swift-ring/-/merge_requests/22 [18:48:44] and btw that ring_manager output above was a dry run, I didn't push it for real [18:50:26] Thanks for spotting that! I was too busy trying to figure out my way through the rest to stop and check if the docs and previous MR were misleading [18:50:26] Yep, that's fine, the timer is due to run in 20 minutes. I will hang on until then as I am still tailing the log [18:51:22] I may as well get the MR merged and then enable Puppet after testing that [18:56:51] OK, that's done and Puppet is enabled again