[07:02:10] marostegui: btw, for main stash data (msX), now there is a weekly reporting dashboard https://grafana.wikimedia.org/d/mainstash-resident-inventory/mainstash-resident-inventory?from=now-30d&to=now&timezone=utc&var-datasource=000000026&var-cluster=$__all&var-keygroup=$__all [07:02:22] T430940 [07:02:23] T430940: objectcache: Enable teams to measure their load in production by adding resident-inventory measurement for DB-backed MainStash - https://phabricator.wikimedia.org/T430940 [07:02:34] Amir1: will check thanks! [08:05:11] dear Dell, why does the first hdd in this new system appear as /dev/disk/by-path/pci-0000:50:00.0-scsi-0:3:86:0 ? [08:07:26] (as opposed to pci-0000:50:00.0-scsi-0:2:0:0 which is what we've had before) [08:09:21] Emperor: update firmware and bios please [08:11:32] XD [08:11:52] [I mean, I can live with the stupid names, it just makes the hosts.yaml file more unwieldy] [08:12:17] I presume they've changed the disk controller again, and/or its wired in differently, but all 3 new nodes have this set of by-path entries [08:13:30] Emperor: Maybe the disk is attached in a different bus? [08:13:46] note 2 vs 3 in the bus address [08:14:10] so maybe they are wired in a different way? [08:14:14] I am just purely guessing [08:15:21] yeah, I'm assuming they've wired them up differently. I am trying not to get rabbit-holed into _why_ [08:17:48] yeah, and also the problem is what to expect on each host [08:20:18] The swift_disks fact is quite agnostic to this, I'll just need to tell the ring manager what the resulting paths are on these new hosts. [09:06:32] oh darn it, this pontoon stack has exploded again, I only rebuilt it yesterday :( [09:09:25] FIRING: SystemdUnitFailed: check-private-data.service on db1270:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:09:43] woooot [09:09:44] checking [09:17:40] Found the issue, it was expected as that db1270 doesn't have x4 yet [09:18:09] I will clone it now so I can clear the alert for the weekend [13:58:53] Emperor: do you have any suggestions aside from duplication the accounts in terms of getting the Swift accounts in profile::mediabackup::worker (https://gerrit.wikimedia.org/r/c/operations/puppet/+/1343056) - the style guide rejects it and I was trying to avoid having them in 2 places [14:05:12] cezmunsta: so I _think_ the usual idiom would be to have the lookup at the top of modules/profile/manifests/mediabackup/worker.pp (where you already have some lookups) and then pass resulting variable into the call to class { 'mediabackup::worker' [14:05:55] but this is the sort of puppet-wrangling I don't do very often [14:10:38] Thanks, yes I will see if that appeases it for now. If it doesn't then my thought was perhaps something like a Swift client role+profile that could be used to avoid too much effort [14:27:49] It did indeed suffice [14:54:24] Hello. I added this patch to fault-tolerance for your kind review, when convenient: https://gitlab.wikimedia.org/repos/data_persistence/fault-tolerance/-/merge_requests/8 Thanks. [14:56:45] btullis: I've added Amir.1 as reviewer to that, since he knows the code best [14:58:20] Emperor: thanks, I was just typing that it might be best for that :D [14:58:36] It is the first time that I have seen that specific repo [15:04:21] Thank you. I didn't have the rights to select a reviewer for it. [23:00:16] hi data persistence! oncall here, heads up that sessionstore1005 flared up again (https://phabricator.wikimedia.org/T437915) so we depooled sessionstore from eqiad. we'll leave it there for the weekend but someone probably needs to have a look sooner than u.random is back from OOO