[07:29:30] Congratulations! Your RIPE Atlas probe 7261 is 3 years old! In the last 3 years, your probe was connected for 99.460000% of the time. [07:29:30] That's esams VM atlas anchor - https://atlas.ripe.net/probes/7261/overview - quite a nice uptime knowing that the current VM the service runs on is only 1 year old (we did a migration last september). [10:10:10] Hii, I have a question about docker base images [10:10:45] Where does the debian base image (e.g. https://docker-registry.wikimedia.org/bookworm/tags/) comes from? Is it imported from upstream, or do we build it locally somewhere? [10:16:13] <_joe_> we do build them, at the moment on build2004 [10:16:35] <_joe_> is there any problem you're seeing? [10:17:50] <_joe_> atsukoito: https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/production/modules/docker/files/build-bare-slim.sh is the script that builds the images [10:18:55] thanks! [10:19:46] Well, it is not really a problem, bookworm image has now added a snapshot, and it makes the image not able to do apt-get update after a week or so [10:20:10] brouberol told me to check #-security [10:21:13] atsukoito: I opened a ticket today, which I think is this issue T437836 [10:21:14] T437836: "apt update" failing in CI, on both bookworm and trixie, and both gerrit and gitlab - https://phabricator.wikimedia.org/T437836 [10:21:30] <_joe_> there were quite a few tasks already, please talk to moritzm [10:21:44] <_joe_> but also, base images are rebuilt weekly at most [10:21:46] (which I think is distinct to the bullseye failure) [10:22:40] s/to/from/ [10:23:31] ahh, i see, it is a known issue [10:23:48] <_joe_> yeah that's why I was asking :) [10:23:49] i'll better adjust my CI, as it is totally doesn't need apt-get in it [10:24:24] _joe_: it is exactly as in the task Emperor mentioned [10:24:25] <_joe_> but yes, I see the snapshots are creating a few issues. We might want to ensure they don't get created [12:55:11] FYI, I plan to start some sessionstore reimages shortly (if there are no objections). No impacted expected. [13:17:41] FYI, I'm rebooting cumin1003 [13:23:25] and it's back [17:51:32] mutante: thanks for all the pings on the access approval requests <3 [17:59:19] sukhe: very welcome! [19:42:32] if a reimage cookbook were started, then killed, but its lock remains, is there an alternative other/better than waiting it out or running with `--no-locks`? [19:45:33] I suppose deleting the lock file directly. [19:46:20] it's in etcd I think [19:46:31] I mean, same principle, but that sounds awful [19:50:36] why not use --no-locks [19:52:22] no reason in this case, I was just wondering if there was a less hammer-y way [20:20:07] --no-locks means your run can start right away, which is good, but it also means your run won't prevent another one from starting, which is bad [20:20:22] probably better to wait for the TTL or release it in etcd [20:20:49] (probably low risk if you're confident nobody else is going to run a conflicting reimage, though) [20:21:00] yeah, I mean, it's foregoing a safety mechanism, hence the "hammer-y" description [20:22:16] yeah, just calling out that it's not only ignoring the max_concurrency (which is totally fine if you know the existing lock is bogus) it's also not going to take out a lock itself [20:23:02] Also, by way of update on the sessionstore reimages: I am stuck on sessionstore1005 because reasons. It's been down for a few hours now, and there is no ETA. With it down, there is no redundancy in eqiad, if there were a failure of any kind it would be impacting. [20:24:30] The "fix" (to the redundancy loss) would be to de-pool eqiad, but that would add cross-DC latency for some, so I don't know whether that's better. [20:25:52] maybe if you elaborate on the reasons part and point out how it's more urgent than standard installs you can get that fixed faster (infra security? dcops?) [20:26:36] dcops is working on it. https://phabricator.wikimedia.org/T437915 [20:28:58] it started out with a cryptic hardware error after rebooting it for the reimage. after getting past that (with no subsequent hardware errors reported), it completes an install, but doesn't make it to the grub prompt. [20:29:01] https://usercontent.irccloud-cdn.com/file/bkhxPGvP/image.png [20:29:04] ah, real hw failure. gotcha.. well then you already did all you could and dcops has to install latest firmware after which Dell will send replacement. :p [20:29:09] should be under warranty [20:29:31] that's where I am now, just hung there, no errors... [20:29:55] yeesh, how many weeks does it take them to send a replacement? [20:30:39] with an undeniable firmware failure under warranty - i don't think that long. but better to ask Willy [20:30:52] "Dell will send a replacement" doesn't inspire confidence [20:30:54] eh, I meant mainboard failure [20:31:37] there are ways to get smart hands to do things quickly [20:31:51] also, we're probably at the point now that I should /cc: wiki_willy [20:32:36] robh: maybe you have some general input on turn around times for mainboard replacement [20:33:03] reading backlog [20:33:05] Dell ,R450 [20:33:13] Error: Backplane 0 [20:33:32] Its under warranty, John posted on task that he is opening a case [20:33:41] its usually a 1-3 days [20:33:43] business days [20:34:10] there you go, urandom. not that bad [20:34:15] John will open case, hopefully he provides the TSR at the time of the case opening or it will take a back and forth to upload that, then they'll figure out whats worng [20:34:20] but this doesn't appear to be mainboard [20:34:22] its a backplane issue [20:34:41] the backplane is where all the hotswap drive bays slot into [20:34:55] which connects to mainboard with a cable (or two depending on number of bays) [20:35:01] so its WAY easier than mainboard swap [20:35:12] minutes once its onsite. [20:35:26] ah, assumed too much that it's all in one. that's good new [20:35:29] news [20:36:01] yup, ive seen odd number of backplane failures aross various models recently [20:36:02] its odd.... [20:36:21] so very much like that one i had an out of warranty host in esams get wonly and have to have the cables reseated and so far its ok.... [20:36:51] urandom: i dont know why you would cc willy yet if the case is opened by john today? [20:37:05] but you can plus we've just pinged him i bet ;D [20:37:29] it seems the hosth ad 2 other failures. i haven't read full backkog [20:37:40] re deciding whether to depool, note we'll be depooling everything from eqiad next week ahead of the dc switchover too, including sessionstore -- so it sort of becomes moot at that point, which is sort of good news [20:38:20] oh, i suppose you eman in regards to the fact it has happened across the fleet maybe... cuz eyah [20:38:35] it seems we dind't open a case the last time this single host had the issue, reseated cables and flashed firmware returned to service [20:38:45] but since its now happened again, john opening a case is the right call. [20:39:06] also dc ops recently instigated opening a case for every single hw error even if we clear it with manual intervention [20:39:13] liek in our last dc ops meeting... [20:39:40] we're also currently working on a way to more accurately track hardware repairs fleetwide to detect patterns and see if we have issues with any given config but thats very recently begun =] [20:41:40] rzl: not really though. I mean, it inverts the traffic balance I think, but it's otherwise the same. [20:42:30] both DCs are getting requests, and those landing in eqiad are at risk if there is any sort of failure there (transient or otherwise) [20:43:21] rzl: unless by everything you mean *everything*(?) [20:43:39] do we normally do that? [20:43:43] we do, yeah [20:43:52] for how long? [20:44:07] for a week -- partly capacity validation and partly to allow for dcops/netops/database maintenance [20:44:18] oh ok, that is sort of good new [20:44:21] news [20:44:34] (we're looking at potentially making that period shorter in future switchovers, depending on the actual maintenance needs -- it's not great for the latency story, for the reasons you point out) [20:44:36] sessionstore1005 seems to be a lemon. it's like the 3rd time and "Dell reviewed logs and found no prior cases on record" but we certainly do [20:45:02] because I am out next week, and out this friday, so if this doesn't get resolved by thursday, I won't be around [20:45:04] that also doesn't mean we don't care about sessionstore in eqiad not having redundancy -- the reason it's a good test is that we can undo it if we need to, and we sometimes do [20:45:10] mutante: where was that? [20:45:30] robh: older tickets https://phabricator.wikimedia.org/T398225 https://phabricator.wikimedia.org/T421297 [20:45:36] (sorry for the triple negative, lol. we do still want to have redundancy) [20:45:59] Ok, so best case if they don't argue and send a replacment backplane tomorrow its going to be down until wednesday at the earliest and i dont htink its to that point [20:46:17] mutante: thanks, i'll paste that speciifcally into the task for john to see so he doesnt have to potentially miss from past tasks [20:46:26] robh: sounds good, thanks as well [20:46:39] i know its in the task summary but not the old case numbers [20:46:42] its best if he has those heh [20:47:35] ok, task edited to make that clear [20:47:56] so yeah, best case it gets swapped mid week, that sounds like a failover to codfw as primary for the sessionstore since its not tomororw? [20:48:20] back-of-napkin, but I typically try to keep maintenance down to low single digit $hours, so $days is disconcerting [20:48:29] the fact it had a previous dell case makes it easier to get hw replacement approved [20:48:38] where they can see we updated firmware then and its happened again [20:48:40] * urandom can hear rzl say "something something SLO" [20:48:48] yessss my work is half-done [20:49:02] well, lets say it was an average memory failure and next day dispatch cuz troubleshooting is easy [20:49:11] its still next day dispatch for easy repairs [20:49:15] so its always days. [20:49:35] You may wish to recalibrate repair times on your calendar =] [20:49:57] I've asked for a 4th host in each DC [20:50:07] the answer has thus far been no [20:50:32] thats above my paygrade i was just letting you know the best case resolution times for a hw repair [20:51:06] we have a host we dont need and are about to give back [20:51:09] while i am the person ordering hte hosts, i go off a set list i dont control or modify ; D [20:51:18] but that's a R470 - vs 450 [20:51:23] ConfigB [20:52:21] sessionstore is a config a, but pelae note you cannot just hand them between teams as there are budget allocions and stuff =]. (it can happen but that is when you get wil ly involved; ) [20:52:29] typically its a rubber stamp yes. [20:52:44] its better a host go into prod use for another team than sit in the rack doing nothing and depreciating [20:52:58] yes, needs to go offical way but this is the host in question to apply for if of any use: https://netbox.wikimedia.org/dcim/devices/6688/ [20:53:07] so, an old config b going into use where an addtional config a is needed could be useful [20:53:17] not old [20:53:40] well we also have many many config b on the actual budget for order [20:53:52] so this may have to go to soemthing esle too but i hav ezero clue thats a mgmt call [20:54:06] i was just seeing if it was a 1:1 or in this case 1:more [20:54:19] with a name like sessionstore my initial fear was many data disks ;D [20:54:39] which makes it having a failing backplane even more annoying its 2 OS disks and thats it! [21:15:47] urandom: if the lock persisted the cookbook it means that it was killed "badly", for normal single ctrl+c (not multiple) or kill (not -9) the lock should get cleared on exit (if not that's a regression) [21:16:35] volans: I figured that was the case [21:19:58] there have been discussions in the past if we should wrap ctrl+c to prevent accidental double ctrl+c but at same time without preventing to be able to abort a cookbook quickly in case that's needed. But no decision/prioritization/implementation on that. [21:21:07] I guess we could add the docs on how to remove a single lock on wikitech. I have an old branch with a cookbook that shows the current locks, but it's RO [21:23:10] to be fair: this is the first time I've ever had it be an issue, so it's probably not a huge issue [21:23:51] I waited for it to expire, and the world turned on :) [21:24:04] at least as much as it was going to... [21:25:28] :D [22:27:11] Ok, so work on sessionstore1005 continues, but hope wanes, and so it begs the question of what to do next. Everything is fine for the time-being, but we have no redundancy in eqiad, so any sort of issue with the remaing two hosts and we have an incident. We could restore redundancy by depooling eqiad and running entirely out of codfw, but requests otherwise landing in eqiad will incur the additional latency. [22:27:39] I'm inclined to chance it, and make everyone aware of what to do in the event something does happen. [22:27:52] Thoughts anyone? [22:31:30] if others can monitor the service and respond with a failover to codfw (if another host in eqiad presents problems) then to me that would be good [22:31:58] but that's my cheap and less relevant two cents :) [22:32:50] It means reacting after the fact, after we've had an incident, but saving everyone the added latency if nothing bad happens [22:33:09] but probably nothing will happen [22:33:12] * urandom knocks on wood