[02:10:39] c.danis: I do to; I mentioned earlier that I've asked for an additional node per datacenter several times in the past, and folks balked at the cost (hardware, but also space in the DC). And this was before the current cost inflation. [02:12:08] so I think the next step is to put together an SLO, an honest one that takes into account the lead time for serious hardware repairs, and see if that doesn't sway opinions. [02:13:21] because I think the difference between expectations of availability, and what is realistic will be...large [04:29:05] <_joe_> I will go in the opposite direction - we're running sessionstore with more or less the same expectations of redundancy as a lot of components that are fundamental to our operations, like the main etcd clusters [04:29:57] <_joe_> so we can tolerate losing one node at most. This is a risk we've always accepted; if we're not ok with it, then let's talk. [06:29:36] <_joe_> I am going to switch shellbox-timeline to use gvisor in production, starting from eqiad [06:40:33] <_joe_> https://trace.wikimedia.org/ is down atm, I guess because of problems with pki? [07:17:08] <_joe_> elukey: ^^ [07:17:29] <_joe_> I think there's been a few changes that weren't reflected in applying them to the k8s clusters [07:27:20] _joe_ see my email to sre-at-large, wikikube and ml clusters have a pending change that I will roll out later on in the morning [07:28:02] but it is not that, I think it is the same issue that happened a while ago - jaeger's setup doesn't automatically reload certs, so when a new one is issued it is not picked up [07:28:27] so usually the "fix" is just to kill/restart jaeger pods [07:28:30] lemme do it [07:29:12] <_joe_> yeah the cert is expired, just found out :P [07:36:43] done! [07:36:56] the cert for jaeger query IIUC [07:37:04] killng the pods "fixes" it [07:42:31] I don't recall if there is already a task about, I think so, maybe it is a matter of upgrading jaeger or similar [07:42:59] the whole setup hasn't been touched in a long time, maybe time to prioritize it during the next quarters [07:44:14] (need to run an errand, will be back in a few, I'll start working on pki after that) [08:40:58] notice for all: eqsin temporary depooled for T438052 [08:40:59] T438052: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052 [09:12:28] all k8s clusters have the pki patch applied! [10:20:32] <_joe_> elukey: thanks! [12:51:39] !log eqsin pooled again (T438052) [12:51:42] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:51:43] T438052: Jin onsite work for 2026-09-16 @ 09:00 UTC - https://phabricator.wikimedia.org/T438052 [12:55:09] slyngs: is https://gerrit.wikimedia.org/r/c/operations/puppet/+/1342135 ok ? [13:08:56] dpogorzelski: +1 [13:10:13] thx [14:21:57] _joe_: I think we're saying the same thing. I want to create an SLO that makes explicit what the expectations are given the current redundancy, and if that's OK, then we're good. [14:22:31] but I think everyone, myself included fills in that missing SLO with a whole bunch of 9s [14:36:01] urandom: that is a good idea and it is inline with our efforts to add SLOs :) [14:38:09] <_joe_> at the same time, how do you account for the random hardware failure? [14:38:18] <_joe_> and, I'd use the etcd slo as a comparison [14:38:41] <_joe_> basically, how do you evaluate the probability of losing two servers in the interval between the first failing and its replacement arriving? [14:40:13] it would also help if our servers were more fungible [14:40:34] <_joe_> yes [14:41:28] I think hw failures and SLOs error budget are two separate things, but they can inform each other. If we aim for session store's availability at 99.9 target then we also need to structure the hardware in a way that the error budget can be realistically sustained with the loss of one node [14:42:33] if you loose even a DC, and we have a single SLO (like I think we should, since it is the pov of the clients) the target should be tailored to support the use case [14:42:42] not super easy but doable [14:44:53] no idea if session store could be dc agnostic or we need to have two separate SLOs, but the main point remains [14:47:04] re : how do you evaluate the probability of losing two servers in the interval between the first failing and its replacement arriving? - I look at the issue from a different angle - Given an SLO target, what should we do to avoid this exact situation to burn the whole budget? [14:47:40] if the answer is we can't because no more hw etc.. we may think about lowering the SLO target expectations [14:48:06] these are my 2c, lemme know your thoughts :) [14:52:46] <_joe_> yeah the "no hw expansions" is kind of a given [14:53:49] the way I was thinking of it was that whatever the probability of losing a node in any given moment, is what I'd continue to apply to the period when we're already down one node. so it feels like it's a remote possibility to lose two, but I don't think it is (especially if you're running in that state for literal days). [14:55:13] <_joe_> that's not how independent probability is calculated :P [14:55:34] and again, I'm not saying we have to get more nodes; it may just be an incongruity between what we're willing to risk, and what I perceive we are [14:56:25] I'm exhausted ever time I shift into high anxiety mode trying to restore redundancy, when it's really not within my power to do so [14:56:42] so maybe the answer is that I shouldn't worry about it [14:56:59] j.oe I think he was hinting at the fact that with one host down the hosts are more under pressure and hence might have a higher probability of failure than in normal redundancy mode [14:57:01] (so much) [14:57:14] urandom: don't worry about it! [14:57:20] * volans sends hugs [14:57:22] boom, there we go! [15:01:02] the rule of thumb that I use with SLOs is - calculate over a window of time (like days in 3 months) you can stay down when you aim to a certain SLO target. It is a high level approximation, but for example 99.9 means 0.1 of error budget, so ~0.09 days of downtime. Is that doable with the current setup? And this means hw redundancy, etc.. if the answer is "probably no, we'll burn the error budget in these cases" then we may need to revise the [15:01:02] target [15:01:23] *how much you can stay down [15:02:30] also I agree, urandom shouldn't go in high anxiety mode for these things :)