[11:30:19] swfrench-wmf: I was wondering if this rings a bell, after the etcd operation, it now takes a lot longer for dbctl to complete any command eg: https://phabricator.wikimedia.org/P95944 [13:29:01] <_joe_> marostegui: try the operation from cumin in codfw [13:29:06] <_joe_> and see if it's faster [13:29:16] <_joe_> assuming now the masters are in codfw [13:44:38] Is there any strong feeling whether we should or shouldn't be running fstrim? I just noticed it's failing on an EFI host due to lack of vfat support for trim. I'm not finding anything in Puppet, so just wondering if anyone's seen this before and has any recommendations [13:49:20] ah, per https://manpages.debian.org/testing/util-linux/fstrim.8.en.html looks like there is an `X-fstrim.notrim` option [14:10:17] marostegui: ah, interesting! yes, I wouldn't say that magnitude of slowdown expected, but is foreseeable for any operation that makes a large number of round trips to etcd that are now going over the WAN. [14:10:17] so, precisely what _joe_ said, and indeed the same config-diff operations takes less than 400ms if executed on a cumin host in codfw [14:12:10] FWIW, we'll be switching back in the next week or two in order to unblock some new hardware integration work in codfw. in the meantime, though, if you're able to use a cumin host in codfw, you should see the latencies you're used to. [14:31:50] _joe_ swfrench-wmf ah I see! Thank you both [14:36:20] marostegui: _joe_: so, there's another reason why cumin1003 is *much* more impacted by this than expected ... it appears to still be running conftool 6.0.1, which does not include a whole bunch of optimizations that landed in 6.1.0 [14:36:48] those cut down massively on the number of round trips [14:37:01] <_joe_> uh heh [14:37:07] <_joe_> please update it [14:37:42] and concerns about doing so ... on a Friday? (I'd say no, but ... Friday) [14:38:43] swfrench-wmf: I'd say we can wait if you feel more comfortable doing it next week. If there's anything super urgent we can just use cumin codfw [14:40:37] <_joe_> zero i'd say [14:40:59] <_joe_> we've been using 6.1.0 in various places for quite some time right? [14:41:02] yeah, I'm not terribly worried about it, since we've been using the same package on 2002 for months [14:41:07] exactly yeah [14:41:14] * swfrench-wmf will do momentarily [14:47:02] marostegui: _joe_: much better: https://www.irccloud.com/pastebin/h9cWmYcZ/ [14:47:59] I knew the impact of https://gitlab.wikimedia.org/repos/sre/conftool/-/merge_requests/102 was going to be impressive in this case, but wow :) [14:48:00] <_joe_> swfrench-wmf: oh wow *that* much better? [14:48:31] Nice!!!!! [14:48:34] <_joe_> I gotta say it was a low-hanging fruit [14:48:43] yeah, you can see the sheer number of round trips on 6.0.1 with `--debug` - it needed to do *so* many [14:48:58] <_joe_> but also when I first wrote conftool it was supposed to do much less than it ended up doing :) [14:49:04] heh [14:49:25] That's usually the case here xdd [16:02:49] swfrench-wmf: fantastic, I had meant to fix that too [16:02:55] for years and years [16:03:13] once we had a dbctl alert start failing after we moved etcd because 10 seconds wasn't long enough anymore to run diff [16:03:21] all just because of rtts to etcd [16:03:21] I'm glad _j.oe_ did for us, then :) [16:03:24] yeah :) [16:03:31] haha, that's amazing [20:04:22] puppetserver1001 is unwell, due to some Java version mismatch. I'm investigating... [20:05:06] "Incorrect Java version: 17.0.18+8-Debian-1deb12u1" [20:11:18] updated (I assume by unattended-upgrades) a couple of hours ago [20:32:51] andrewbogott: that is my fault, as I performed the upgrade, happy to deal with the consequences [20:33:37] jhathaway: I'm restarting all the puppetservers, that's likely all we need. [20:34:19] ah we found a scapegoat! [20:34:25] I was just starting to dig in [20:34:38] goat indeed [20:45:44] 🐐 [20:55:57] I am off for a month on PTO and will quit IRC today. But will be back early September. cya [20:58:06] cya! [20:59:54] o/, enjoy! [21:11:44] thank you:)