[10:02:31] lunch [12:40:41] \o [12:44:42] o/ [13:04:52] o/ [14:00:43] for some reason i thought having claude write most of the safe delete script would be easier....but now it wrote a mountain of code i have to review and trim down :P [14:00:56] :/ [14:01:20] I bet it took the "safe" a bit too seriously [14:02:08] i did a grilling session first which already reduced the scope, it wanted to re-do a bit of the expected_indices, check cirrusDumpQuery outputs to see what prod was querying, etc. I cut must of that off [14:02:12] but it still wrote ~650 lines [14:02:37] ouch... [14:04:24] Heads-up that I'm restarting CODFW for the ECS log changes [14:04:26] i let it keep the idea of having a timed stats monitor, where it checks stats for index_total, search_total, etc. with a delay between them to verify idleness, but i'm tempted to dump all of that and use the cirrus rule "if no alias, delete" [14:04:43] it's depooled at the DC level, so no user impact outside of latency [14:04:52] kk [14:05:25] ebernhardson: yes... I did this check in the the stale alias cleanup for the completion suggester but ended up dropping because there was monitoring queries running [14:05:44] and the stat check was not very reliable in the end [14:05:55] oh, i forgot about that. Yea that makes sense. Ok will kill that bit [14:06:12] the index stat check could perhaps be nice? in case it wants to delete a running re-index? [14:06:33] yea i suppose index_total is at least a better indicator. [14:06:50] maybe delete_total as well, as long as it's there [14:07:05] sure [14:47:43] hey, if anyone has time to check search-loader2002 today LMK, I reimaged to trixie yesterday, just wanna make sure it didn't break anything. Ref T433890 [14:47:44] T433890: Migrate search-loader and apifeatureusage hosts to Bookworm or later - https://phabricator.wikimedia.org/T433890 [14:51:29] looks fine to me (https://grafana-rw.wikimedia.org/d/000000591/elasticsearch-mjolnir-bulk-updates?orgId=1&from=now-7d&to=now&timezone=utc) [14:55:05] well... it does not have much to do since yesterday 21:00 so hard to tell [14:56:30] it's reporting metrics at least so it's running [14:57:01] dcausse ACK, I screwed up the ticket association, but it would have come back up yesterday around 22:16 . We can wait another day or however long it takes if that would help [15:06:13] inflatador: pretty sure it's working OK, logged there and I see the daemons running [15:37:02] codfw-psi is unhappy ATM, working on it [15:41:18] OK, it's back...I really need to get a ticket started on using dedicated masters. Will do that by EoD [16:06:20] codex gave me a nice one-liner to check opensearch service start time (well, really JVM start time but same diff) ` curl https://search.svc.codfw.wmnet:9243/_nodes/stats/jvm?pretty | jq -r '.nodes[] | [.name, ((.jvm.timestamp - .jvm.uptime_in_millis)/1000 | todate), (.jvm.uptime_in_millis/1000|tostring + "s")] | @tsv' | sort -rnk 3` [16:08:40] * ebernhardson struggles, still, with indices vs indexes. What craziness allows a word have 2 plural forms... [16:10:40] hmm, 172 emails. Yup servers are restarting :P [16:13:31] indices is better, it's valid in french :) [16:14:10] hmm, i could go with that. I often write indices, but then think maybe it's supposed to be indexes [16:16:54] in french "indexes" is the conjugated form at the second person singular of the verb "indexer" (to index), so that does sound correct to me :P [16:17:07] does *not [16:18:46] opensearch is not happy :/ [16:19:11] nope, codfw-omega is acting up now [16:19:35] Yea opensearch doesn't like when we restart masters. Our idea was to follow best practices and finally setup dedicated masters. [16:19:36] I've stopped one of the masters, let me try a second [16:19:53] we never needed them in the past, but it seems like if we are having problems we should get on the expected path [16:20:07] sure [16:20:21] yeah, at the moment we are 1 server restart away from 5-10m downtime for an election ;( [16:33:13] OK, CODFW is done and repooled [16:52:31] dinner [17:43:32] I created https://docs.google.com/document/d/16d5KZe34ldvRNJXUiUgk5Gt215fgWTVyK-x1Kz0gsmY/edit?tab=t.0 for brainstorming the master performance issues, feel free to add any ideas or respond to mine [19:46:07] sigh...i installed a synaptics specific driver for wayland...but the config util apparently needs qt >=6.4 and i have 6.2, and it made the mouse pointer move extra slow, but the scroll insanely fast...fun :) [20:01:23] get to pretend it's a few decades ago and do it all through piping into sockets, and writing config files by hand...somehow it was more fun back then : [20:03:51] LOL, somehow getting paid to do stuff changes things [20:28:17] lol, figured out what i needed to get scroll under control, but it's insane :P Basically the driver has a bunch of config options that go straight to the trackpad driver, those can all be configured over a socket (why i installed it, to get access to the fine-detail config options). But it had a default 5x multiplier on scroll speed that can only be changed on the driver command line [20:36:49] 5x multipliers are good for my pinball score, bad for scroll wheel settings [20:49:41] pretty clear view of our quorum issues here: https://grafana-rw.wikimedia.org/goto/ffu77w16xpo8wf?orgId=default [21:22:15] https://grafana-rw.wikimedia.org/goto/efu7agxac1am8c?orgId=default looking at 2 chi masters around the time of the outage . 2084 has the performance governor, 2061 does not. Not seeing much of a difference [21:27:03] :S [21:27:32] i suppose thats good though, it means we have enough capacity, it's just the software being janky [21:33:07] my trackpad troubles summarized in one graph..send to office it and will see if they have ideas. I have a feeling it's going to be either accept touch-to-click (i disabled that, this is the physical touchpad clicks), or send it in for warrenty: https://phabricator.wikimedia.org/F97197630 [21:33:10] I really need to fix these dashboards so it's easy to select multiple hosts based on cluster or master eligibility [21:33:54] I haven't had much luck with LLMs and grafana yet, but I'm pretty incapable of visualizing stuff. I wonder if I should ask it to write jsonnet instead of json [21:34:51] oh wow, that plot is pretty telling. If only they had something like this during their QC process :P [21:35:17] lol [21:35:47] it's worked fine for a couple years, my theory is that the trackpad is wearing out and is more flexible than it used to be, such that edge clicks are bending the pad enough to not click the physical switch...maybe [21:36:11] that bottom left area moving into the center is i think why i noticed, i suspect the right side is "working as expected" [21:58:45] I'm wondering if it's possible to reduce our cluster state. https://opster.com/guides/elasticsearch/capacity-planning/elasticsearch-large-cluster-state-post-mortem/ has some suggestions but I think we're already doing most of that [22:02:59] gotta run, but i'll check that in the morning [22:03:13] np, I'm done pretty soon too