[10:22:36] fixed drop_old_data_daily and idwiki embeddings extraction dag failures [10:22:40] lunch [13:14:44] o/ [13:41:06] \o [13:41:43] .o/ [13:44:20] are we doing backlog or dpe-mtg today? [13:45:26] I'm not sure myself [13:45:49] Unrelated, but I'm merging the semantic search SSD changes now, ref https://gerrit.wikimedia.org/r/c/operations/puppet/+/1345171 [13:50:52] o/ [13:51:55] ebernhardson: either ways, but I think this'll be only you and me in triage (Trey's out and Peter declined) [13:55:51] hmm, i know my first steps this week are load testing, can always find more things. i guess need to review if anything in triage is important [13:58:12] same happy to pick-up anything that seems urgent from the backlog if running out of work [14:00:50] we seem to have a 3h deploy window reserved next monday (14-17 UTC) for shifting traffic and doing some perf analysis, added this to my calendar, ebernhardson & inflatador do you want to me to invite you? [14:04:31] dcausse yes please. I was gonna ask at the SRE meeting, but is there a place where we can keep track of the SRE follow ups from the DC switchover incident? I'm looking at T433363 but not seeing much [14:04:31] T433363: 🏖️ Southward Datacenter Switchover (Sept. 2026) - https://phabricator.wikimedia.org/T433363 [14:07:02] Peter made https://phabricator.wikimedia.org/T438984 for this experiment, as to where to track anything search related that happened during the switch-over and the codfw incident I don't really know [14:17:20] dcausse: sure [14:17:47] ok invite sent [14:18:46] dcausse: question in #semantic-search about things like anchor tags in semantic search responses [14:18:57] i suspect thats quite difficult based on the enterprise dumps [14:19:01] but unknown [14:19:05] looking [14:19:24] I merged the DNS and puppet changes for opensearch-semantic-search-ssd so y'all should be able to deploy now, I can give it a try if you like [14:20:13] inflatador: thanks! I actually already did the data load, just had to do a stupid thing in python where i replace gethostbyname [14:20:17] but will test in a few [14:20:25] ebernhardson: yes that's very tricky even if we still had the HTML version... I'm not even sure where we would start [14:21:33] indeed, would almost need yet-another-markup-language (not yaml) for a pass-thru... [14:21:41] and even if we kept the html I doubt they'll want the same HTML presented during page views [14:22:40] they suggest this was documented in the tasks "Inline reference markers and links inside the highlighted run stay legible and is not stripped out" but i suspect we never noticed that before [14:22:43] so that'd need a specific rich representation for search responses where we'd pick what we want to keep from the original HTML output and what to remove [14:23:16] no and I think we told them early in the process that the response is going to be plain text [14:23:34] if it's strictly links..i could imagine some sort of side-car data with the passage that says "chars 40-48: link to article Foo`. But not sure that's worthwhile. and it's only anchors [14:24:47] i can imagine the use case though, with longer text they would like the user to progress from there instead of needing a click to article first [14:26:27] school run, back in ~30 [14:26:34] possibly... but seems somewhat tricky and very involved just to keep links and it's a data available in the source we read for now [14:26:46] *not available [14:31:41] it could well be that embeddings model do not care much when reading HTML but the number of tokens is going to skyrocket [14:32:38] at some point I was wondering if an intermediate format like markdown could be a good fit but this needs a highly specialized HTML -> markdown conversion [14:53:25] yea..its hard to say :S [14:53:35] or at least, awkward to provide with lots of funny edge cases [14:55:37] yes... and at this point that's a complete redo of the pre-processing pipeline using the HTML dumps as a source [14:57:44] yea, that's a total non starter then [15:38:15] ebernhardson, dcausse o/ would you have some time for a short triage meeting? [15:38:26] sure [15:42:17] sure [16:08:02] that reminded me...with the goal of comparing ceph with local ssd, i should probably scale down the semantic-search cluster to match the -ssd cluster. [16:16:03] ebernhardson: opensearch-semantic-search@codfw should not be reachable by cirrus IIRC in case that helps [16:16:18] oh right! That's much less risky [16:16:38] annoying because of the added latency to liftwing@eqiad tho [16:16:51] hmm, hopefully i can run locust from the codfw deploy server [16:17:15] oh, right. i did start up a thing to record a bunch of embeddings to a file, then it queries with embeddings instead of liftwing [16:17:17] or you could ask to route the ns entry to codfw directly and use eqiad [16:17:38] but i haven't finished that, just started wiring it into locust. i have to do the data collection side still [16:17:43] yes capturing the embeddings out of the benchmark should be equivalent [16:19:54] also if you have time could measure what it'd cost to run the knn with min_score=0.5, research is looking into a reasonable threshold and that'd be easy to implement if min_score is too costly [16:20:01] *is not too costly [16:20:22] sadly you can't set k & min_score [16:20:23] hmm, is the min_score a short-circuit? Don't think i'm familiar with what that does in context [16:22:05] reading the doc, k feels right. a min_score or max_distance seems too variable given our input queries [16:22:20] but maybe research finds something more concrete [16:24:39] hm... maybe I was reasong min_score the wrong way... it's not supported with our current index shapes anyways... [16:24:52] s/reasong/reading/ [16:25:27] Martin is trying to come-up with a thresold to avoid returning results from gibberish queries [16:25:50] i'm otherwise having annoying problems with python...on stat1009 i have the old directory i used for loadtesting. But it says silly things like `ModuleNotFoundError: No module named 'locust'`. Attempt to install locust, `ModuleNotFoundError: No module named 'pip'`...something has changed :P [16:26:02] (using a custom venv, same as worked before) [16:27:33] so far this is somewhere 0.5 and somehow I was hoping that the knn search could support that out of the box but no... so that filtering will have to happen elsewhere [16:28:01] pip install pip? [16:28:17] tryed hacking pythonpath, only gets to: ModuleNotFoundError: No module named 'pip._vendor.packaging' [16:28:41] probably just needs a new env initialized, i was trying to be lazy [16:31:35] yea, created a new env and it worked...no clue what went wrong with the old env [16:45:12] first requests .. 12s :P Hoping its just filling caches [17:03:48] :/ [17:11:41] so far not getting better, and it's only querying frwiki :S [17:11:45] will have to poke around [17:40:50] hmm, it appears we have 128kB readaheads. And iiuc the set-rdb-readahead.py we did for k8s readaheads locks to /usr/share/opensearch/data, but here we mounted it to /var/lib/opensearch [17:41:57] i suppose it also only drops ra to 64kB, i would have expected even lower [17:49:25] inflatador: when you get back from lunch, can i get you to set readaheads to 16kB on the -ssd hosts? Should be dse-k8s-worker100[1234] [17:49:59] i suspect this is just not nearly enough memory, but we don't have the page-types kernel tool to inspect how much of the memory is being wasted on readahead [17:58:24] all cpu usage is iowait https://grafana-rw.wikimedia.org/d/000000377/host-overview?from=now-3h&to=now&timezone=utc&var-server=dse-k8s-worker1004&var-datasource=000000026&var-cluster=dse_k8s&refresh=5m [18:00:02] not a great sign, frwiki is ~250gb including replicas, we should have ~110Gb for disk cache according to graph, and i'm getting about 1 req/s [18:00:46] although it just got faster with no obvious reason :S always fun :) [18:00:52] not hugely, but 1.5-1.7req/s [18:06:19] ebernhardson 👀 [18:28:44] Ah, looks like I'll need to tweak my playbook to work on local storage stuff too [18:42:19] well, I crapped it out with `for n in $(lvs | grep 170 | awk '{print $1}'); do y=$(ls -l /dev/vg1/${n} | awk -F "->" '{print $2}' | awk -F\/ '{print $2}'); echo 16 > /sys/block/${y}/queue/read_ahead_kb; done` [18:42:59] :( [18:43:11] ebernhardson readahead should be set to 16kb for all `-ssd` pod storage on dse-k8s-worker100[1-4], LMK if ya need anything else [18:43:34] I'm not complaining, just showing how awful my bash is ;) . I'll fix the playbook too [18:48:17] looks like req rate is up a bit, but still pretty mid :S Will let it stabilize a bit on the read IOPS and see if it gets better [18:48:28] like, ~2req/s [18:51:10] inflatador: not sure if it matters, but i still see 128 on `kubectl exec opensearch-semantic-search-ssd-data-0 -- cat /sys/class/block/sdb/queue/read_ahead_kb` which i think is the disk behind the mapping [18:51:44] ebernhardson interesting. I wonder if Puppet is setting it back or if I just missed something. Will check [18:51:45] but i see 16 on /sys/dev/block/254:6/queue/read_ahead_kb [18:51:54] not sure which it uses :( [18:54:03] oh, duh...i have to restart the nodes after readahead changes [18:54:15] because it's all mmap'd...after like 5 rounds of this over the years i should remember [18:54:27] lemme run a rolling restart first [18:54:38] no worries [18:55:48] randomly curious, do you have bpftrace access on these nodes? In theory that can reach in and trace to see the real readahead values that are being run [18:56:00] maybe something like `sudo bpftrace -l 'tracepoint:*readahead*'` [18:56:33] it's the modern version of what we did long ago when first diagnosing readaheads, print out the real readaheads as executed to verify whats happening [18:56:46] (we used perf and kernel probes, but same idea) [18:59:34] wow, even restarting single nodes at a time totally tanks the rate. [19:00:23] actually i think we looked before and bpftrace wasn't available [19:01:41] * ebernhardson realizes since this isn't a live cluster could have closed and reopened the index...oh well [19:06:14] we don't currently have bpftrace installed but I could work up a patch, otherwise I might look at the mwdebug container to see if it's already there (although troubleshooting that way is a lot more awkward than from the host) [19:08:41] something doesn't quite add up in the graphs...i was looking at the cluster overview dashboard for dse_k8s in eqiad. dse-k8s-worker100[1234] show 50-80% disk utilization, but only ~1.5k IOPS and <10MB/s of transfers [19:10:24] I wonder where the disk utilization % comes from. The OpenSearch storage comes from a different LVM volume group than the main OS [19:12:16] not sure, but the timing does match when i started running things [19:12:44] PSI metrics look bad regardless https://grafana.wikimedia.org/goto/sxpr6v?orgId=default [19:15:56] anyway, I'm happy to be your hands and eyes if you wanna do a live sync or something [19:17:45] if you have time, warning i have no clue what i'm doing and just guess at things and collect data :) [19:17:54] usually the data gives ideas though [19:18:04] but right now it's still restarting nodes, up to 11 of 16 [19:31:49] ebernhardson sure, I'm up in https://meet.google.com/fde-tbpf-wqh?authuser=0 [19:54:46] https://github.com/louwrentius/fio-plot/tree/master [19:59:09] https://phabricator.wikimedia.org/P96532 [20:21:04] summary: some data is on 7.2k rpm spinning rust [20:36:43] hmm, i did find some suggestions that the 'rotational' flag might not be valid if it sits behind MegaRAID/PERC (which it does) [20:39:30] ebernhardson interesting. I'll check on the DRAC web UI too [20:49:00] so the /dev/sda is SSD (despite being show as rotational) but /dev/sdb is definitely made of HDDs, `megacli -CfgDsply -a0` shows it from the host [20:49:40] interesting! I double checked across the data nodes, they all report being on sdb1 [20:50:29] err, actually one host says sda1, but they all report the underlying disk at 1999GB, so as long as those hosts still have the disks from the purchase order, they should all be 7200 rpm drives