[00:36:18] proceeding shortly. with luck we should have the updater back after this [00:44:20] updater's flowing again :D [00:52:47] 12 shards away from green status [01:02:11] dangit, realized i missed a few hosts, the 12 shards are perma-stuck. chasing down the remaining hosts needing restart [01:17:47] back to green [10:27:22] lunch [13:36:08] \o [13:47:39] o/ [15:07:51] inflatador: do i need to do anything special to deploy opensearch-semantic-search-ssd, or should a simple helm apply do it? [15:08:29] i guess i can just try, but didn't want to set off random alerts for a failing service if it needs more [15:12:20] to update on my complaint earlier this week about Valid-Until in docker images preventing builds, it looks like we are getting rid of that requirement: T438866 [15:12:21] T438866: Ignore Valid-Until for WMF container images - https://phabricator.wikimedia.org/T438866 [16:15:20] meh...looks like i have to default security plugin to on for the opensearch operator bootstrap to work. We define env vars per-node pool, and there is no existing option to provide the bootstrap env vars. [16:16:38] oh right, I remember Brian stumbling on this too... [16:18:18] i'm considering...would it be silly to let the entrypoint grep opensearch.yml for security config, and enable by default when it's there? It can just be default on, but i find that annoying [16:20:12] are there some env vars specific to k8s that the entrypoint could inspect and possibly set this? [16:20:28] there is KUBERNETES_SERVICE_HOST which should be set everywhere [16:20:56] re: opensearch.yml can't really remeber what's in there in the image [16:21:24] perhaps if you run k8s that's a strong indication that you must run the sec plugin [16:21:37] yea that's probably a bit more concrete and less magical [16:21:59] plain docker should not set KUBERNETES_SERVICE_HOST so hopefully that's a noop for existing mwcli/mw-docker setup [16:54:21] (brian’s out today fwiw) [17:04:15] so much fun :P bootstrap has gotten further but i don't know if the current logs (varied ports with Authentication finally failed for secret from [::1]:32888) is an error or just random parts of init... [17:05:35] sadly i expect it means some request is being made with wrong secrets... [17:16:08] sigh, yea declaring it stuck :S trying to find the operator logs to understand what it thinks [17:47:25] dcausse: can we merge https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/2693? [17:50:46] pfischer: yes [17:51:58] I've kicked a manual run in the backgound, if all is good the dags in airflow will pick-up the right snapshot to analyse only the weekly delta [17:53:14] I find knn very slow compared to what it used to be after Erik's optimizations on readahead, I wonder if that setting could have been lost after the restarts for the 3.7 & 3.8 upgrades [17:58:17] ryankemper: i've realized whats wrong, the secrets in /etc for this namespace are empty. Didn't see anything in operations/puppet, expecting that's in the puppet private repo. hoping it's easy to dupe the opensearch-semantic-search config? [17:59:23] ebernhardson: I can take a look in 5'. what's the task exactly, copying secrets that are already in the private puppet repo for `opensearch-semantic-search` to `opensearch-semantic-search-ssd`? [18:00:05] ryankemper: yes, i mean perhaps newly generated passwords are appropriate, but generally copying whats already in puppet private for the new namespace [18:00:33] dcausse: thanks! [18:01:24] see also https://wikitech.wikimedia.org/wiki/Data_Platform/Systems/OpenSearch-on-K8s/Administration#Deploying_a_New_OpenSearch_on_K8s_Cluster (which i should have re-reviewed before spending time guessing whats wrong..) [18:03:16] perhaps not... RA at 64 for rbd1 (/var/lib/opensearch) on a random opensearch-semsearch node in codfw... so not that... [18:03:28] hmm [18:03:56] possibly I feel it's slow because it was never properly warmed up? [18:04:50] anyways curious to see what the bench comparison will tell us [18:04:54] dinner [18:11:54] alright looking at the secrets now. and right, ill generate new ones in same format [18:12:15] thanks! [18:13:40] ok hold that thought, it's time to repool wdqs/search apparently xD will circle back to this [18:16:18] certainly [18:29:07] a reminder that we still have all more_like queries pointed at eqiad, i can undeploy that in a few hours in the deploy window [18:31:41] ebernhardson: that's not a blocker to repooling cirrus is it? in fact I'd think that we actually want that because the cache is prob cold now? [18:32:05] (meanwhile wdqs-main is back up and scholarly following rn, then wcqs. i'll pause atp before cirrus ofc) [18:32:13] ryankemper: not a blocker, just an oddity of how things are now. I expect that once you get both clusters serving traffic we can undo it and let codfw have higher miss rate [18:32:27] got it yeah that makes sense [18:33:10] the deploy window is fairly full, it might not make it and have to be deployed tomorrow [18:33:23] (also not a big deal, but eqiad will be more loaded) [18:47:22] ebernhardson: okay pws should be live now. I used `hunter2` for both [18:50:01] pfischer: [18:50:03] https://grafana.wikimedia.org/d/opensearch-k8s-comparison/opensearch-on-kubernetes3a-a-b-comparison?from=now-3h&to=now&timezone=utc&var-datasource=000000026&var-k8s=$__all&var-cluster_a=opensearch-semantic-search&var-site_a=eqiad&var-device_a=rbd.%2A%7Cdm-.%2A%7Cnvme.%2A&var-cluster_b=opensearch-semantic-search&var-site_b=codfw&var-device_b=rbd.%2A%7Cdm-.%2A%7Cnvme.%2A&var-nodepool=$__all [18:50:05] &var-worker_cluster=.%2Ak8s.%2A&var-search_cluster=semanticsearch_test&refresh=1m [18:51:55] * ryankemper laughed a little bit at the `Reading it honestly` subsection. never change, claude [18:52:06] "The genuinely loadbearing way to read it" [18:54:36] Alright, I'll bring cirrussearch back up in 5' [18:55:47] thought the same [19:12:39] repooling codfw, chi first [19:18:47] nice! can see latencies falling back into line [19:48:12] stepping out to meet with landlord, cirrus looks good [20:46:45] hmm, was trying to start the -ssd cluster with 16g*16nodes, but stuck at pending after allocating 4 :S [20:47:19] couldn't find a good grafana dashboard on per-node memory limits/allocation, but there is one that shows topolvm-balanced has plenty of capacity and shouldn't be the limiter [20:59:46] i'm going to randomly guess someone in sre can move pods to a different host and free up memory, but don't think i have access to see the exact details [21:00:06] for now i've undeployed it, will check in with -sre tom morning and try again [21:04:27] more_like is also now back to standard routing through dnsdisc, initial latencies look within normal bounds.