[10:12:21] lunch [13:15:40] The three new nodes are ready in dse-k8s-eqiad (dse-k8s-worker104[2-4]) - but your existing pods aren't going to be moved there. If you `helmfile destroy` and then `helmfile apply` it, you should be deployed onto the new nodes, with SSDs. [13:24:57] o/ [13:40:01] \o [13:40:06] btullis: thanks! I'll work that out today [13:41:04] o/ [13:43:27] btullis, brouberol: by any chance is at least one of you two be able to join the WMCS ingress logs meeting in ~37 minutes? [/me just checking that we have quorum from all teams involved, sorry for the cross-posting, I'm never sure which channel the people follow more] [13:52:51] volans: Yes, I'm here. Balthazar is off today due to torrential rain in his part of the world. I'm swotting up the pre-read now :-) [13:56:14] great, thanks a lot, sorry to hear about the rain for Balthazar though :( [14:01:26] inflatador: I don't have much for pairing but please let me know if you have something [14:02:07] dcausse not much. Mainly I wanted to ask about T438977 and if there was anything I could do to help [14:02:07] T438977: Query degradation routing - https://phabricator.wikimedia.org/T438977 [14:02:36] inflatador: sure we can have a quick chat about that if you want [14:02:50] dcausse cool, I'm in the Meet if you wanna join [14:50:46] https://opensearch.org/blog/opensearch-kubernetes-operator-3-0-is-now-generally-available/ [14:52:49] nice! [15:00:16] Yeah, I feel bad 'cause I said I'd help test the alpha version but never ended up doing much ;( [15:02:54] dcausse: ping for wed meeting [15:03:02] oops, joining [17:09:32] dinner [17:23:45] hmm, tried deploying opensearch-ssd and right now it's leaving them pending: Warning FailedScheduling 63s default-scheduler 0/40 nodes are available: 2 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 30 node(s) had volume node affinity conflict, 4 node(s) had untolerated taint {dedicated: wdqs}, 4 node(s) were unschedulable. preemption: 0/40 nodes are [17:23:47] available: 40 Preemption is not helpful for scheduling. [17:25:15] We just added the new nodes with the SSDs, sounds like maybe they aren't fully joined to the cluster? Will take a look [17:26:47] thanks! [17:32:51] The node labels look OK at 1st glance w `kubectl describe node dse-k8s-worker1042.eqiad.wmnet`. I see that the PVCs pre-date the new hosts though. I'd say undeploy and then delete all the disks with `kubectl delete pvc data-opensearch-semantic-search-ssd-{masters,data}` or similar after undeploying, then try redeploying [17:33:08] oh! i'm a dumb dumb..yea i didn't delete the PVCs [17:34:37] Maybe once we get on operator 3.x it will do it for us. That'd be nice anyway ;) [17:34:45] ok deleted pvcs, trying again [17:35:33] yup, got ContainerCreating now! thanks [17:35:41] well see if it manages all of them [18:32:05] huh, from T438896, this would also probably effect our morelike hit rate: "eqiad has double the memcached capacity codfw has atm, as we are waiting for the refresh. While not the sole reason for the swithover issues, it is a contributing factor" [18:32:05] T438896: Investigate high saturation in codfw - https://phabricator.wikimedia.org/T438896 [18:33:24] !!! [18:47:53] wonder what the warning should say when the query building profile selects CONTEXT_NONE and we return an empty result set.