[10:07:23] lunch [12:53:56] errand [13:22:33] \o [13:42:07] so much better, first run on the -ssd hosts hit 180 req/s [13:42:32] err, actually table is off :P it's 18. but still better :) [14:01:33] o/ [14:01:43] nice! [14:05:17] what's the RPS on the ceph-backed cluster? [14:07:23] inflatador: higher, but i haven't tuned anything or ramped up yet [14:08:56] looks like on the old benchmark we peaked with 35 req/s at 270ms mean latency [14:09:23] ramping up the concurrents, this is now at 51 req/s [14:09:51] i have to look into how k8s/containers perform with disk cache though. I wonder if unused memory on the host is allowed to expand the expected disk cache (basically it doesn't get evicted) [14:11:57] I would expect the file cache behavior to operate at the same level with or without a container but I honestly have no idea [14:12:54] i would hope so too, but i can kinda imagine if the host just has free memory sitting there maybe it doesn't hit the eviction routines? But maybe they do fire per-container based on cgroups [14:24:28] 46k requests, mean 190ms, 53 req/s not too bad. Did have 35 failures communicating with liftwing, i guess i'll also need to figure out if we are overloading that (and run the variant with pre-calculated query vectors) [14:24:41] that was upped to 10 concurrent req's [14:24:45] could suggest that you could give a container mem limits more or less equals to the jvm max heap? [14:25:19] dcausse: hmm, yea i suppose i could try that should give some better hints. Claude claims that it should evict based on the configured limit, but that should verify [14:26:08] yeah, to me that's the thing I'm not sure about. Does normal kernel file caching behavior help the container in any way, even indirectly? Or is the file cache part of the container's cgroup memory? It sounds like Claude is saying the latter [14:26:25] from what i can tell in grafana i don't think we were overloading it. I haven't 100% verified it's the right gpu, but the graph loks right for one gpu that ramped up to 20-25% utilization and fell off when i stopped [14:27:19] inflatador: claude is saying the later, i just don't know if i trust it enough :) will verify [14:29:45] yeah, it makes sense to me too...looks like the node exporter has a metric for it https://grafana-rw.wikimedia.org/goto/sk84mn?orgId=default [14:31:38] in separate news, thankful for writing that bit to configure the cluster for ml from cirrus-toolbox. So much easier this time around [14:31:46] just apply and it works [14:32:17] And cadvisor has a `container_memory_file_dirty_bytes` so that implies the file cache is indeed within the cgroup (ref https://github.com/google/cadvisor/blob/master/docs/storage/prometheus.md) [14:33:05] dcausse: this patch probably needs a review, it's a security ticket (patch in ticket): https://phabricator.wikimedia.org/T437940 [14:33:34] not critical, i don't think that path runs at wmf, but sec asked us to resolve it [14:56:45] looking [16:38:26] I'm not sure that we have the readahead configured correctly for opensearch on SSD, at the moment. [16:38:50] btullis: it might be ok, i'm seeing 150k iops and 600MB/s, i think thats 4k/iop roughly? [16:39:43] it might be that opensearch 3 is finally properly instructing the kernel about how it intends to read randomly. It's been a thorn in the side of lucene for awhile but they finally added some support for it [16:40:33] perf is way up vs ceph though, i have half the memory and seeing higher qps [16:40:40] Great! [16:52:03] The Ceph CSI provider also assigns a much higher readahead value than the default for non-RBD devices, I guess that is an optimization for other workloads [17:02:46] also getting some cpu throttling, restarting now with 6 cores instead of 4 to see if it does better. Not clear we actually need more cores for the deploy, just trying to understand the limis [17:02:48] *limits [17:46:04] curiously, adding 50% more cpu cores increased cpu usage, decreased disk needs (~500MB/s -> ~300MB/s), and increased throughput (~60->90qps) [17:46:19] i don't know what to make of the disk iops/throughput going down :P [17:47:24] also store size somehow dropped ~10% going through the restarts, something i don't understand happening there. Would have expected no change [17:51:16] could be wrong but the indices have max_thread_count to 1 for merge.scheduler, wondering if some merges were still happening before your first round and the last restart? [17:52:53] shouldn't be, we have a graph for bytes currently merging and it's been flat, show merging from indexing completed ~16 hours ago [17:54:18] i can't explain the drop in store size, but i wonder if perhaps starting the pods didn't evict the page cache, so the pages cached by the previous pods are still sitting in memory but not attributed to the container anymore [17:54:22] not sure how to get around that [17:54:27] (also not sure it's the cause, just a theory) [18:02:03] inflatador: random idea, could you 'echo 1 > /proc/sys/vm/drop_caches' on dse-k8s-worker104[1234]? That should drop the page cache and rebuild it [18:02:54] although it may or may not help with mmap, unknown [18:03:19] i think i used that back when evaluating readaheads on the bare metal instances though, used to work [18:09:38] ebernhardson sure, will give it a shot [18:10:16] inflatador: actually just 104[234], 1041 is listed but thats a completed pod (and not a pvc host) [18:11:39] ebernhardson got it, I just applied. Stepping away for lunch, but should be back in ~45 [18:11:49] thanks! [18:15:08] no change, throughput staying pegged at 300MB/s (vs 500MB/s before)...no idea :P [18:32:18] https://www.linkedin.com/blog/engineering/data-streaming-processing/overcoming-challenges-with-linux-cgroups-memory-accounting (esp the behavior after a restart) my rough understanding is that the page cache survives a restart but the new container does not own it and thus actual usage is not visible from the container perspective [18:33:07] interesting! reading [18:37:06] yea that does seem to be highly likely [18:51:12] hmm, restarting the cluster and one of the pods is failing to come back, not even in the `kubectl get pod` list, instead the operator is regularly logging a reconcile error against the cluster name: failed to delete resource: getting resource failed: failed to get restmapping: failed to find API group "monitoring.coreos.com" [19:02:03] oh nevermind, that error is a red herring. Instead the StatefulSet failed to recreate due to hitting cpu quota (i'll turn it back down, we don't actually need that many cpus was just seeing if the disks support more without cpu throttling) [19:25:43] Random thought: we had a bot trying to use WDQS to run AI benchmarks (ref https://docs.google.com/document/d/1u9ZQshc2U9g1VbOgCUlx3vT2VSKOozjlcu4YWa7sN2M/edit?tab=t.0 ) , I wonder if the increased fulltext search traffic could be explained in the same way? [19:34:35] hmm, certainly seems plausible. [19:36:34] fwiw, after another round of restarts throughput and iops are back to previous levels. So seems there will be some uncertainty with benchmarks, but probably still mostly valid [20:09:30] Interesting article! I guess their use of mlockall() followed by an OOMkill can be problematic [20:09:46] Pretty sure opensearch wants mlockall to function as well