[08:04:59] unblocked cindy, was failing because of new ssh gerrit finger prints [10:41:42] lunch [12:30:59] updated T420230 with some bullet points [12:31:00] T420230: Migrate to OpenSearch 3.x - https://phabricator.wikimedia.org/T420230 [12:33:50] \o [12:33:59] o/ [13:59:09] o/ [13:59:50] .o/ [14:00:08] I would skip this week’s retro so we can attend the staff meeting. [14:14:26] same [14:23:57] sounds good [14:28:23] kk [14:42:30] meh... can't convince blubber to build an image on ml-lab, failing with apt unable to access deb.debian.org... tried https://doc.wikimedia.org/releng/blubber/configuration.html#proxies with no luck... [15:01:58] dcausse: can you build it locally and scp it? [15:05:28] that's what I'm looking into and then realized that the image is close to 50Gb... [15:05:38] !!! [15:06:57] yes... not sure why... using the base image docker-registry.wikimedia.org/ml/amd-vllm022:gfx90agfx942rocm7.2.0pytorch2.10.0flash-attn2.8.3aiter0.1.13vllm0.22.1-3 as that's what other service are using [15:18:53] i don't really know what to do about that... [15:20:33] I feel like it's because we copy multiple times the python venv and that adds up in different layer, the python venv is close to 9gb apparently [16:21:48] oh wow, that's craziness [16:26:15] trying with another baseimage with pytorch+rocm, it's about 14gb... a bit more manageable if I have to do one or two roundtrips... [16:26:28] after that I'll bother ml folks to understand how they test :) [16:32:12] At least we're not running Windows. Dunno if they've gotten any better, but our cloud Windows images started at 30G [16:37:31] ouch... [16:47:55] dinner [20:40:19] shipped the cirrus and sup config patches, page_rerender_upsert jumped to ~70/s from basically 0/s, so that's most likely our redirects [20:44:30] saneitizer avg completion is like 98%, so it might go up a bit higher when it gets back to 0 (by 98% all the small wikis have fallen out and completed, mostly)