[07:52:16] hey folks! [07:52:38] I am prepping the last docker image copies before switching the docker registry from swift to s3 [07:52:45] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1333237 [08:48:11] elukey: naive an non-urgent question: should we expect faster/slower/unchanged pull times? [08:49:44] brouberol: you sould expect less swift but more SSStrong pull times :-P [08:49:47] * volans hides [08:50:06] brouberol: far from naive - the registry metrics show that S3/apus seems much faster, but we'll need to check it with a sustained pull load [08:50:13] so I expect/hope quicker pulls [08:50:36] volans: _slow claps_ [08:50:55] it'll be interesting to get numbers for the chonky ML images [08:52:29] correcting myself - it should be faster for high percentiles [08:52:59] brouberol: that is already there, ML now has a separate bucket (already migrated weeks ago) https://grafana-rw.wikimedia.org/d/StcefURWz/docker-registry?forceLogin=true&from=now-6h&timezone=utc&to=now&var-datasource=000000005&var-instance=$__all&var-docker_distribution_instance=.%2A:5005 [08:53:23] under "storage" you should see S3 timings [09:09:14] if someone has time for a quick +2 in puppet, https://gerrit.wikimedia.org/r/c/operations/puppet/+/1333291 would appreciate it [09:19:00] urbanecm: I can take care of it [09:19:11] ty ❤️ [09:20:02] done :) [09:20:23] thanks! [09:20:24] elukey: CR thief [09:25:52] quick update on the docker registry - I am still finishing up with the latest copies, so I'll kick it off right after lunch! [09:26:13] ack [10:39:01] moritzm: cookbooks.sre.dns.netbox with changes for ldap-replica1006, I'm approving given it's a new one [10:47:31] ack, thx. that's expected, I'm currently creating new LDAP replicas for the trixie migration [11:05:39] FYI; I'm disabling puppet merges for ~ 10 mins for a Puppetserver reboot [11:11:44] and Puppet merges are re-enabled [12:04:28] moving the registry to S3! [12:09:35] ok don [12:14:53] elukey: is that now all of docker registry stored in apus? [12:15:08] Emperor: correct :) [12:16:37] cool. [12:17:48] Emperor: we have also latency metrics about every S3 action from the registry to apus, from a quick glance I think we'll have a faster service at high percentiles compared to Swift, but I'll confirm after a day of datapoints :) [12:20:00] now that 've said that, I noticed high latencies [12:20:02] sigh [12:20:32] I would be a little surprised if the tiny apus cluster (without any of the NVME-pools as yet) could out-perform the giant ms-* swift clusters [12:22:26] in theory the S3 protocol should be more efficient, and at steady state (redis cache warmed up etc..) I don't expect problems.. I think there are some "move" "blob_upload" ops that may be more expensive, but I'll dig a bit more into it [12:23:07] there's a bunch of expansion of apus coming this FY which should improve the overall performance of the cluster [12:23:12] regular ops etc.. appear to be as fast as swift [12:28:54] Emperor: having NVMEs or similar would surely help in the future when we'll implement garbage collection workflows, since docker distribution uses a mark and sweep algorithm [12:29:18] yeah, it'll make any operations that want to list bucket contents faster. h/w is on order now at least... [12:30:15] super [15:45:05] hnowlan and claime: It's working on staging \o/ [15:45:27] 200, no issues [15:46:05] going to codfw [15:46:45] Nice [15:46:55] Always works better with a Service [15:48:11] I'm shocked to hear that [15:49:58] https://grafana.wikimedia.org/d/b1jttnFMz/envoy-telemetry-k8s?orgId=1&from=now-1h&to=now&timezone=utc&var-datasource=000000026&var-site=codfw&var-prometheus=k8s&var-kubernetes_namespace=thumbor&var-app=$__all&var-destination=$__all [15:53:30] \o/ [15:55:00] I am going afk but I'll check the registry's stats later, if there is anything weird ongoing and you are not sure call me :D [15:55:27] my team already knows etc.. but if there nobody around, I'll do a quick check [15:55:29] o/ [15:57:21] I'm going to eqiad [16:40:40] "MW Deploy" annotations don't seem to work anymore in Grafana. Data source not found, it says. [16:41:56] Looks like they came from graphite [16:42:08] where OPTIONS https://graphite.wikimedia.org/events/get_data?from=-6h&until=now returns HTTP 502 [16:43:19] I'm guessing Scap was sending them to both Loki and Graphite and so there was never a push to fully update the settings from https://wikitech.wikimedia.org/wiki/Grafana#Show_deployments everywhere [16:44:48] Hm.. well, graphite is still defined as datasource so that's not the reason for the first error [16:44:54] "Public Logs (Loki)" isn't listed anymore as datasource [16:44:58] I guess that's the first error [16:45:00] theyre both gone? [16:47:54] That might be related to the recent change in plugin management in grafana (cc denisse) [16:48:22] A minor version bump of grafana changed how plugins worked and subsequently removed a bunch of plugins [16:49:21] Graphite-web has been decommed as of august 31st, and was RO for a year beforehand so it's less likely that it's at fault [16:56:22] ref T406478 [16:56:22] T406478: Scap logs on Grafana dashboards are broken - https://phabricator.wikimedia.org/T406478 [16:56:38] OK. I guess maybe Loki got disabled then? [17:10:21] T436045 is the ticket for reenabling loki [17:10:22] T436045: Grafana no longer includes many datasources but adds them as plugins - https://phabricator.wikimedia.org/T436045