[01:51:44] go.dog, the ceph cluster was briefly reporting healthy so I ran the ceph reboot cookbooks. [06:25:51] nice, thank you a.ndrewbogott [07:38:46] morning! [08:41:06] morning [10:10:57] rebooting clouddbs was a breeze today, thank you dhinus for the excellent cookbook [10:11:13] cookbook being sre.mysql.multiinstance_reboot [10:11:27] I have filed T435178 to mitigate some of the alert spam [10:11:27] T435178: Make sure wmf-pt-kill service waits for mariadb to be ready - https://phabricator.wikimedia.org/T435178 [10:12:17] pretty sure this is the first time we get such spam since we unified SystemdUnitFailed with production [10:24:37] hmh, why did the clouddb reboots trigger HAProxyWikiReplicaSectionUnavailable alerts? [10:30:31] good question, not sure [10:30:41] * godog lunch [10:45:01] taavi: maybe it did not repool correctly? checking [10:46:08] ah I think x4 currently has only one copy [10:46:17] ah that'd do it [11:02:44] I'm discussing this in #-data-persistence, I think we can fix it in confctl [11:02:57] * dhinus lunch [12:34:26] bd808: we should look at the stashbot elastic->opensearch migration some day this week :P [12:40:59] hmpf... I did not mean to push https://github.com/toolforge/quarry/commit/5b58f83b4a39106bac1d929838a486c448a3ed13 [12:41:55] reverted [13:10:48] what if in striker you could go to a gitlab repository details page and had a single button to set up a deployment token for that tool and add it as a secret to that repository [13:12:16] 🎉 [13:12:46] i think the main problem in doing that is authenticating to the api from striker [13:13:30] but, like, api-gateway now supports service accounts with restricted routes, so I don't think it would be super high risk to come up with a system for static API tokens to map to service accounts, and we could give striker a token that lets it call the deployment token APIs only [13:13:33] yep, it's a pending issue authentication, though I was able to get it more or less running locally at some point [13:13:39] and the rest is writing a bit of glue code [13:13:56] that's also an option [13:15:09] does the ssl termination get all the way to the api-gateway for external urls? [13:16:15] no, for api.svc.toolforge.org the tls is terminated at the haproxy layer with haproxy in https and then the request is forwarded to api-gateway [13:16:29] since the certs our cert-manager can issue are self-signed [13:16:45] there's an anchient task to get publicly trusted certificates for the api-gateway nginxes [13:38:45] dhinus: merged your metricsinfra patches, let me know when you want to try re-enabling the alerts [13:54:58] taavi: thanks! no rush, I can add the custom route to the db [13:58:58] I was thinking of adding a label "{audience: tool-maintainers}" to the new alerts, and use that as the custom route filter [14:27:08] I'm having issues to connect to cloudcumin1001, it just seems to hang, anyone having similar issues? [14:27:20] (I did connect without issues ~1h ago) [14:30:35] for some reason my ssh-agent was stuck ;/ [14:30:49] (first time it happens) [14:39:05] I haven't used it in a while, I tried now and it worked fine [15:26:45] hmm... I got another timeout deploying loki in lima kilo [15:43:46] toolsdb history length growing again :/ I opened a new task T435220 [15:43:46] T435220: [toolsdb] 2026-08-18 Transaction History Length growing too much - https://phabricator.wikimedia.org/T435220 [15:44:20] thanks! [15:51:57] taavi: I would love to get to it this week. I would also love a pet unicorn. Hopefully the first is easier to source than the second. [15:53:04] * dcaro looks at his dog with wondering if he can put a horn-looking party hat on without getting bitten [17:21:07] * dcaro off [17:26:01] cya tomorrow!