[03:26:06] Hello team, we have a quota request pending: https://phabricator.wikimedia.org/T437117 [03:26:39] And this: https://phabricator.wikimedia.org/T437088 [06:47:02] greetings [06:47:05] will take a look [07:53:40] morning! [09:20:00] do we know what's up with these probes flapping? ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes [09:29:28] there was a task for it iirc [09:30:01] ™ [09:32:10] T426584 maybe [09:32:11] T426584: [toolsbeta] probe flapping on ipv6 only - https://phabricator.wikimedia.org/T426584 [09:43:11] ack thank you ! [09:44:44] I've put in a silence [10:24:18] Good morning cloud admins. I have a couple of patches for your review, since they touch your Ceph clusters, as well as ours. [10:25:59] They come now because we're currently upgrading our clusters from reef to squid, but this requires rotating all of the CephX keys. See https://phabricator.wikimedia.org/T428445#12294192 and https://docs.ceph.com/en/latest/security/CVE-2025-30156/#cve-2025-30156-upgrade-steps [10:27:21] I believe that we have bug in the way that puppet management of the cephx keys works, so I made this ticket: T437233 [10:27:21] T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material - https://phabricator.wikimedia.org/T437233 [10:30:08] The script added by this CR should allow us check for any divergence between the on-disk keys and the key that the ceph clusters themselves know about: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1337875 - It's read-only and just gets put on any host with a ceph keyring. [10:32:06] This CR starts using it when puppet is actually writing the keys, so that they get updated properly when the puppet private keydata gets changed. https://gerrit.wikimedia.org/r/c/operations/puppet/+/1337876 [10:56:35] taavi: thank you for creating the News page <3 https://wikitech.wikimedia.org/wiki/News/2026_Commons_links_tables_database_split [10:56:47] I wrote a draft for the cloud-announce email: https://etherpad.wikimedia.org/p/wikireplicas-x4 [10:58:20] dhinus: did you check from data persistence that 'is expected to take several days to complete' is accurate this time as well? [10:59:06] I copied it from the old x3 announcemente... let me ping them about it :) [11:01:48] I did some minor tweaks, otherwise LGTM after the timeline is checked. just send it to the existing cloud-announce thread please [11:02:44] yes that was the plan. I'll wait for a +1 from data-persistence [11:14:20] manuel +1d the "several days", taavi would you wait for amir/zabe or can I send it now? [11:17:15] dhinus: I would assume manuel knows what they're saying here, so if they don't have concerns I would send it :) [11:17:34] thx I'll post it then :) [11:23:01] I forgot to add wikitech-l as well, but I think cloud-announce is probably enough for this one? [11:27:03] yeha [13:54:00] dhinus: I tried to write some docs at https://wikitech.wikimedia.org/wiki/Portal:Data_Services/Admin/Wiki_Replicas#Add_or_split_a_new_section, could you have a look at them if I'm missing something when you have a moment? [13:59:19] Could someone review my MR for T436614 [13:59:28] T436614: Request increased quota for wiki-economics Toolforge tool - https://phabricator.wikimedia.org/T436614 [14:01:22] taavi: thanks! looks good at a first glance, I'll re-read it more carefully later/tomorrow. maybe one thing to note is if data-persistence is involved, or if all the steps are delegated to WMCS. [14:01:39] komla: LGTM [14:01:47] dcaro: thanks! [14:30:08] btullis: Thank you for the very complete write-up on T437233 -- it makes sense and also sorts out my confusion about T399594 [14:30:08] T437233: ceph: Fix the CephX key guard so that Puppet can correctly update the key material - https://phabricator.wikimedia.org/T437233 [14:30:09] T399594: Proposed improvement: Manage CephX users via exported/collected Puppet resources - https://phabricator.wikimedia.org/T399594 [14:30:49] why does a version upgrade also require key rotation? [14:31:15] oh, nm, didn't finish reading apparently [15:44:42] I think that the last dep upgrades broke the debian builds [15:44:45] https://www.irccloud.com/pastebin/8KatlSXn/ [15:45:01] thilp: dhinus Raymond_Ndibe ^ did you see this error when merging those patches? [15:45:36] maybe it was broken before? I did a rollout of all the clis not long ago without issues :/ [15:45:44] no, that’s new to me, but I didn’t merge anything last week I think [15:46:12] ooohhh, it might be the image selection thingie that I changed [15:46:25] let me check [15:46:33] (I tested tox/tests but not package building) [15:47:50] hmmm.... it's like it overwrote the image? [15:48:31] dhinus: I merged in https://gerrit.wikimedia.org/r/c/operations/puppet/+/1335918, if you want to take a look at any erb differences [15:50:14] oh no, even simpler... I copied it to the wrong place xd [15:51:50] 'ProjectProxyMainProxyDown' taavi is that expected? [15:52:23] e.g. shutting down an old one? [15:53:13] no [15:53:24] 2026-09-08T15:53:00.571856+00:00 proxy-5 haproxy[556700]: Connect() failed for backend dynamicproxy-backend-http: no free ports. [15:53:32] quick review? https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/99 [15:53:56] does this mean we're literally running out of port numbers? [15:54:41] ^ first time I see it in the wild (not intentionally created) xd [15:55:10] jhathaway: thx, I'm afk but will check later! [15:56:11] maybe it doesn't recycle them for some reason? [15:56:11] sample-complex-app failed to deploy, harbor having issues? Unknown error (HTTPSConnectionPool(host='\''tools-harbor.wmcloud.org'\'', port=443): Read timed out. [15:56:29] oh, that's the same as the proxy port exhaustion probably [15:56:42] yeah [15:56:54] let me know if I can help [15:56:57] I see lots of open sockets to 172.16.18.225 wsexport-main01.wikisource.eqiad1.wikimedia.cloud [15:57:09] I restarted haproxy for now [15:58:24] I wonder if there's some haproxy timeout that needs tuning a bit more [15:58:26] the machine seems pretty idle [15:58:57] a lot of established sockets [15:59:29] (<800, so depends what you compare it with) [15:59:32] could be related to the shower of alerts and recoveries in -ops ? [15:59:37] there's a serverfault example where people are getting that failure due to confusion between v6 and v4... that doesn't quite sound like us but https://serverfault.com/questions/1041526/haproxy-shows-connect-failed-for-backend-no-free-ports-but-there-are-plenty [15:59:51] I haven't followed it but there have been some network-related issue there [16:00:12] maybe, most of the currently open connections are to an anubis process in the VM [16:00:25] 👀 [16:00:38] (but most seems codfw related AFAICT) [16:00:44] I think that's all codfw? [16:02:45] main thing I'm confused about is that on https://grafana.wmcloud.org/d/iPCm_L94z/web-proxy-usage?from=now-3h&to=now&timezone=utc I don't see anything out of ordinary [16:04:32] hrm, toolforge sets `timeout http-keep-alive` in the haproxy that I missed when porting the timeout configs to the cloud vps haproxy, that should be fixed at least and seems potentially relevant since that controls idle http2 sessions [16:07:03] review for https://gerrit.wikimedia.org/r/c/operations/puppet/+/1337960/? all of those settings are set for toolforge [16:09:39] taavi: the timeout server 1h was also tere in toolforge? [16:09:47] (LGTM) [16:09:48] yes [17:11:04] * dcaro off [18:40:32] 'no free ports' alert again [18:43:42] taavi: ^ [18:49:42] not sure if related, but am hearing from the devex team that the beta cluster is getting scraped pretty badly today [18:51:58] I don't think beta cluster uses the same proxy very much, but it could be the same scraping episode [19:12:44] looking at https://grafana.wmcloud.org/d/iPCm_L94z/web-proxy-usage?orgId=1&from=now-30d&to=now&timezone=browser, is there any reason why all active connections are going to proxy-5 and not proxy-6? (see graphs at bottom) [19:13:06] I do not think they're meant to be active/active [19:13:12] ah. [19:16:52] It's very unlikely any of these connect attempts are actually doing anything, I just need haproxy to close them out faster... [19:18:30] Yep, Beta Cluster has been very unstable since yesterday [19:20:19] The Beta Cluster has its own (cache) proxy, which should be able to sustain the load, but even the cache proxy - which should be able to withstand the traffic, even if the appservers are getting flooded - seems to be unstable... [19:33:14] taavi and/or bliviero, my not-great mitigation is now documented on https://phabricator.wikimedia.org/T437349 -- we should be good for now but I also hope we can make the proxy do something smarter in this scenario. [19:38:11] Talking about this scraping, could I be added to https://phabricator.wikimedia.org/T393487? [19:39:25] done [19:40:08] Thanks [19:40:30] Let's continue fighting fires... [19:40:36] thank you andrewbogott [19:41:15] Speaking of fighting fires, now I'm going to open up my chimney and release a soot-covered squirrel into my living room. Back soon! [19:42:45] :-)