[06:42:36] greetings [07:30:49] morning! [08:16:42] godog: there's a bunch of libvirt tls alerts, is that you? [08:17:14] dcaro: ah yes indeed, thank you [08:17:20] 👍 [08:18:01] did you mention it somewhere? (I might be missing looking somewhere for the actions xd) [08:18:14] no I forgot to !log and I did it just now [08:18:27] on -operations [08:18:32] okok, np [08:19:14] morning [08:19:18] oh, interesting, we usually put in in the admin project (mention in -cloud), any reason that changed? (maybe I should be looking in -operations now) [08:20:38] I don't mind either way tbh, can do admin project [08:21:08] I'm ok too, just trying to figure out where to look for them :) [08:24:14] ok I did admin too just in case [08:25:41] thanks :), might be good to clarify which one to write to though, I can add a note to the team meeting just so everyone is in sync [08:27:12] sure sgtm, thank you [08:28:52] hmm... I think most cookbooks that do stuff on cloudvps log in admin by default, though the cookbooks that are `sre` log on prod, so there might be a spread of things [08:29:38] yeah I usually make sure I include the task so it gets logged there [08:30:22] the admin !log on -cloud is what i use for everything that's limited to 'our' hardware [08:37:39] yep, I do the same too [08:39:00] SGTM, I'll do the same [09:48:31] * aputhin is back [09:48:37] hello ya'll, good morning [09:49:13] good afternoon! [09:51:06] welcome back aputhin [09:51:10] erm, I see our weekly has moved 1h earlier permanently. unfortunately that now conflicts with our OKR bi-weekly. I'll need to check if we can move that. [10:04:19] aputhin: sorry, I think we missed that overlap when checking if it could be moved. the reason for moving was avoiding the overlap with the monthly staff meeting, without having to adjust the time once a month [10:15:01] * dcaro lunch [10:40:46] i am merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328644 to move novaproxy backend selection and proxying to haproxy, as with last time puppet is disabled on -6 to allow a quick rollback [10:41:00] ack [10:53:03] seems like targets given as v6 addresses are broken. that only affects tofu.wmcloud.org, I will see about fixing [11:09:59] fix is https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329253 [11:31:34] LGTM [12:33:12] bah, the proxy is hitting fd limits [12:34:16] of course, I copied all the in-haproxy tuning settings from toolforge but not the systemd unit override which is configured separately [12:35:58] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329271 [12:56:25] * dcaro paged [12:56:30] what's going on? [12:56:50] alertmanager is giving me 503s [12:58:36] taavi: project proxy is having issues? [12:58:49] are you working on it? [12:59:38] oh, codfw [12:59:43] ? [13:00:07] still? [13:00:15] https://usercontent.irccloud-cdn.com/file/y92qEILK/image.png [13:00:17] not only codfw it seems [13:00:39] the page is for codfw though [13:00:46] MetricsinfraAlertmanagerDown [13:00:55] dcaro: will be 10-15' before I can look, stopping keepalived abd puppet will fail over traffic in the meantime [13:01:10] on proxy-5, that is [13:01:14] ack [13:02:07] (👋 JFYI I'm getting this when I try to see a project's proxies on Horizon) [13:02:11] https://usercontent.irccloud-cdn.com/file/aDBPFvZN/image.png [13:02:47] (assuming it's related to the ongoing issue) [13:03:09] I think it is yep [13:05:22] jnuche: sholud be working now, still having issues? [13:05:54] dcaro: back to normal, thanks! [13:08:38] okok, things are back to normal (as far as I can tell), will wait for taavi to get back to properly fix it :) [13:15:00] FYI I'll be doing a dumps-nfs failover in 5 [13:17:41] it is still seemingly complaining about sockets: [13:17:45] Aug 25 13:02:29 proxy-5 haproxy[631997]: [ALERT] (631997) : socket(): not enough free sockets. Raise -n argument. Giving up. [13:17:53] looking why my previous patch was not enough [13:28:59] turns out there is one more config file with a limit than i remembered https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329280 [13:30:05] LGTM [13:36:27] godog: I assume the paws alert is related to your work? [13:36:52] taavi: it is yeah, checking though it should self-recover [13:37:05] cool, just checking that it isn't haproxy related [13:37:18] indeed thank you for the heads up [13:40:21] hub is back on the browser, standing by for the alert to resolve [13:41:10] there it goes [13:41:23] k8s took a little to notice what was going on and then worked as intended [14:28:21] bliviero: just double-checking: the new 'clouddumps' servers (currently clouddumps100[12]) should be named dumps100[3-5]? [14:28:46] or dumps-distrib*? [14:29:46] * andrewbogott has no opinion [14:37:44] Southparkfan: you're working on deployment-prep upgrades right? Is there anything I can do to help/support you? Have enough quota headroom &c? [14:58:55] andrewbogott: can i get back to you on the clouddump* naming until Ben is back? when do you need that info [15:16:16] andrewbogott: that's correct. I haven't encountered quota problems, but have to admit I haven't started migrating yet. [15:16:52] We upped the quota last time, hopefully it's sufficient for this migration :-) [15:16:56] ok! thank you in advance, and please ping if there's anything I can do to help [15:18:46] Thanks for the offer! I'll need your help for merging Puppet and mediawiki-config patches [15:19:31] If you want to review some or all tofu-provisioning MRs as well, happy to ask you. [15:41:11] you should definitely ask although I am typically baffled by tofu syntax [15:51:07] since the proxy issue we have no alerts at all? https://alerts.wikimedia.org/?q=team%3Dwmcs [15:51:17] * dcaro happy but also suspicious [15:52:46] nope, it seems [15:52:54] and https://alerts.wikimedia.org/?q=%40cluster%3Dwmcloud.org still works so it's not an AM communication issue [15:54:26] I think no alerts is right. The last time I looked it was all cloudvirts in flight and filippo finished them up today. [16:31:20] andrewbogott: bd808: can either of you think of a reason not to do T431284? [16:31:20] T431284: Fold special maps proxies back to the normal proxy - https://phabricator.wikimedia.org/T431284 [16:32:41] Now that the maps project is not running a tile server it should be fine I think. The public tile server was the reason for the split as far as I remember [16:32:58] I commented on the ticket. It sounds great but I'd like to see the new proxy burn in a bit. [16:35:18] taavi: have time for another glance at https://gerrit.wikimedia.org/r/c/operations/puppet/+/1318779 before you go for the day? (possibly you have already gone for the day) [16:35:20] bd808: I don't believe that is stopping some old clients from trying to load tiles :/ [16:36:52] I fell into the rabbit hole of reading about the historic toolserver.org & OpenStreetMaps collabs. [16:37:28] https://phabricator.wikimedia.org/T190451 has a bunch of history if anyone cares. [16:38:17] * taavi cares [16:39:33] andrewbogott: +1. I might send some stylistic fixes later (notably, the singular tab with the rest of the spaces in the script, plus lack of stdlib::ensure()), but it's good enough so that I can disappear and make dinner instead [16:39:53] ok! I will for sure knock out that tab [16:40:25] apparently my vim install is haunted [16:40:31] I think the really short TL;DR is that WMDE has OpenStreetMap ties that extended into toolserver.org and that led to WMCS having a poorly documented OpenStreetMap relationship as the successor to toolserver.org [16:41:18] where do on-wiki map tiles come from now? I kind of thought we ran a tileserver in prod... [16:41:52] maps.wikimedia.org [16:42:01] we do run a prod tileserver. it serves completely different tiles than the former Maps project tileserver [16:42:23] different renderer, and different origin data [16:42:31] OK, but I guess that's not what you mean by 'relationship' then? [16:44:30] I guess if we were providing tiles to OSM services that weren't otherwise WMF services, that's more of a tangle [16:47:04] One of the things that the Cloud VPS maps project used to host was tiles for the https://wiki.openstreetmap.org/wiki/Hike_%26_Bike_Map project. [16:48:01] https://hikebikemap.org seems to be still up, but very broken. (like mixed TLS/non-TLS levels of broken) [20:31:03] !log deployment-prep T401839 - provisioned deployment-docker-wikifeeds01 (trixie) using Tofu (cloudvps-repos/deployment-prep/tofu-provisioning) [20:31:04] Southparkfan: Not expecting to hear !log here [20:31:04] T401839: Migrate deployment-prep away from Debian Bullseye to Bookworm/Trixie - https://phabricator.wikimedia.org/T401839 [20:31:14] shrug... [20:38:19] Could someone force a puppet run on deployment-docker-wikifeeds01? Its initial puppet run failed, and the instance does not let me ssh in. I have fixed the Hiera, puppet should run now. [20:55:39] Southparkfan: I'll try [20:57:55] Southparkfan: can you access now? [20:58:26] * andrewbogott does the puppetcert dance while he's in there [20:58:46] Yes, that has worked. Thanks! [20:59:40] What kind of puppetcert dance do we need to perform? [21:00:18] hmmm did you reuse an existing hostname for this? [21:00:23] I'm seeing unfamiliar puppetserver warnings. [21:00:58] Not to the best of my knowledge [21:01:18] ok it's happy now [21:01:27] so normally when a VM first comes up it uses the central puppetserver... [21:01:32] Ah, looks like Tofu happily accepts the complex profile::docker::runner::service_defs object. [21:01:46] but most things in deployment-prep have hiera set up to point them to deployment-puppetserver-1.deployment-prep.eqiad1.wikimedia.cloud [21:02:15] (so the puppet failure won't occur anymore) [21:02:32] Is autosigning enabled on the puppetserver? [21:02:43] so once that hiera setting applies, then puppet breaks because the host wants to talk to a puppetserver that has never heard of it... [21:02:46] https://wikitech.wikimedia.org/wiki/Help:Project_puppetserver#Step_2:_Setup_a_puppet_client [21:03:15] so then you rm -rf /var/lib/puppet/ssl on the client and then puppetserver ca sign --certname on the project-local puppetserver [21:03:46] hm, it looks like autosign is enabled [21:03:57] so in that case it's probably just removing the busted certs on the client [21:04:00] and re-running [21:04:28] I can vaguely recall doing that for the auditlogging VMs, yeah [23:31:32] andrewbogott, Southparkfan: I figured out just a bit ago that the necessary magic is for the second puppet run (after the first changes the config to point at the project local puppetserver) to be run via `sudo -i run-puppet-agent`. That helper script knows how to fix up the certs when needed. [23:32:04] Ahmon updated that stuff for T429413 [23:32:05] T429413: Eliminate sudo rm -rf /var/lib/puppet/ssl step in new deployment-prep WMCS project (and others) - https://phabricator.wikimedia.org/T429413 [23:33:57] I think that also means that waiting for the puppet timer to trigger the second run would "just work".