[00:28:18] @lucaswerkmeister: I would say very safe that ToolsDB will stay MariaDB until and unless the WMF moves off of MariaDB and back to MySQL. [00:34:26] which, at this point, seems unlikely unless big things happen [00:35:10] Let's switch to Microsoft SQL Server [00:35:18] wait no, MongoDB—it's webscale after all [07:32:06] csv files on NFS seems like the logical next step here [07:32:08] xd [08:12:51] bd808: sounds good, thanks ^^ [08:22:02] !log admin put cloudvirts back in service T431682 [08:22:07] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [08:22:08] T431682: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682 [10:41:47] !log project-proxy merging https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328644 to move novaproxy backend selection and proxying to haproxy T429930 [10:41:51] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Project-proxy/SAL [10:41:52] T429930: Convert dynamicproxy from nginx to haproxy - https://phabricator.wikimedia.org/T429930 [12:00:33] !log lucaswerkmeister@tools-bastion-15 tools.quickcategories deployed 5cc772f346 (use MariaDB INSERT…RETURNING) [12:00:38] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.quickcategories/SAL [13:04:29] !status cloudvps proxy dows [13:04:33] !status cloudvps proxy down [13:06:02] !log admin stopped keepalived on proxy-5 to force a handover to proxy-6 [13:06:10] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:06:34] !log admin stopped puppet on proxy-5 to avoid restarting keepalived [13:06:39] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:06:58] !status ok [13:18:03] !log paws cordon and reboot paws-127d-mq2566hee6zs-node-1 [13:18:05] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Paws/SAL [13:22:57] !log admin failover dumps-nfs from clouddumps1002 to clouddumps1001 [13:23:01] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [13:28:47] !log paws reboot remaining worker nodes to pick up nfs changes [13:28:49] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Paws/SAL [14:06:48] !log project-proxy update metricsinfra scrape names to not reference nginx specifically T435917 [14:06:52] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Project-proxy/SAL [14:06:52] T435917: rspec-puppet does not support Debian Trixie (13) facts - https://phabricator.wikimedia.org/T435917 [14:50:20] !log project-proxy update metricsinfra scrape names to not reference nginx specifically T429930 [14:50:24] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Project-proxy/SAL [14:50:24] T429930: Convert dynamicproxy from nginx to haproxy - https://phabricator.wikimedia.org/T429930 [16:22:19] !log bd808@tools-bastion-14 tools.wheres-my-code-running Built new container for 6da7257e [16:22:21] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.wheres-my-code-running/SAL [16:23:54] !log bd808@tools-bastion-14 tools.wheres-my-code-running `toolforge webservice restart` to pick up new container. [16:23:54] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.wheres-my-code-running/SAL [18:08:26] !log bd808@tools-bastion-14 tools.versions Restarted webservice after reports of slowness. Still seems slow. [18:08:28] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.versions/SAL [18:14:43] !log bd808@tools-bastion-14 tools.versions Update to 24fb23a6 (T401818) [18:14:49] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.versions/SAL [18:21:23] Hello! I'm trying to launch a new acme-chief instance in deployment-prep but it's failing to resolve a hiera value (profile::acme_chief::cloud::designate_sync_password). That value is in labs/private but it's seemingly not picking it up. [18:52:53] brett: the easiest thing is for you to add an override in horizon. You can add it to the acme-chief prefix or just to your VM [18:53:09] andrewbogott: not for a secret password [18:53:21] true! [18:53:29] but labs/private is not secret [18:54:02] brett: if you're talking about a real password then don't do what I said :) [18:54:08] except when it is [19:58:36] andrewbogott: It is a secret password [19:58:43] Why is it not picking up the value? [19:58:51] There's no override in the existing acme-chief hosts [19:58:54] ok. So you're talking about labs/private on the puppetserver? [19:58:58] yeah [19:59:17] what file is the pwd in? [20:00:49] git/labs/private/hieradata/common.yaml right at the top [20:00:56] brett: this is your problem: [20:00:57] taavi@deployment-acme-chief07:~$ rg server /etc/puppet/puppet.conf [20:00:57] 15:server = puppetmaster.cloudinfra.wmflabs.org [20:01:07] the server isn't actually configured to point to the deployment-prep puppetserver [20:01:24] I wonder why that is [20:02:34] AFAICT the old acme-chief hosts aren't using any overrides, they're just using the prefix [20:04:28] presumably the puppet compilation failure means that the change to the project puppetserver isn't getting applied either [20:05:25] So it's a circular dependency, and a bogus secret needs to be added to the puppetmaster.cloudinfra.wmflabs.org labs/private? [20:06:37] that way it applies anything at all to compile, which will then set up the proper puppetserver endpoint, which will then give the proper password? [20:07:20] that's one way of solving it, yes, and would also fix these hosts in PCC and similar [20:07:31] Or is this an issue in which the puppetserver endpoint needs to be updated somewhere and that puppetmaster is since dead? [20:08:17] you're right that it's a chicken/egg thing. A dummy dependency is probably easiest, that or remove the prefix config while standing up the host and then replace it. [20:09:07] I guess the dummy is the way to go. I don't have access to the puppetmaster, could someone do that for me? [20:09:30] no, i mean add it in the git repo [20:11:27] oh, gotcha [20:31:25] !log deployment-prep T401839 - provisioned deployment-docker-wikifeeds01 (trixie) using Tofu (cloudvps-repos/deployment-prep/tofu-provisioning) [20:31:30] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [20:31:31] T401839: Migrate deployment-prep away from Debian Bullseye to Bookworm/Trixie - https://phabricator.wikimedia.org/T401839 [20:44:13] All good now. Thanks for the help! [20:48:07] brett: hm, I am not seeing any changes for review on labs/private? [20:52:02] er, I did a direct push, not with refs/for/master [20:52:29] https://gerrit.wikimedia.org/r/plugins/gitiles/labs/private/+/cdc68fb6637c27d5893879d5c65478bc68845f26%5E%21/#F0 [20:52:44] I mirrored what was in the private private one [20:56:54] apologies for that, I was in gitlab-mode and not gerrit-mode. I didn't even know it would let me push like that. Do I need to do anything to remedy that with gerrit? [21:00:30] brett: the general problem you hit here is the the steps to switch to a project local puppetmaster are often needed before applying a role and yet we have Beta Cluster setup to switch everything to project local and also name based config to apply roles. [21:00:51] puppet can be a real pain outside of prod (and inprod too, but for different reasons) [21:03:25] !log deployment-prep Switch acme-chief active to deployment-acme-chief07, passive to deployment-acme-chief08 [21:03:28] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [21:06:34] Okay, the new acme-chief instances are up and running, and I've updated the hieradata for the project puppet to use acme-chief07 now [21:06:56] great! [21:22:11] having some acme dns issues, so I'm reverting it for now [21:23:13] !log revert deployment-prep acme-chief switch: acme-chief active to deployment-acme-chief05 and passive to deployment-acme-chief06 [21:23:14] brett: Unknown project "revert" [21:23:20] !log deployment-prep revert deployment-prep acme-chief switch: acme-chief active to deployment-acme-chief05 and passive to deployment-acme-chief06 [21:23:22] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [21:29:34] ah, it looks like custom authdns_servers were set on the previous primary [21:30:43] reverting the revert now - before I needed it to be set to primary so that puppet would run but now I was able to do it with a disabled puppet so it still thought it was primary [21:35:08] !log deployment-prep Switch acme-chief active to deployment-acme-chief07, passive to deployment-acme-chief08 [21:35:08] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Deployment-prep/SAL [21:47:37] I'm keeping acme-chief05/acme-chief06 around. If there's an issue please do poke me but just in case I'm not around you can change the hiera values for the acme-chief prefix to 05 as primary and 06 as secondary, then the project prefix to 05 [21:47:46] (and then run puppet on em all)