[00:15:01] taavi: Something changed mid US day today that is making the deployment-prep tofu automation fail. https://gitlab.wikimedia.org/cloudvps-repos/deployment-prep/tofu-provisioning/-/jobs/941926 shows the "Plugin did not respond" error. [00:16:39] I was using tofu to test the php 8.5 stuff I was working on and built the instance I wanted with success. Destroying it by applying the config before I added it is where things started failing. [00:17:22] I'm mostly asking you because I am not very clear on how to debug the problem myself. [00:18:52] I think I am just going to manually delete the openstack bits that are not wanted for now and hope the crash is scoped just to this deletion. [07:44:36] greeetings [07:54:37] bd808: ugh. haven't seen that before. [11:40:05] Hello, I’m taking care of the pywikibot upgrade T431056, could someone please add me to https://gitlab.wikimedia.org/groups/toolforge-repos/-/group_members so I can update our mirror? (following https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Pywikibot_image) [11:43:28] T431056: [pywikibot-buildservice] New upstream release for Pywikibot - https://phabricator.wikimedia.org/T431056 [12:56:32] thilp: doing it now, sorry for the delay [12:58:10] done! [12:58:29] dhinus: thank you! [12:59:49] is that not a part of the onboarding checklist? [13:00:16] it was not, I’m adding a comment in mine for posterity [13:01:48] can you please add it here as well? https://www.mediawiki.org/wiki/Wikimedia_Cloud_Services_team/Onboarding_template [13:09:26] done! Thanks, didn’t know we had this template, I thought we copied from previous phab tasks :) [14:35:52] taavi: did you already chase down the deployment-prep tofu thing? [14:35:59] no [14:37:01] ok, I'll poke at it [15:13:28] I tentatively added a "check helm" step to https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Kubernetes/Upgrading_Kubernetes#Check_everything_looks_good [15:14:13] please adjust as you like... in theory it shouldn't bite us now that we locked the helm version, but we're gonna upgrade helm and/or helmfile again at some point [16:02:58] bd808: my first guess is that something is messed up with volume attachment (because that's the weakest link in the things that tofu run is doing.) Did you already get things working by manually removing/detaching things? [16:12:07] andrewbogott: I did clean it all up manually. And I could see it having something to do with cleaning up the volume. I have dropped instances with volumes before, but I haven't tried it for a while. [16:13:37] OK. There have been a few different issues with sticky volumes; I thought they were all fixed. Was this a mount that existed for a long time? [16:46:52] andrewbogott: no, the instance was created around 18:00 UTC and things were failing to clean up by 21:0 UTC [16:47:08] that's concerning [16:47:35] Is it easy for you to retry and see if it happens again? I don't want my curiosity to get in the way of deployment-prep progress elsewhere... [16:50:18] I need to work on some other things today, but yeah I can probably run stuff in the background to make a similar instance and then try to kill it again. [16:50:40] * bd808 starts making the thing that may or may not delete later [16:52:50] thx [20:02:03] andrewbogott: create and delete worked as expected today. I did find a novel new way to make tofu sad, but it only ended up making it impossible for me to adjust hiera settings on the instance. [20:33:09] ok. thanks for retrying [21:26:45] andrewbogott: I'm back to a "plugin6.(*GRPCProvider).ApplyResourceChange' error. The steps this time were 1) provision 2 new hosts, 2) try to provision local hiera settings for those hosts. Step 1 worked, step 2 is failing. [21:27:45] I think that the first time this all happened I had basically done the same steps, but step 2 worked that time and a step 3 of deleting started the failures. [21:28:15] But I'm not thinking this is more likely something in the hiera settings interactions than volumes. [21:28:22] hm, so no volume attachment involved this time? [21:28:30] https://gitlab.wikimedia.org/cloudvps-repos/deployment-prep/tofu-provisioning/-/jobs/943112 is the current failure [21:28:44] andrewbogott: there are volumes, but they happened in step 1. [21:28:55] I am overall surprised that tofu can interact with hiera. I guess someone wrote a custom plugin to talk to our hiera api? [21:30:11] t.aavi wrote the tofu logic to talk to our hirea stuff and the WMCS http proxy [21:30:44] https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-cloudvps [21:32:27] The hiera things look like https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-cloudvps/-/blob/main/examples/resources/cloudvps_puppet_prefix/resource.tf [21:33:42] is this run happening on some server where I can see the logs, or is happening in magical CI cloud container land? [21:34:19] in deployment-prep that is inside the "cluster" stuff at https://gitlab.wikimedia.org/cloudvps-repos/deployment-prep/tofu-provisioning/-/blob/main/modules/service/main.tf?ref_type=heads#L50-67 and then configured like https://gitlab.wikimedia.org/cloudvps-repos/deployment-prep/tofu-provisioning/-/blob/main/main.tf?ref_type=heads#L56-70 [21:35:18] The opentofu part can be run locally or in the GitLab CI. [21:36:19] The tofu side logs are in https://gitlab.wikimedia.org/cloudvps-repos/deployment-prep/tofu-provisioning/-/jobs/943112 I think. I might be able to make them noisier locally. [21:36:48] `2026-08-28T21:20:55.472Z [ERROR] provider.terraform-provider-cloudvps_0.4.0_linux_amd64: panic: runtime error: index out of range [0] with length 0` looks like where things go wrong. [21:37:39] It says 'Plugin did not respond' -- I assume 'Plugin' is taavi's provider in this case? [21:37:46] Ah, ok. But not stack trace huh? [21:38:01] That is in the stuff t.aavi wrote. I will make myself a bug report and then try to get done with other things so I can make time to dig in the golang [21:38:08] bd808: does the extra egress firewalling on gitlab runners include rules for the enc and proxy apis? [21:38:18] andrewbogott: yeah there is a trace after that error line. [21:39:03] and yeah, now that you mention it I suspect I forgot to set any kinds of explicit HTTP timeouts in the code in the provider/go library :/ [21:39:31] taavi: that wouldn't cause the 'it worked for a while and then broke' behavior though... [21:39:36] taavi: They must somehow. The hiera interactions do not always fail. It seems to have something to do with adding the tofu side hiera after the initial instance was provisioned. [21:41:05] https://phabricator.wikimedia.org/P96274 is just the error log bit from the failing run [21:41:27] listElmType, err := ctyTypeForValue(value.([]any)[0]) [21:41:30] so probably value is null there [21:42:23] my current hunch is the instance provision makes empty hiera settings on the enc side and then later tofu doesn't deal with that properly. [21:43:33] I tried in a run earlier today to make explicitly empty hiera setting as part of a tofu build and tofu did not like the data that came back from the enc. [21:45:01] So far the trace has me in yamlEncodeHiera(prefix.Hiera) [21:45:17] So... maybe there's a prefix that's defined but empty? Although that shouldn't be a problem ideally [21:45:46] andrewbogott: yeah, that's my guess about what makes it blow up [21:46:20] what's the name of the VM that's choking on t his? [21:46:23] Possibly just needs a guard to see if the API response is empty and do something different [21:46:49] deployment-mediawiki15.deployment-prep.eqiad1.wikimedia.cloud and deployment-mediawiki16.deployment-prep.eqiad1.wikimedia.cloud [21:47:24] the deployment-mediawiki prefix sure isn't empty [21:48:02] andrewbogott: it would be the instance local prefix that is named with the FQDN of the instance [21:48:19] oh right, I forgot that that is also a 'prefix' [21:49:17] ...also not empty though [21:49:31] well hiera isn't, I assume this is all hiera we're talking about [21:51:05] "also not empty though" -- what does it look like? [21:51:52] bd808: I have adjusted the task description in T401839 to include per-server notes. To avoid conflicts caused by deploying VMs using the share OpenTofu state you're using for the new appservers, I'll refrain from using OpenTofu today - it's yours. I'll work on creating a few subtasks instead. [21:51:53] T401839: Migrate deployment-prep away from Debian Bullseye to Bookworm/Trixie - https://phabricator.wikimedia.org/T401839 [21:51:55] well, horizon at least shows deployment-mediawiki16 with [21:52:12] https://www.irccloud.com/pastebin/ygWGLRJf/ [21:55:09] Southparkfan: thanks for the public plan there. and sorry I have tofu jammed in a way that blocks you. I was hoping that we could both use it via the gitlab-ci to keep the locks in sync. And then I found a new way to make it sad. [21:56:18] andrewbogott: oh that is interesting. That is the hiera that my patch was trying to provision. So it worked outbound to the enc, but something went wrong later. [21:57:04] Are we seeing this after only one attempt? [21:57:26] yes, one run of the plan that starts at https://gitlab.wikimedia.org/cloudvps-repos/deployment-prep/tofu-provisioning/-/jobs/943112#L1015 [21:59:28] the plugin is failing at the very end, at the validation step. So it wrote the thing, then checked to see if it was written, and crashed out. That shouts race condition... [22:00:02] It gets the prefix content which doesn't return an error code but does, presumably, return emptiness. [22:00:29] https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-cloudvps/-/blob/main/internal/provider/puppet_prefix_resource.go?ref_type=heads#L393 [22:02:31] I think if I run the plan again it will blow up differently. I think a second run will get an error back from the enc side that the prefix already exists. [22:03:53] meaning the plugin tries to create the prefix even if the prefix is already there? [22:04:09] Not an issue, take your time. I didn't anticipate getting much of the work done anyways. And I do not want to slow down your debugging process :-) [22:04:16] `// The create API does not support setting hiera or class data (at least at the moment), // so do it with an update request afterwards` sounds like it could be racy [22:04:33] andrewbogott: hang on... let me check something... [22:04:59] (I should disclose at this point that I'm very unlikely to try to learn how to build and deploy this go plugin anytime soon) [22:05:18] the plugin does a create and then an update... [22:05:57] oh, and oddly it does the same validation check after the create, and that passed. [22:06:10] I built it locally and attached it to tofu once. There should be notes someewhere in phab. [22:06:28] So an empty prefix must work ok... [22:06:58] andrewbogott: so back to the "meaning the plugin tries to create the prefix even if the prefix is already there?" question-- right now the tofu graph does not know that the prefix exists. [22:07:05] mostly we need the plugin to log "Here's what that prefix contains" before trying to parse it. [22:07:29] bd808: it doesn't know because of the previous failure? It doesn't re-check the state before applying changes? [22:07:33] so the next run will walk the graph and say "I need to make one of these" and I know that errors. [22:07:47] * andrewbogott kind of thought that tofu updated the graph before doing anything else [22:08:22] andrewbogott: it checks its graph, but does not try ask OpenStack about all the things it find missing, it just tries to make them. [22:08:48] that's silly :) [22:08:52] the 'graph" here is the tofu maintained state, not the actual state [22:09:05] yeah, that always bothers me. 'tofu init' refreshes that right? [22:09:36] I think that dealing with a "oops that exists already" error is a thing tofu expects our plugin to do [22:10:02] `tofu init` loads the state from wherever it was saved [22:10:32] `tofu deserialize-graph-from-stroage-or-make-empty` was too long ;) [22:11:07] But... I'm sure that I've started with a fresh install and checkout of the tofu code and it didn't say "well now I have to create every single thing" [22:11:21] is there another tofu state stored someplace between openstack and my local tofu install? [22:11:39] because you fetched a prior state graph from somewhere, like our S3 buckets [22:12:19] ... [22:12:20] I hate this [22:12:22] https://gitlab.wikimedia.org/cloudvps-repos/deployment-prep/tofu-provisioning/-/blob/main/main.tf?ref_type=heads#L4-6 is where the beta cluster state is [22:12:31] I mean, it's not related to your immediate issue. But I hate it. [22:12:43] Why have an intermediate source of truth? [22:13:45] because if you have a deployment with 1,000,000 existing AWS/OpenStack/Whatever objects there is no "ask the provider to list all the things that exist in the realm" api call [22:14:29] and you could have N different tofu states to build out your full enterprise [22:14:40] It doesn't need to check everything that exists, only everything that is defined in the current tofu catalog. [22:15:20] I mean, it's fine to say "that's expensive so we need a cache" but that's different from saying "that's expensive so we need a cache and also there's no way to refresh the cache" [22:15:33] I think the system design is expecting "create a thing that already exists" to be idempotent and our plugin is not [22:15:51] hmmm [22:16:04] yeah, that's fair, although it still means that 'tofu plan' will lie to you [22:16:30] So indeed the plugin does not seem to check whether the prefix exists before creating it. So that's an easy fix, modulo build-and-deploy [22:16:46] That wouldn't fix the most recent crash but it would allow you to just re-run and get past it. [22:17:58] that's a matter of perspective :) `tofu plan` will tell you what it needs to CRUD in its local state graph to achieve the plan. It just doesn't know at that point if the outside world matches it's state graph. [22:19:09] So tofu plan means "I will create some subset of the following" not "I will create the following things that don't exist" :/ [22:19:24] do we have a task for tracking fixes to the plugin? [22:19:54] We just make tasks like T398643 I think [22:19:55] T398643: [tofu-cloudvps] cloudvps_puppet_prefix.hiera settings show dirty diffs based on YAML canonicalization - https://phabricator.wikimedia.org/T398643 [22:21:44] grrr... I was hoping https://phabricator.wikimedia.org/T398643#11067045 would actually have steps for the local testing. I wonder if I have local notes still. [22:23:34] T436412 [22:23:35] T436412: [tofu-cloudvps] cloudvps_puppet_prefix.hiera is not idempotent and should be - https://phabricator.wikimedia.org/T436412 [22:26:05] I can't tell if we will also wind up adding the same keys multiple times... but probably [22:34:47] andrewbogott: you have definitely helped me see the problem better. I need to hop back to a different thing, but I will try to circle back to this over the weekend. If I can remember how I wired a custom build of the plugin into my local tofu I think I can find a hack at least. [22:35:23] ok! I hope that I helped and did not just distract [22:35:35] andrewbogott: when you looked at the enc data to see what was there, how did you do that? [22:36:28] Just looked at the 'puppet' tab in Horizon like a chump. [22:40:59] heh [22:43:05] I know the prefix stuff is exposed as an openstack api endpoint. if I can learn how to use the openstack cli to talk to it then I can at least make an unblock runbook to hack around thing by manually updating the tofu state graph. [22:43:47] I'm not sure that there's a cli for the heira stuff even though there is a standard API. You might need to curl your way to victory. [22:43:59] But you can use the cli to get your initial token which is the annoying part [22:44:06] :nod: [22:45:11] andrewbogott: can you teach me / point me to where to get a bearer token using the cli? [22:46:07] root@cloudcontrol1011:~# openstack --os-cloud novaadmin token issue [22:46:20] do you have a clouds.yaml to work with --os-cloud already? [22:47:21] clouds.yaml docs are here although oddly there is no example... [22:47:22] https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Openstack_cli#clouds.yaml [22:47:34] yeah, my `make shell` working environment has all of the betadevopsbot credentials in place [22:48:03] `openstack token issue -f value` dumps out what I need [22:48:14] ok, then it's really as simple as 'token issue' [22:49:12] past me already knew this stuff... https://wikitech.wikimedia.org/wiki/User:BryanDavis/OpenStack#API_access_with_curl [22:49:27] (found while going to make a note) [23:05:26] `curl -H "Accept: application/json" -H "X-Auth-Token: $(OS_DEBUG= openstack token issue -f yaml | yq .id)" "https://puppet-enc.cloudinfra.wmcloud.org/v1/deployment-prep/prefix?detailed" | jq` -- that gives me the things I need to update the state graph!