[07:50:33] Morning! [08:10:13] morning [15:32:31] andrewbogott: is there a way to force a new openstack server to run on a certain hypervisor? trying to set up the cloudinfra etcd cluster and by default the scheduler tries to run two servers at the same host (which fails the server group check) [15:33:37] huh, the server group should prevent that at schedule time... [15:33:47] but yes, let me make sure I'm getting the syntax right... [15:34:45] yeah let's at least get the 3 etcd on different hosts :D [15:35:24] Try adding "--availability-zone host:cloudvirtlocal1001" to the cli [15:35:31] I'm not sure there's a way to do that in horizon [15:36:41] alternatively you can target a host when you migrate a server (which in this case would be cold migrate but I think I've done that with etcd nodes) [15:37:01] hrm, now how do I pass that through the cookbooks :D [15:38:14] With godog's rack work we may want to use --availability-zone more often in the future so might be worth adding to the cookbook ui [15:39:48] well, this sounds like an excuse to plumb that through the wrappers :D [15:39:54] :S [15:39:58] *:D [15:40:16] it's too easy to nerdsnipe you [15:40:35] I figured you were going to do that anyway, I'm just trying to provide you with extra justification [16:04:49] neat, that worked, thanks andrewbogott [16:06:48] I think I still need to make the cookbooks manage a SRV DNS record for confd to use, but otherwise there is now an etcd cluster [16:10:48] \o/ [16:13:07] andrewbogott: https://gerrit.wikimedia.org/r/c/cloud/wmcs-cookbooks/+/1341932 [16:48:53] I kept scrolling expecting to see it in a command-line format someplace but that cookbook actually uses the API! When did that happen? [16:54:51] no, the cookbooks still use the openstack CLI, it's just that the ensure canaries cookbook had a use for that so they implemented it to the api wrapper [16:55:23] alas [16:56:05] godog: any idea about T438062? (haven't looked into it any further than noticing the ticket) [16:56:06] T438062: Dumps NFS mount down - https://phabricator.wikimedia.org/T438062 [17:01:03] taavi: no, taking a quick look [17:02:06] thanks! [17:05:10] * dcaro off [17:05:12] cya tomorrow! [17:11:33] taavi: no tbh I can't quite figure how/why the kernel got wedged and now can't mount with [Errno 19] No such device: '/mnt/nfs/dumps' [17:12:57] but yes as the user said the nfs server is working and serving data, I was looking at https://grafana.wikimedia.org/d/nfsd-overview/nfsd-dumps-overview?from=now-6h&to=now&timezone=utc [17:13:19] looking at paws too [17:18:44] godog: is user id 400 correct? [17:18:48] drwxr-xr-x 2 400 400 4096 Jul 14 08:26 dumps [17:18:53] at least on tools-bastion-15 [17:19:03] yes that's dumpsgen user on clouddumps [17:20:40] but yes the mountpoint is wedged on paws too, I noticed both clouddumps are pooled, I'll depool one [17:27:54] ok clouddumps1002 depooled, restarted nfs-kernel-server there and no change whatsoever [17:28:48] godog: shout if you need a hand for anything [17:29:12] thank you volans ! [17:29:14] will do [17:30:40] could be related? [17:30:41] 2026-09-15T17:27:16.983988+00:00 tools-bastion-15 puppet-agent[1039325]: (/Stage[main]/Profile::Dumps::Nfs_client/File[/mnt/nfs/dumps]/owner) owner changed 'root' to 400 (corrective) [17:31:40] I doubt it, I suspect that's because the directory was unmounted [17:31:51] i.e. owner root on the bastion's filesystem [17:32:10] ok, syslog is full of nfs related logs, trying to make sense of them [17:34:04] on the bastion? yes that was me temporarily turning on nfs client debug [17:34:34] except that there are other nfs mountpoints so the log is hard to make sense of [17:35:05] see https://etherpad.wikimedia.org/p/volans-tmp2 [17:35:18] it tried to remount it so in some way it failed [17:35:21] btw did it alert? [17:35:36] what's "it" ? [17:37:01] mnt-nfs-dumps.automount [17:38:15] it seems that was the first one on tools-bastion-15 [17:39:08] no alert afaics no [17:44:06] ok still no idea what's going on btw, I see clients connected to 1001 and yet no nfs operations [17:45:01] I'm draining and rebooting paws-127d-mq2566hee6zs-node-0 as a test [17:45:14] is nfs working fine on toolforge? [17:45:18] I can check the tracing stuff [17:45:36] yes please check [17:45:58] the TimeoutSec=3s seems pretty low btw on the mount unit in case there are slowness, but nc connected to port 2049 immediately [17:46:32] try to either mount it manually (even on another path temporarily) to bypass the timeout or change it live [17:47:43] just to be clear, is that an attempt you are doing or a suggestion ? [17:47:46] nfs tracing data doesn't show anomalies in the last 6h [17:47:52] a suggestion for you [17:48:02] not making as I don't want to step on your tests/tries [17:48:07] but I can try on another host if you want [17:48:43] sure yes please do try on another host, thank you [17:48:52] which one are you on?] [17:49:10] tools-bastion-15 and paws-127d-mq2566hee6zs-node-0 [17:49:13] k [17:50:06] ls /mnt/nfs is stuck on tools-bastion-14 (as opposed to -15) [17:51:43] and no mnt-nfs-dumps.mount tries in syslog [17:53:34] ack thank you [17:55:41] so can't try that there :/ [17:57:34] somehow I suspect this may have to do dumps refresh from airflow [17:58:36] judging from /srv/dumps/xmldatadumps_airflow_temp/xmldatadumps/public/ on clouddumps100[12] empty but today's mtime [18:02:50] ack [18:03:08] * volans 's dinners almost ready, not sure how long I'll be able to stick around [18:03:39] ack thank you anyways, I'll flip the active server to 1002 [18:06:29] what Idon't understand is why it looks like it's working fine on toolforge and broke elsewhere [18:07:13] like on tools-k8s-worker-nfs-66 it works fine [18:07:49] indeed [18:08:28] volans: does it work even if you say 'find' in the directory to make sure we're not seeing the cache ? [18:09:14] I did an ls -la inside itwiki/latest/ that shows a lot of files, but I can try other thigns [18:09:25] it took the usual ~5 seconds as there are a lot of files [18:09:32] yeah that's probably fine [18:09:45] didn't seemed cached [18:10:03] yeah I think we may be back, I flipped to 1002 [18:10:20] tools-bastion-14 now responds [18:10:38] hope I didn't jinx it [18:10:49] confirmed that tools-bastion-14 can see fine files [18:10:57] where ls was stuck before [18:11:19] yeah paws is back too [18:11:20] tools-k8s-worker-nfs-80 keeps working [18:11:58] tools-bastion-15 too works [18:14:01] I have this hunch that the airflow -> clouddumps process somehow messes with the nfs export directory inode and that greatly confuses things [18:14:43] it's totally possible if we export/mount the same directory they touch, hopefully they should just touch sub-directories [18:15:47] was this the first time airflow completed since the switch to the balanced service? [18:16:03] I don't know how to check for that, totally possible though [18:16:26] that's a problem for tomorrow-me, since it looks like we're okay I'll go to dinner [18:16:27] on ariflow? there should be the list of dags and for each the list of runs [18:16:33] +100 [18:16:34] :D [18:16:41] thank you for your help volans ! [18:17:56] I've not done much :)