[07:20:38] Hey DSE and AUX maintainers, if you haven't seen yet: You got unlucky and one of your etcd instances was running on ganeti2046 when it failed (https://phabricator.wikimedia.org/T434681) so your etcd is currently in degraded state (since we don't use DRBD there) [07:43:24] <_joe_> that means you have to likely remove the node and add a new one to the cluster [07:48:05] what that guy said [13:31:22] Oh hai! I just spotted this. Checking now. [13:34:47] Was there an alermanager alert for a degraded etcd cluster? I'm looking for one, but alerts.w.o makes my head spin. [13:42:11] btullis: I was thinking the same. Mainly because I could not find any alerts for the aux and dse nodes being down [13:42:21] but that was because I wasn't looking at icinga :D [13:42:54] heads-up: https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1325484 - I will change the deployment strategy for calico-node. heads-down: that will only affect wikikube prod :) [13:50:07] btullis if I can help with that LMK [13:50:23] Also stopped by to plug Ben's new k8s dashboard ;) https://grafana-rw.wikimedia.org/goto/afv2aidpw7pc0c?orgId=default [13:54:11] much color :D [14:00:35] inflatador: Yes, let's work together on this. [14:02:19] I don't think we have documentation for k8s etcd, but https://wikitech.wikimedia.org/wiki/Etcd#Adding_a_new_member_to_the_cluster might be enough. [14:02:59] Great, yes that's just what I was looking at. [14:03:18] or https://wikitech.wikimedia.org/wiki/Kubernetes/Clusters/Add_or_remove_control-planes#Add_stacked_control-plane for stacked controll planes. If the process heavily derives, a dedicated k8s documentation would be a blast [14:10:36] Please make sure you plan on upgrading istio in your clusters to 1.29 (https://phabricator.wikimedia.org/T427401) - running on wikikube since a week [14:11:41] and also plan on upgrading calico to 3.30.7 (https://phabricator.wikimedia.org/T427400), running in staging since a week. Just rolled out to wikikube codfw [14:13:21] Cool, will do. This is our ticket for the etcd node work. https://phabricator.wikimedia.org/T434793 - inflatador is going to drive it. [14:14:50] Sweet. If you happen to have another hour to spare I bet aux people would be delighted if you would do one for them as well :) [14:15:58] Depending on how much time it takes, I can probably do that. Easy for me to say now, but I kinda wish we had ceph storage and images for our VM environments ;P [14:19:30] well...you can have DRBD if you want to live with the additional etcd latency [14:22:10] DRBD is not on my list of things I want for many reasons ;P [14:26:48] number one would be supporting it at Rackspace for several years [14:59:34] moritzm or anyone else in IF, is there a way to remove `ganeti2046` from the ganeti cluster? The makevm cookbook is failing from `Can't get data for node ganeti2046.codfw.wmnet` [15:00:38] ah, looks like https://wikitech.wikimedia.org/wiki/Ganeti#Failed_hardware_node [15:17:30] I can't see any SAL entries for the failed hardware node stuff, so I'm gonna be bold and do it now [16:02:30] cdanis jayme as y'all are probably aware, `aux-k8s-etcd2003` and `kubestagemaster2005` are on the failed ganeti node (ref T434681 ). It looks like I'm gonna have to remove those instances completely before we can start to replace them. I'm happy to run the `makevm` cookbook after I remove, I just wanted to make sure it was OK frist [16:02:31] T434681: ganeti2046 doesn't come back after reboot - https://phabricator.wikimedia.org/T434681 [16:03:20] If y'all are the wrong people to ping for those VMs LMK, sorry [16:03:37] you'll probably need to silence a few things but in theory all of those hosts are disposable [16:05:43] https://wikitech.wikimedia.org/wiki/Etcd#Reimage_nodes_a_cluster has the right idea, remove, reimage, etcdctl add, hopefully that Just Works [16:05:55] (or instead of reimage, makevm and install, etc) [16:06:14] shouldn't need to mess with etcd::cluster_bootstrap at all [16:08:02] I think the new control plane will just work? https://wikitech.wikimedia.org/wiki/Kubernetes/Clusters/Add_or_remove_control-planes doesn't seem to be updated with the etcd-split-out (which I think applies to staging but not 100% on that) [16:08:44] oh, you'll also have to destroy their current puppet cert, I think [16:09:33] ACK, I'll give it a shot. My main goal is to unblock VM creation for my own selfish purposes. And I'm happy to do the etcd stuff for aux as well. But I'm not sure I'll have time to deal with `kubestagemaster` [16:10:21] ack, maybe someone from serviceops should deal with that [16:11:00] Agreed. I'll do my best and keep the ticket updated, but might have to handoff the kubestagemaster stuff [16:11:28] swfrench-wmf rzl ^^ FYI [16:13:07] kubestagemaster is not a real problem for now since it's just staging-codfw which is basically a k8s test cluster [16:13:10] have to go AFK for the next ~30 or so, will get started after that [16:13:28] * swfrench-wmf is reading backscroll [16:13:38] I'd leave it like this and take the chance of the ganeti node coming back :) [16:14:17] +1 to jayme's point, that kubestagemaster2005.codfw.wmnet isn't a problem per se [16:27:07] cool :) [17:12:42] Back...starting now. If anyone is aware of any dragons waiting for me LMK. I'm curious to see what happens when I run `makevm` after I `gnt-instance remove` [17:46:13] I was afraid of that...`makevm` is unhappy because the existing VM is in netbox. I'll try a full decom of `dse-k8s-etcd2001.codfw.wmnet` and then a makevm and see what happens [17:47:04] did you already etcdctl remove it? [17:48:16] no, but I'm about to do that first [17:51:49] hmm, the curl command in the docs doesn't seem to work, wonder if that's for an older version [17:54:52] inflatador: probably the thing to do is to crib some etcdctl incantations from the cookbooks repo [17:55:15] https://codesearch.wmcloud.org/operations/?q=etcdctl#operations/cookbooks [17:55:27] cdanis ACK, I'm RTFM at the moment. I've done the RAFT thing with consul/nomad/vault so I don't feel completely out of my depth ;) [17:55:52] I think what happened was "we don't have to update the docs, we fixed the tooling" [17:56:42] Puppet is nice enough to automatically set up the endpoints and mTLS stuff on the etcd hosts [17:56:47] sre.k8s.reimage-stacked-control-plane calls `member remove {m}` via a wrapper that is literally just return f"ETCDCTL_API=3 /usr/bin/etcdctl --endpoints https://$(hostname -f):2379 {command}" [17:56:59] so I think just do that on the cli :D [17:59:09] yup. I'm looking for a command that will show the active master, I guess that concept doesn't exist anymore? Maybe just voter and non-voter [17:59:59] I think you can point it at any of the active replicas [18:00:25] or rather, at either of them lol [18:00:50] `etcdctl member remove $UUID` from `dse-k8s-etcd2002` did the trick [18:01:06] ahhh also I see in sre.k8s.__init__.py that you could run `etcdctl endpoint health --cluster` [18:04:20] well that's interesting...decom cookbook says `dse-k8s-etcd2001.codfw.wmnet` is not in Netbox, maybe a cleanup happened. Trying `makevm` again [18:12:50] made it past the netbox/DNS steps this time [18:19:35] oops, forgot to remove the ganeti node before trying the makevm again. And that's not working either...`Instance aux-k8s-worker2002.codfw.wmnet is still running on the node, please remove first`. But the output of `aux-k8s-ctrl2002.codfw.wmnet` shows it's associated with 2 other hosts. Hmm [18:22:31] ah OK, looks like I'll have to tell Ganeti to pick a new secondary for a few more VMs [18:42:13] Looks like the command to pick a new secondary can only work on one VM in the whole cluster at a time ;( . On that note I'm grabbing a quick lunch [18:42:41] is DRBD enabled on the other etcd nodes? [18:42:45] is that what's taking a long time? [18:43:39] no, I can't make a new VM until 2046 is completely out of the cluster, and ganeti won't let me remove it until I reconfigure the VMs that used it as a secondary [18:44:09] ah [18:44:38] https://phabricator.wikimedia.org/T434681#12213759 list of VMs here, I imagine it will take a few hours if these hosts are 1 Gbps (haven't checked yet) [18:48:02] looks like they are 10Gbps. Still looks like pretty slow going though, maybe because of the software RAID5 [18:50:43] inflatador: just saw this. I need to leave in a bit, but to move the secondaries you can do: [18:51:06] sudo gnt-instance replace-disks --new-secondary ganeti2045.codfw.wmnet $FQDN_OF_VM [18:52:37] we still have a handful of legacy nodes with just 1G NICs in the codfw cluster, which sadly caps the migration bandwidth in total, but on the upside Rob kicked off the procurement for the new hosts to replace these with modern servers today [18:53:01] I can also take care of the remaining drain steps tomorrow morning, but I won't have time for it now [18:54:04] also, if a downtime of the dse etcd nodes has such immediate blast radius, let's maybe move to a four node etcd cluster to improve redundancy? we can have one node in row A-D [18:54:33] never do 4 for paxos/raft, always do 3 or 5 [18:54:40] instead of ganeti2045 any other nodes in row A works, so you can just as well also use 2027, 2028, 2029, 2030 [18:54:54] ah, right. we can also do 5, these VMs are small and cheap [18:56:24] sorry, need to leave now, I can check IRC later if there's any more questions, or you can also leave it to me to drop the node from the cluster for the European morning [19:34:12] back [19:39:13] as far as Ganeti, I get that we can do 5 VMs, but that's not free either. I'm looking for a more cloud-like experience where I can immediately reprovision if the host dies and let y'all work on the host as time permits. We have some unused hosts that were gonna form a "ganeti-jumbo" cluster, but we've also been talking about evaluating some other VMMs. If anyone is interested in looking at Openstack/Incus/Proxmox whatever LMK [19:40:57] inflatador: I think if we had taken a different path it would have been a fair bit faster [19:41:50] but also I'm not sure, my mental model of Ganeti isn't great [19:42:21] I really want to encourage you to spend some more time going deeper on getting to know the technologies we do have installed already, though [19:44:07] I've gotten pretty deep with Ganeti today [19:45:01] yeah, I totally understand this has been a lot more trouble than anyone would reasonably expect [19:46:03] Openstack certainly hasn't been a free lunch either, in my experience, and I don't have any personal experience with Proxmox, but I've read more than a few upsetting threads while looking up grungy ZFS stuff [19:52:31] I was a virt engineer doing this kinda stuff all day every day at Rackspace for years, so 1) this is fun in a weird way for a weirdo like me, and 2) I'm very sensitive to getting deeper on technology that doesn't seem to have an ecosystem and a future [19:54:41] Openstack is horrible, that I fully understand...I just don't know how long we're gonna be able to keep the lights on with Ganeti. Google divested from it in 2020 and based on the mailing list it seems like a one-person shoq [20:00:56] anyway, almost done reconfiguring disks...will let you know when I try `makevm` again [20:53:32] OK, looks like 2046 is out of the cluster now, trying a `makevm` again [21:02:16] we're getting somewhere [21:33:02] I think the most interesting thing I learned is that the lack of a single Ganeti node can block VM creation. I'm guessing that is probably true of most VMMs (may have seen it on Proxmox), whereas with k8s or openstack the control plane is more separated and a single worker outage doesn't affect too much [21:54:19] `dse-k8s-etcd2001` is back in its cluster, I'll do `aux-k8s-etcd2003` once it's done reimaging [22:23:49] OK, `aux-k8s-etcd2003` is back in its cluster as well. I'm currently reimaging `kubestagemaster2005.codfw.wmnet` but I'll leave it up to Service Ops to restore that one as my day is almost done [23:02:33] I'm done, I left T434844 (kubestagemaster2005 restore) for j-ayme and sw_french-wmf [23:02:33] T434844: Restore kubestagemaster2005 to service - https://phabricator.wikimedia.org/T434844