[08:28:31] looking at the grafana dashboards, it's me or now the timezone is always set to browser timezone (it's possible that before upgrade that was pinned to UTC or am I still sleeping?) [08:43:21] hey on-callers, I merged https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329572 that is related to cfssl::cert, so very broad scope. If anything doesn't work it will be visible in puppet failures, in case lemme know! [08:46:25] fabfur: I checked https://grafana.wikimedia.org/d/slot-pilot-slo-detail/sloth-slo-detail? and it is in UTC, but it may be depending on the pre-defined time window saved? Is it a specific dashboard or all? [08:47:27] nope, it's only the home page dashboard, others are in UTC as expected. At this point is most probable that it's me (I remember also that dashboard being UTC but may be wrong...) [14:47:37] I have a server that, after a reimage, is now in a crash loop. You can re-reimage it, so it will boot PXE, but afterward when switched to boot from disk it's back to the death loop (and if an error is displayed, it's not for long enough to read over the serial console). No errors in the drac log either. Does this ring a bell for anyone? [14:50:43] urandom: does it have software raid [14:51:04] eys [14:51:05] yes [14:51:49] it smells like the root partition is not recognized or similar, so the host keeps rebooting etc.. but if you connect to the mgmt console you should see the error right when booting [14:51:50] what partman configuration are you using [14:53:04] elukey: I can't see any errors before it reboots again... there might be one, and it's just not there long enough to paint over the console [14:53:26] what is the host? [14:53:26] cdanis: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329659 [14:53:33] elukey: cassandra-dev2002 [14:54:25] now that I think of it, maybe there is a bios setting for halting on errors? [14:55:23] I am attached to the console, but I only see blank [14:55:40] I powered it off, rather than let it loop in perpetuity [14:55:47] can this be an efi/non-efi boot issue? [14:56:30] oh, yeah, good point [14:56:50] yeah, I don't think it's the software raid grub thing [14:56:58] I am booting it up [14:56:58] although, I'd successfully used this same partman recipe for cassandra-dev2001 and the hardware is the same [14:58:09] urandom: we saw that as well for the cp nodes and then moved the recipe to UEFI and it worked [14:58:40] UEFI is the default now so you have to manually specify BIOS (I don't recall exactly where) [14:58:55] so perhaps adapting that recipe and then re-attempting a reimage may work [15:00:09] sukhe: this is a Dell r440, so BIOS I'm sure [15:00:45] yeah the boot is cryptic, it ends up in a blank screen and then it restarts [15:00:53] from the boot menu it seems that the disks are recognized [15:01:00] this partman recipe is UEFI though [15:01:15] elukey: yeah, not helpful at all [15:01:15] ah ok then it is a problem :D [15:01:26] are you using a UEFI partman recipe? [15:01:30] yes [15:01:52] yeah then you need to run the provision cookbook to set UEFI [15:02:15] otherwise the host will try to do a legacy boot not finding what to start [15:02:55] sudo cookbook sre.hosts.provision cassandra-dev2002 --no-user --no-dhcp --no-switch [15:02:59] urandom: [15:02:59] 'cassandra-dev*': [15:02:59] - reuse-parts.cfg [15:03:00] - partman/custom/reuse-cassandra-jbod.cfg [15:03:00] this will set the host to UEFI [15:03:06] not entirely sure but this is not UEFI? [15:03:12] I may be wrong [15:04:02] no idea if reuse-parts supports uefi [15:06:31] reuse-cassandra-jbod.cfg corresponds with cassandra-jbod.cfg, which is uefi [15:08:18] worth mentioning though that these partman recipes were all tested (base and reuse) when created on this cluster, and previous to this I did cassandra-dev2001, which worked fine [15:10:15] elukey: any reason *not* to set it uefi? [15:14:29] uefi is preferred, what are you setting to uefi in this context? [15:14:44] cassandra-dev2002 [15:15:15] thanks [15:29:33] I'm happy to take a look as well urandom if you are unsuccessful [15:57:13] urandom: sorryyy I was in a meeting [15:57:39] ok, I ran the provision cookbook to set it uefi, then re-ran the reimage cookbook, and it's erroring because no efi partition was found [15:57:50] which makes sense I think [15:58:10] yeah I think you'd need to use the non-reuse recipe [15:58:26] but if it worked for cassandra-dev2001 that is the same hw, maybe there is something wrong with the host itself [15:58:39] yeah [15:59:28] we can also check if they run the same idrac/bios firmwares, just to be sure [15:59:56] but in general enabling uefi may or may not end up to be the right bet, since we are not sure what the problem is now [16:00:31] I can re-run the provision cookbook w/ --legacy to get it back to bios, yes? [16:00:37] correct [16:03:08] fwiw the bios versions of 2001,2002 is the same [16:04:59] at this point you may try to install the os with a non-reuse recipe even for legacy, and see if it fixes [16:30:53] elukey: so if that also fails, hardware, but if it doesn't...? :) [16:31:16] but if it doesn't... "confusion". [16:37:30] what in the actual...? [16:38:52] Ok, so provision cookbook to change it to uefi and reimage, which did not work for obvious reasons, then cookbook it back to bios, and reimage just to verify a return to the status quo... and it has booted [16:39:40] I retried the reimage many times prior, the only thing that is different was the roundtrip through the provision cookbook [16:39:55] gawd I hate hardware [16:45:39] lol [16:45:55] so with re-provision (legacy) + reimage it worked? [16:46:07] maybe there was something misconfigured, who knows [16:46:10] good that works! [16:53:27] I guess. I mean "works" is the state I was after. I hate not understanding why tho! [17:51:54] Am I the last SRE to still use keepassx 0.4.4 or are others still loyal? I ask because last I checked (probably a decade ago) the newer versions were considered untrustworthy due to an un-reviewed rewrite... but 0.4.4 is about to stop working on apple silicon [17:52:49] I use keepassxc [17:53:07] I am not aware of the differences but it serves me well since it doesn't have any of those sync features that I don't want [17:54:08] I take it that's the successor to what I'm running... do you just run the latest version? Does it come from the apple store or is it a direct download? [17:56:16] which version is in Debian unstable, which is what I run on the laptop :) [17:56:21] Version: 2.7.10+dfsg1-2.1 [17:56:42] and then you can sync the passkey with syncthing or something if you really want that [18:00:24] I don't need syncing, I'm just hoping for something that will import my existing db. Will give it a go! ty [18:04:20] gl : [18:04:20] :)