[00:03:00] FIRING: [2x] SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:12:28] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12173259 (10Papaul) [01:22:50] 10SRE-tools, 10Cumin, 06DC-Ops, 06Infrastructure-Foundations, and 3 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12173276 (10wiki_willy) Thanks for looping me in. Before I approve though, just wanted to provide space for any comments or concerns tha... [02:23:00] RESOLVED: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:26:25] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:36:43] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12173316 (10Papaul) [04:20:15] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12173344 (10Papaul) [06:26:40] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:37:41] 10SRE-tools, 10Cumin, 06DC-Ops, 06Infrastructure-Foundations, and 3 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12173444 (10elukey) @Dzahn technically we are adding a sudo rule so the I/F team should +1 it, but I already anticipate it is a +1. I'll... [07:39:44] elukey: thanks for looking at T433635 ; if I do a factory reset of the BMC (via the web UI), is there a non-destructive way to restore our configuration to the BMC? And if so, shall I try that before poking dc-ops? [07:39:45] T433635: Unable to reimage or reprovision ms-be2082 due to redfish connection errors - https://phabricator.wikimedia.org/T433635 [07:41:14] Emperor: I was about to ask if I could do it, in theory everything can be restored via the provision cookbook so it is ok to do it [07:41:36] I am also unable to use ipmi with all accounts, there is something weird on that bmc [07:42:09] elukey: so hard reset and then run the provision cookbook without the usual (--no-dhcp --no-users)? ISTR that produces some alarming warning about the host not being in NEW state [07:45:03] elukey: alternatively, since you know what you're doing, shall I leave this to you? I'm happy if the host itself ends up getting rebooted a few times. [07:46:30] Emperor: yeah I can do it :) The trick for provision is to set the host's status in netbox to something different than active, so you can fully reprovision [07:46:32] lemme try [07:47:59] cool, thanks. [07:48:54] I am not able to use ipmi with others ms-be208x that is weird [07:49:19] anyway, one thing at the time :D [07:50:27] The first try is with a BMC reboot from the UI, maybe it does something different/more than redfish [07:55:01] 10netops, 06Infrastructure-Foundations: Investigate Nokia Apply-path for prefix-lists - https://phabricator.wikimedia.org/T433662 (10ayounsi) 03NEW p:05Triage→03Low [08:04:12] nope same [08:07:03] doing the reset preserving users [08:11:24] same, no improvements [08:13:44] Emperor: ok if I powercycle the host? [08:15:29] elukey: yeah, go ahead [08:47:44] Emperor: ok the reboot seems to have done the trick, I am provisioning [08:48:12] my understanding is that on supermicro x12 motherboards the bmc is just a bare proxy to bios, and sometimes a reboot is enough to retrigger a sync [08:48:45] once up I am fairly confident that you can reimage [08:52:19] elukey: cool, thanks. Is it useful if I try a reimage today? I usually avoid on Fridays, but if it's helpful to check this has worked I'm happy to [08:55:38] Emperor: I think it should be ok in this case, if the host stays down we are not risking much IIUC, plus we have both a bit of time to debug this (I am going on holiday for two weeks starting monday) [08:55:52] (provisioning is completing now) [08:57:53] cool. I have a meeting for the next ~hour. [08:58:24] 10SRE-tools, 10Cumin, 06DC-Ops, 06Infrastructure-Foundations, and 3 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12173638 (10elukey) 05Open→03Resolved Done! [08:58:42] Emperor: ah lovely now it fails for ipmi, sigh [08:58:54] I'll check what's wrong, I feel it is the same for other ms-be nodes [09:03:59] so the config J supermicro seem to not have any IPMI allowance now, both ADMIN and wmfroot user [09:07:46] I mean, presumably this host used to do IPMI, right, since we've (re)imaged it before [09:08:09] ? [09:56:26] 10netops, 06Infrastructure-Foundations, 06SRE: Consider removing BFD on datacentre IBGP peerings - https://phabricator.wikimedia.org/T433675 (10cmooney) 03NEW p:05Triage→03Low [10:22:45] elukey: just so I can check I've understood correctly: ms-be2082 currently has no useful IPMI user, so any cookbooks that rely on that (like reimaging) won't work. And this is true of all the Supermicro Config-J systems? [10:22:46] Emperor: correct yes, but for some reason that I don't understand both ADMIN (SM pre-existing user) and wmfroot don't work with IPMI [10:23:35] Emperor: no sorry the correct was for the sentence before :D ipmi is not used for reimage if you use UEFI, and I added a provision patch to skip a failure if ipmi doesn't work at the end [10:23:50] so in your case, in theory with UEFI you should be able to reimage without issues [10:23:57] IIRC those are all UEFI right? [10:24:40] We definitely still have some BIOS-booting Config-Js. I don't _think_ any of them are supermicro, but that's not entirely straightforward for me to check. LMS... [10:25:37] elukey: so, err, it would be worthwhile my checking I can reimage ms-be2082 now, then? [10:25:45] Emperor: we have a cumin alias for uefi-enabled hosts, so if you know the hostnames we are good [10:25:52] Emperor: I think so yes [10:26:08] elukey: will do; there's a cumin alias for swift backends. [10:26:40] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:29:03] elukey: 34 uefi backends, 38 bios ones [10:30:40] 17 SM backends (per A:swift-be and P{F:boardmanufacturer=Supermicro} ) [10:31:41] 54 Dell backends (per A:swift-be and P{F:boardmanufacturer='Dell Inc.'} ) [10:32:24] relevantly, all 17 SM backends are uefi booted (A:swift-be and A:uefi-boot and P{F:boardmanufacturer=Supermicro} ) [10:39:48] okok the dells should all be fine, but I'll re-run the cookbook to rollout the wmfroot later on since it now has an extra check for ipmi [10:40:01] so I'll confirm that only SMs are problematic, for some reason that I don't know [10:40:28] fun fact - sretest2010, that is the new SM config j, doesn't have problems [10:40:39] Thanks - ms-be2082 looks to be reimaging OK, I'll confirm here when it's done [10:40:58] perfect, I am going afk for lunch but I'll read later! [10:41:00] elukey: I need to get to that (I made puppet run OK, I think next is get dc-ops to swap/wipe some disks) [11:02:52] elukey: ms-be2082 reimaged OK 👍 [11:02:56] * Emperor -> lunch [11:56:25] FIRING: [2x] SystemdUnitFailed: gitlab-package-puller.service on apt-staging2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:16:52] 10netops, 06Infrastructure-Foundations, 06SRE: JunOS: Investigate BGP PIC Edge / Protection - https://phabricator.wikimedia.org/T432381#12174142 (10cmooney) >>! In T432381#12131427, @ayounsi wrote: > Another optimization to consider: https://www.juniper.net/documentation/us/en/software/junos/bgp/topics/topic... [12:23:13] Emperor: \o/ [13:19:22] Emperor: I left a note for sretest2010 in https://phabricator.wikimedia.org/T394357#12174444, in case you need to reimage and I am not around [13:23:56] 10SRE-tools, 10Spicerack: Clarify when spicerack/cumin is going to retry - https://phabricator.wikimedia.org/T433698 (10fgiunchedi) 03NEW [13:23:59] 10SRE-tools, 06Infrastructure-Foundations, 10Spicerack: Clarify when spicerack/cumin is going to retry - https://phabricator.wikimedia.org/T433698#12174470 (10fgiunchedi) p:05Triage→03Low [13:31:13] elukey: ta, thanks [13:40:38] 10netops, 10Cloud-VPS, 06collaboration-services, 06Data-Persistence, and 7 others: codfw: rack B5 maintenance - https://phabricator.wikimedia.org/T430918#12174538 (10brouberol) [13:40:55] 10netops, 06cloud-services-team, 06Data-Persistence, 06Infrastructure-Foundations, and 5 others: codfw: rack B7 maintenance - Tuesday July 21st 14:00 UTC - https://phabricator.wikimedia.org/T430928#12174539 (10brouberol) [15:32:39] 10SRE-tools, 10Cumin, 06DC-Ops, 06Infrastructure-Foundations, and 2 others: add dcops group to run sre.hosts.downtime cookbook - https://phabricator.wikimedia.org/T433409#12174984 (10Dzahn) @elukey ACK, after making that comment I realized there is a difference between "add a member to the group"-appro... [15:56:40] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:55:11] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12175477 (10Papaul) [19:56:40] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:26:25] RESOLVED: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:27:55] FIRING: SystemdUnitFailed: generate_vrts_aliases.service on mx-in2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed