[05:54:05] federico3: when decommissioning a host I got: https://phabricator.wikimedia.org/P96464 [05:54:52] I've deleted it manually from zarcillo but can you take a look? [05:54:58] ok [05:55:03] thanks [05:55:36] maybe were you not logged in? [05:55:51] in zarcillo? [05:56:05] yes in the web ui [05:56:13] I am not, but that's needed to run the decommissioning cookbook? :-/ [05:56:30] I think that should either be specificied when starting the run or fixed (if that's fixable) [05:56:34] ah it was through the cookbook, ok [05:56:38] yeah [05:56:49] That's the output of the cookbook [05:57:04] It is in the paste, at the end: END (FAIL) - Cookbook sre.mysql.decommission (exit_code=99) [05:57:28] ah could it be perhaps a new cumin host/ipaddr [05:57:51] ah that could be, because it was from cumin1004 [05:57:58] where is that handled? [05:59:09] yep https://gitlab.wikimedia.org/repos/data_persistence/zarcillo/-/blob/main/zarcillo/webapp/routers/auth.py?ref_type=heads#L30 here - I'm adding 1004 now [05:59:57] Is there way to avoid hard-coding IPs there? [06:00:07] if we could have a trusted list from some API ideally [06:00:09] btw, you can remove cumin1003 entirely [06:00:13] yep [06:00:17] it is going to be decommissioned very soon [06:01:07] yes, I'm replacing it with 1004, I spoke with Moritz [06:01:13] thanks [06:03:20] maybe the list of ipaddrs could be generated by puppet? but puppet cannot configure k8s containers... [07:56:49] moritzm: all the database backups have run fine from cumin1004 [07:59:08] that's how I like DB backups! [13:54:55] dhinus: how do you feel about starting to decommission clouddb hosts after the DC switch? [13:55:03] They've been depooled for a month now I think? [13:55:21] I can do it slowly [13:55:52] yes I think we can proceed with the decom [13:56:37] the new ones seem to be working fine... the only issue I saw is the "OOM" alert that triggered yesterday, and did not trigger before (at least not in a long time) [13:56:51] dhinus: which host? [13:56:57] clouddb1023 (it did not actually OOM but it went to 95% usage) [13:57:14] I put some graphs in T438200 [13:57:15] T438200: clouddb1023 memory alert - https://phabricator.wikimedia.org/T438200 [13:58:11] I am not too surprised that happens on clouddb hosts [14:01:15] dhinus: did anyone stopped clouddb1023? [14:02:15] yes andrewbogott yesterday debugging the alert, and he didn't know that he had to restart replication manually, I told him later [14:02:40] also not too surprised it happens, but surprised that specific alert never triggered in the past year IIRC [14:02:51] ok just to understand the graphs [14:02:54] I am checking a few things [14:03:48] thx [14:23:19] dhinus: I've found nothing too relevant to be honest, I wanted to discard OS stuff, which I have sort of done by checking node_memory_Slab_bytes and node_memory_SReclaimable_bytes metrics in grafana [14:23:58] dhinus: I think it was just a heavy query, we can monitor this host and see how it goes after the restart, but this is not uncommon on clouddb* hosts given their very different workload from prod [14:24:03] Do you want me to comment on the task? [14:26:24] marostegui: yes please, and I agree if it does not re-occur we can ignore it for now, and continue with the decomms. [14:26:46] ok I will comment there [14:27:11] thx! [15:43:15] Hi folks, could I get a +1 to https://gerrit.wikimedia.org/r/c/operations/puppet/+/1342718 please? Prep for the next lot of ms-be2* nodes [15:46:31] Done [15:47:06] Thanks :)