[11:51:16] alas, I think my pontoon stack has bitrotted :( [13:07:28] heh, trying again in the right channel this time: [13:07:33] I'm seeing several alerts about systemd unit failures on clouddb1024 and clouddb1025: wmf_auto_restart_prometheus-mysqld-exporter@s4.service and wmf-pt-kill@s4.service -- I assume that's fallout from the x4 change, should I just go through and remove those units by hand? [13:25:25] andrewbogott: that (fallout) would indeed appear to be the case. I didn't spot anything in Puppet to remove stale ones, so I would guess so. marostegui: is that correct? [13:27:20] I'll removing things and see what happens :) [13:35:01] andrewbogott: yeah those can be removed. I can do it later if you want, I'm busy at the moment. [13:35:18] np I'll do it [13:37:31] Thank you! [13:41:28] hm, puppet replaced it [13:50:22] andrewbogott: which host had it replaced? [13:50:59] both clouddb1024 and clouddb1025. But you don't need to dive in unless you feel like it, I'm just thinking out loud. [13:53:37] well, wait... what /should/ be on those servers? Hiera shows only x4, which means those failing alerts need fixing, not removing [13:59:56] cezmunsta: I'm withdrawing my offer to fix this myself because the problem is not what I thought it was. Shall I open a task instead? [14:02:09] Don't the alerts just need to be deleted? i.e. they are resurfacing - I didn't see anything in the Puppet changed output for s4, only x4. I also don't see s4 as a service on there [14:02:25] i.e. "sudo cumin clouddb1024\* 'systemctl list-units prometheus-mysqld-exporter*'" shows only x4 [14:02:58] The alerts look to be from last week, is that right? [14:03:43] You can of course create a ticket if that is not the case [14:03:58] oh you're right I'm confusing s3 and x4. [14:04:21] > codfw snapshot very_wrong_size 13 hours ago 1.2 TB -45.2 % The previous backup had a size of 2.1 TB, a change larger than 15.0%. [14:04:28] so... I will have another go and try not to confise those :) [14:05:00] gonna drop the tables in a couple more replcias [14:25:45] [x] s4 alerts fixed [14:26:18] *\o/* thanks! [15:07:55] and now... clouddb1023 is alerting with "WARN Memory 96% used" [15:09:38] cezmunsta, any idea what to do about that? [15:23:38] Sorry, no. I think that is WCS' realm, but could be wrong. That said. Grafana doesn't seem to show the same usage. marostegui: any ideas? [15:26:04] anybody has a good understanding of mediabackups.backups ? [15:26:30] federico3: why? [15:26:57] Have you sent the MR for the DB backups yet for review? [15:26:59] cezmunsta: https://phabricator.wikimedia.org/P96463 [15:27:15] no, I'm investigating what queries to add hence the question [15:27:52] I thought that I asked for the MR and to leave those for now? [15:32:30] https://phabricator.wikimedia.org/T438003#12327417 [15:34:43] I want to understand what is generated by backups to avoid duplication of effort later on [15:35:04] andrewbogott: that's fairly normal on cloudb hosts due to the nature of some of those massive queries they get. It's not a big deal but if you want to clear the warning just restart the process [15:35:19] federico3: it is still WIP hence the comment on the ticket [15:36:27] Please send the MR as requested and I can then comment