[05:18:57] FIRING: SystemdUnitFailed: check-private-data.service on db1270:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:24:06] ^-- expected? [07:36:06] yes [07:36:13] I will work on it later, noting urgent [07:36:24] I am buried into things, I will get to it later [07:53:02] 👍 [07:57:43] marostegui: I will take a quick look [07:58:08] it is a new sanitarium host, so not yet in production, but thank you :)) [08:01:42] It looks like it is the x4 part that failed attempting to connect to the DB [08:02:46] I will dig later, it is not urgent, it is not being used at the moment [08:03:55] I'll leave it for you then. All of the others ran OK, just the x4 one that failed. For reference, all runs of the timer to date had been successful [08:04:15] yeah, I added x4 yesterday, so may be grants or something [08:04:34] ah I know what it is hehe [08:05:37] x4 was added to db1155 (the original sanitarium) and later it will be transfered to db1270 as part of https://phabricator.wikimedia.org/T436920 [08:05:55] so I will just silence that alert for now [08:08:54] +1 [09:52:27] dhinus: I am going to depool clouddb1024 from s4, as I am converting that one to x4 [09:57:56] marostegui: ack! [10:27:11] Amir1: why would x4 have this on the binlog: #260909 10:07:14 server id 171970717 end_log_pos 511316513 CRC32 0x8a90e755 Update_rows: table id 14341 flags: STMT_END_F [10:27:11] ### UPDATE `testcommonswiki`.`user` [10:27:23] That's from yesterday, why do we have updates arriving to that table [10:27:42] that is on sanitarium, which replicates from its master in x4 [10:27:44] testcommonswiki wasn't migrated until late last night [10:27:52] :_( [10:28:04] node_disk_written_bytes_total [10:28:09] wrong paste [10:28:12] https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1338142 [10:28:30] ok thanks [10:28:46] it's testcommonswiki, it doesn't really have anything honestly, closed for long time [10:29:01] that user update it probably me even when I was testing things [10:29:31] right yeah ok, that explains it, thanks [11:59:15] https://phabricator.wikimedia.org/T437581 Amir1 will this ticket be closed? I'd expect the bot to create it [11:59:24] Or will it be reused? [11:59:42] I run it, let me see [11:59:50] yeah, but do't create the wiki [11:59:57] this needs us to add the filters and restart sanitarium [12:00:00] yup, private wikis are fun [12:00:07] I need to check ACL for swift too [12:00:14] I will try to get it sorted today but no promises [12:00:21] nah, don't rush [12:00:33] Amir1: ok, what's the ETA for creating this? [12:00:44] one week at least [12:00:55] ok I will add a note formyself for monday [14:17:18] Amir1: I thought https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329558?usp=dashboard was waiting until at least observability had OK'd it and Cormac was back and I'd been able to finish reviewing it? [14:22:13] Amir1: relatedly, it it urgent enough it can't be reverted? [14:22:24] I don't think so. This is required for our OKR which we will have to deliver by end of this month and Cormac is out for two months [14:22:30] it is urgent [14:23:03] if you want to, we can limit it to one of the 16 groups, to reduce the metrics but talking to Chris, he said there are mitigations [14:23:47] I would rather wait to see if it causes issues and take action if it does [14:24:09] That's not what I thought I'd agreed with Cormac. [14:25:24] And I was part-way through writing a review asking for a simpler approach and probably a more restricted set of containers, because I thought I had time to review before it got deployed because of where I'd left my discussion with Cormac. [14:28:34] we can look at the cardinality right now, can't we? [14:29:44] is it causing issues? I gave Hugh heads up. If something becomes a problem, I can take a look [14:34:54] swift_per_container_stats_bytes_total has cardinality 512 in thanos [14:35:12] same swift_per_container_stats_objects_total [14:36:48] FIRING: PuppetFailure: Puppet has failed on ms-backup1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:37:50] Hugh said it's okay, just keep an eye on it for a day and ping Keith tomorrow to double check [14:40:03] ^-- I'm looking at the puppet failure [14:41:33] Emperor: lmk what you find :) [14:41:48] FIRING: [2x] PuppetFailure: Puppet has failed on backupmon1001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:46:48] FIRING: [3x] PuppetFailure: Puppet has failed on backupmon1001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:51:48] FIRING: [4x] PuppetFailure: Puppet has failed on backupmon1001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [15:01:48] FIRING: [5x] PuppetFailure: Puppet has failed on backupmon1001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [15:09:27] ^ checking [15:31:17] dhinus: clouddb1024 and clouddb1025 are now replicating x4, but nothing should be sent to 3364 yet (it doesn't have the views yet anyway) [15:32:40] marostegui: great, thanks! [15:33:19] I will resume the work there tomorrow [15:35:35] btw ms-backup have also puppet failing due to the account removal [15:35:40] cezmunsta Emperor [15:35:56] marostegui: yes [15:36:07] rewording the task [19:02:03] FIRING: [4x] PuppetFailure: Puppet has failed on backupmon1001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [23:02:03] FIRING: [4x] PuppetFailure: Puppet has failed on backupmon1001:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure