[01:55:03] FIRING: PuppetFailure: Puppet has failed on ms-be1090:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [05:55:03] FIRING: PuppetFailure: Puppet has failed on ms-be1090:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [09:55:03] FIRING: PuppetFailure: Puppet has failed on ms-be1090:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [10:17:13] federico3: do you happen to know much about the ms-be servers? Puppet looks to be failing due to I/O error on one of the XFS volumes [10:18:51] "Metadata corruption detected at xfs_dinode_verify.part.0+0x46e/0xa80 [xfs], inode 0x301cde42c dinode" [11:03:38] ms-be are swift. Emperor are you around? [11:04:26] No, OoO today [11:44:57] federico3: I have added some updates to the essential work log, do you have anything to add to it? [11:46:19] the work for the testbed, but does it actually qualify as essential? [11:46:54] Yes [11:46:54] probably not [11:46:56] ok [12:59:46] federico3: db2902 keeps firing the predictive disk space alert. Is there a reason that the new servers aren't equal in their disk usage? I see variation from 68-76% usage [13:03:27] Also, db2230 is showing as not using GTID [13:07:01] I don't see anything in Server Admin Log, but have you been doing anything with the test-s4 cluster recently? [13:55:03] FIRING: PuppetFailure: Puppet has failed on ms-be1090:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:18:19] * Emperor not here today, but: ms-be1090 has failed a disk, feel free to silence the alert until Monday, I'll deal with it then [14:53:43] cezmunsta: they seems to be keeping the binlogs for a good while but I'm not sure why they grow so fast [14:58:32] What do you mean "keeping the binlogs for a good while"? There are only 2 and they are set to expire after 30 days [14:59:17] Also, what about db2230 and its lack of GTID? Any ideas why that has happened? [15:01:03] Have you been running anything against the cluster? [15:01:08] no [15:01:40] db2230 was only stopped/started by the cloning script I think, maybe gtid needs to be enabled back by hand? [15:02:22] marostegui: if you get a moment today can you review, https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329665 or if someone else on the team would like to [15:05:13] I can delete some stale files