[09:33:00] hnowlan: https://alerts.wikimedia.org/?q=alertname%3DMediaWikiCronJobFailed&q=team%3Ddata-persistence&q=%40receiver%3Ddata-persistence-task I don't think we own that alert or we certainly have no actionables for it. I am not even sure who created it :) [09:34:55] marostegui: serviceops did that one, when migrating the old mw cronjobs to k8s a team was required as a param. cezmunsta asked about it last week also, some context there [09:35:38] usually I think the job failing would be of interest to yall, but your call - I'd check in with serviceops to see if they'd know of a better owner [09:35:59] for now I will delete the failing job (which failed because of etcd, nothing to do with the DBs) which should resolve the alert [09:37:09] hnowlan: I think I've discussed those jobs with effie before and most of the time the failure was on different layers than the DBs, so we (dba) couldn't do much about it [09:37:13] hnowlan: also I just merged: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1343590 [09:37:25] federico3 cezmunsta ^ if you see something weird in orch let me know [09:37:40] ok [09:37:57] marostegui: thanks! [09:38:21] deleted the job, the alert should resolve soon [09:38:27] thank you hnowlan [09:38:51] I imagine there is probably a mediawiki-focused team who would be better owners of the job [09:39:00] yeah I think so [09:39:08] I don't think serviceops can do much about it either [09:41:42] might be worth asking (or asking serviceops to ask) on slack to see if someone would take it - the git blame on the actual maintenance script doesn't give any clear indications of who to ask, it's very old and is gets 2-3 small changes a year [10:01:02] I'm rolling forward Zarcillo again for a quick test [13:55:54] the split version of Zarcillo is deployed and the metrics are working [14:20:42] +1 [16:00:48] FIRING: [5x] MysqlReplicationLagPtHeartbeat: MySQL instance db1901:9104 has too large replication lag (11m 3s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [16:01:58] cezmunsta: related to your on-´going work on test-s4? ^ [16:05:48] FIRING: [6x] MysqlReplicationLagPtHeartbeat: MySQL instance db1901:9104 has too large replication lag (15m 3s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [16:10:48] FIRING: [6x] MysqlReplicationLagPtHeartbeat: MySQL instance db1901:9104 has too large replication lag (14m 39s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [16:15:48] RESOLVED: [6x] MysqlReplicationLagPtHeartbeat: MySQL instance db1901:9104 has too large replication lag (14m 39s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [16:42:04] FIRING: [2x] MysqlPredictiveFreeDiskSpace: Host db2901:9100 predictive low disk space on root - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [16:47:04] FIRING: [3x] MysqlPredictiveFreeDiskSpace: Host db2901:9100 predictive low disk space on root - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [17:07:04] RESOLVED: [3x] MysqlPredictiveFreeDiskSpace: Host db2901:9100 predictive low disk space on root - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace