[01:09:44] hello data-persistence - friendly reminder about the etcd primary switchover coming up today (Wednesday) at 13:30 UTC. that will involve ~ 15m of read-only time, during which dbctl updates will not be possible. thanksĀ in advance for your patience! [01:11:46] oh, and after the switchover is complete, I'll follow up here about a Zarcillo restart. that won't need urgent action, but would be good to do soon so we don't lose track of it. [02:21:48] FIRING: [4x] MysqlReplicationLagPtHeartbeat: MySQL instance db2225:9104 has too large replication lag (11m 0s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [02:25:25] FIRING: SystemdUnitFailed: pt-heartbeat-wikimedia.service on db2207:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:26:48] FIRING: [10x] MysqlReplicationLagPtHeartbeat: MySQL instance db2175:9104 has too large replication lag (15m 15s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [02:36:48] FIRING: [10x] MysqlReplicationLagPtHeartbeat: MySQL instance db2175:9104 has too large replication lag (18m 30s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [07:14:35] marostegui: I haven't depooled any clouddb instances, unless a switchover or rolling restart did them via detection - when was the log entry that you found? [07:14:52] cezmunsta: no, it was related to the puppet disablement on db2230 [07:15:02] it is all solved now, no need to worry :) [07:36:33] Hmm, still puzzled about that, but &>/dev/null :) [08:15:25] FIRING: SystemdUnitFailed: swift-object.service on ms-be1065:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:02:36] ^-- node being converted to new-style storage [10:01:10] db2230 has needrestart installed, which we usually don't use (since debdeploy has it these features built-in): https://debmonitor.wikimedia.org/packages/needrestart I suppose this was for some debugging, I would go ahead and remove it? [10:07:46] moritzm: I think that is fine, thanks [10:11:19] ok, done [10:16:08] Hi folks, could I get a +1 to https://gerrit.wikimedia.org/r/c/operations/puppet/+/1327070 please? I've converted the last 3 old-style storage nodes to new-style storage, so they're ready to go back into the rings [10:18:54] Emperor: done https://en.wikipedia.org/wiki/Get_in_the_Ring :) [10:26:22] XD [11:10:25] FIRING: SystemdUnitFailed: mariadb.service on db2902:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:12:02] federico3: ^ [11:12:18] yep, It's deploying [11:37:04] FIRING: MysqlPredictiveFreeDiskSpace: Host db2902:9100 predictive low disk space on root - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [11:57:04] RESOLVED: MysqlPredictiveFreeDiskSpace: Host db2902:9100 predictive low disk space on root - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting - https://grafana.wikimedia.org/goto/Jdz2PnLNg?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DMysqlPredictiveFreeDiskSpace [12:39:41] heh :D [13:09:12] s4 selects on master went from 700 per sec to 500 https://grafana-rw.wikimedia.org/d/000000378/ladsgroup-test?orgId=1&from=now-3h&to=now&timezone=utc&viewPanel=panel-33 [13:09:25] nice thanks! [13:11:31] this is the last large db work I'm going to do :') [13:48:46] hello data-persistence - as promised, I'm back to request a zarcillo restart whenever is convenient for you. no disruptive work planned in codfw etcd, so this is not urgent, but it would be good to do before we lose track of it :) cc: federico3 [13:49:05] federico3: ^ [13:52:48] yup, restarting in 1 sec [13:56:59] awesome, thank you very much [13:57:24] federico3: is the zarcillo restart process documented somewhere? [14:06:26] yes but there's an open task to improve the doc and scripts [14:19:35] Emperor: o/ is it ok if I reimage sretest2010 for a test? [14:22:47] elukey: please go ahead [14:49:36] Emperor: as FYI I merged the new reimage code to allow sretest2010 to be reinstalled without test-cookbook and/or spicerack checkouts tricks [14:49:39] seems to work fine [15:04:05] cool [19:39:58] hello. is it okay if i do some updates to db2207 in https://phabricator.wikimedia.org/T435271 ? will require reboots. [19:40:52] JennH: let me just check the server for you [19:45:00] I have extended the downtime and I am just stopping MariaDB for you [19:48:43] JennH: all ready for you [19:49:43] okay cool. i shouldn't need it for long. I'll let you know once i get this done [19:51:13] Thanks! [20:44:44] cezmunsta: finished! you can have it back. thank you very much! [20:46:24] JennH: thanks!