[02:00:40] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db2201:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:00:40] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db2201:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:15:14] can someone give 1 month of silence to those ^ [07:15:19] that host is broken [07:19:22] marostegui: silenced for 28 days [07:19:33] thank you Emperor [10:51:37] marostegui: just noticed clouddb1025 is still serving s4/s6, but the plan in T409557 was to have x4/x1 in pair with clouddb1024 [10:51:38] T409557: Productionize new clouddb* hosts (clouddb1022-1033) - https://phabricator.wikimedia.org/T409557 [10:53:01] dhinus: yes, but x1 isn't in wikireplicas and x4 is still not a reality [10:54:20] is x4 already visible to users as an alias? [10:54:34] is it? [10:54:38] not sure, checking :) [10:54:51] dhinus: In terms of wikireplicas, they are still called s4 [10:54:53] I'm asking because when rebooting clouddb1024 we got an alert that x4 was not available to haproxy [10:55:11] maybe there's something there, but internally it is all s4 still [10:55:13] but if no users are connecting to x4 that's fine [10:55:14] including sanitarium etc [10:56:04] if users start using 'x4' as the DNS, we should have two copies pooled in haproxy [10:57:04] sure, but underneath it is all s4, there's not sanitarium with x4 yet [10:57:06] I don't think there was any public announcement yet, so that's fine [10:57:25] yep, just a pointer to s4 [10:58:48] do you know why s4 is not pooled in clouddb1032? [10:59:08] it should be [10:59:15] if it is not, it should be :) [11:00:01] it was pooled here https://sal.toolforge.org/log/cFFDop8B8tZ8Ohr0dhaj [11:00:14] but currently confctl says {"clouddb1032.eqiad.wmnet": {"weight": 0, "pooled": "inactive"}, "tags": "dc=eqiad,cluster=wikireplica-db-analytics,service=s4"} [11:00:56] I will try to debug why [11:01:42] Yeah I remember pooling all of them [12:47:13] Hi folks, could I get a +1 to https://gerrit.wikimedia.org/r/c/operations/puppet/+/1326826 please? The last 3 old-style storage nodes are drained, so this removes them from the rings and sets them up to be converted to new-style storage. [12:48:33] checking [12:50:08] TY :) [13:15:30] I couldn't find what caused clouddb1032 to be depooled, I repoooled it: https://phabricator.wikimedia.org/T409557#12226198 [13:16:23] I also created https://gerrit.wikimedia.org/r/c/operations/puppet/+/1326830 to sync the config of clouddb1025 with clouddb1024 [13:22:46] FYI, db2230 had Puppet disabled since July 28 and is now evicted from puppetdb, known issue? [13:47:15] not sure, I'll investigate [15:08:25] federico3: did you find why it was disabled? [15:08:47] And, can we get it back to puppet? :) [15:08:53] no, I can look in a bit [15:09:00] Ok! [15:09:08] Thanks [15:09:11] probably a script that disable puppet during the run that crashed out [15:13:42] I've reenabled it [15:17:07] i haven't seen a cb run in SAL [15:17:35] I think it was disabled by ceri based on the logs [15:17:39] Anyway, it's been reenabled now