[03:15:01] FIRING: [3x] DiskSpace: Disk space backup1010:9100:/srv/objectstorage 2.069% free - https://wikitech.wikimedia.org/wiki/Monitoring/Disk_space - https://alerts.wikimedia.org/?q=alertname%3DDiskSpace [06:00:00] https://phabricator.wikimedia.org/T437833 created to track ^ [06:00:04] I am checking [07:03:16] dhinus: can you create views on clouddb1024:3364 (x4)? [08:18:42] that's going to need https://gerrit.wikimedia.org/r/c/operations/puppet/+/1341041 in first [08:26:42] marostegui: views created, seemingly including a lot of testcommonswiki tables that shouldn't be on x4? [08:27:25] taavi: Since dhinus asked, I just moved the testcommonswiki to x4 as well [08:27:32] to make things simpler [08:27:49] https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1338142 [08:29:00] Amir1: that patch is about links tables and I don't think, say, translate_metadata is a links table? :P [08:29:25] I did clean it up IIRC [08:29:43] Amir1: clouddb1024 for testocmmonswiki has lots of non link tables [08:29:45] at least on db1261, maybe not on 1260 [08:29:48] is that expected? [08:30:06] taavi: can you test if things looks ok on your side and pool the host? if that works, we'd have to do the same in clouddb1025:3364 [08:30:52] nope, we should clean it up. I can take care of it. I came actually here to ask whether it's okay to drop the tables on a couple more replicas [08:31:00] (both s4 and x4) [08:33:56] Amir1: yes please, go for it [08:34:27] https://giphy.com/gifs/muppets-chainsaw-beaker-0eVM7GVxTDDKxn7OyX [08:34:30] Amir1: also if you can check clouddb1024 and clouddb1025 for testcommonswiki and clean what is needed there, that'd help a lot [08:34:34] haha [08:34:37] definitely [08:35:43] Amir1: which replica is fully clean now? I need to copyu things to dbstore1007 so I can benefit from the one that is fully clean now [08:36:07] I think db1261 shoudl be, there should be only 13 tables in it [08:36:10] ok! [08:36:15] in testcommons too? [08:36:35] Can't say 100%‌sure but I think so too [08:36:41] let me check [08:36:57] mmm [08:37:00] it has also user table [08:37:07] I think it needs to be cleaned, it doesn't look clean for testcommons [08:37:24] could you do that so I can copy a fully clean copy to dbstore1007? [08:39:15] sure [08:40:50] marostegui: the views on clouddb1024 seem ok except the extra tables which I assume can be cleaned live so I'll pool it, one moment [08:45:28] taavi: thanks [08:45:32] Amir1: let me know when done, thanks [08:54:21] marostegui: db1261 should be clean now on both testcommonswiki and commonswiki [08:54:51] thanks Amir1 [08:55:31] marostegui: turns out I forgot about creating user accounts on the replicas, that script is now running but might take a while [08:57:07] taavi: ok thanks! [08:57:26] taavi: I think I saw some accounts there though [08:57:50] yeah, I saw some earlier today [08:58:58] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db1199:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:59:57] that is probably from past week when I was working on that host ^ i will fix it [09:00:13] huh [09:00:13] taavi: on db1260, testcommonswiki is clean, which hosts exactly need clean up? [09:01:01] marostegui: https://phabricator.wikimedia.org/P96437 are we missing the role definition or something then? [09:01:24] let me check [09:01:28] Amir1: clouddb1024/5 [09:01:47] thanks. Will clean it up after meeting. I probably do it from sanitarium master [09:02:22] taavi: please retry [09:02:50] works now, thanks [09:02:54] I will pool it for real then [09:02:54] thanks [09:02:57] added it also to 1025 [09:03:03] ty! [09:03:10] :* [09:03:58] taavi: there will be abit of lag on x4, as I need to stop db1260 (sanitarium master) for a bit of maintenance [09:04:15] ok! [09:04:53] the new x4 endpoints are now pointing to clouddb1024, let me know when I can create views on and then pool 1025 [09:05:17] taavi: if 1024 works fine, you can go ahead any time [09:07:10] cool, done [09:07:18] thank you [09:07:33] both hosts pooled? [09:07:35] yes [09:07:43] mind commenting on https://phabricator.wikimedia.org/T409557 for the record? [09:08:00] sure [09:08:58] RESOLVED: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db1199:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:09:03] thank you taavi [09:09:30] likewise [09:46:13] https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&refresh=5m&var-server=db2172&var-datasource=000000026&var-cluster=mysql&viewPanel=panel-28&from=now-1h&to=now&timezone=utc [10:35:47] taavi: ok to depool 1025:x4 so I can clone an-redacteddb? [10:36:01] I can depool myself [10:36:16] marostegui: one is fine, just not both at the same time please [10:36:25] yeah, no, just one [11:23:58] FIRING: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:53:58] RESOLVED: SystemdUnitFailed: swift_rclone_sync.service on ms-be1069:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:57:25] taavi: clouddb1025 is back up - I will repool once it is caught up with its master [12:37:00] taavi: host repooled