[03:23:03] FIRING: PuppetFailure: Puppet has failed on ms-be2089:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [07:06:17] marostegui: morning, FYI that I set weight of db1261 (x4 host based on topology) to zero on s4 so I can capture queries and make sure nothing is queying tables it shouldn't. If you want to, revert it or completely remove it from s4 (and other x4 dbs) [07:06:56] Amir1: we already have a host with 0, which is db1263 [07:07:02] why don't you use that one? [07:07:21] or you want just one in s4? [07:08:05] Amir1: keep in mind I will remove the hosts in x4 from s4 [07:08:16] in fact, I am going to do it now, to simplify things [07:08:42] I assumed the db1263 will be the master [07:08:55] yes it will be [07:09:09] ok removing x4 hosts from serving s4? [07:09:17] Yeah. That's what I did for db1261 (weight zero though. Not removal) [07:09:53] Yeah. The other way around too. If we have hosts in x4 it shouldn't, let's remove it too [07:11:02] Amir1: so db1261 would also need to be removed from s4 [07:11:17] That should give a good view of how much traffic it's going to have [07:11:34] Amir1: what I mean is db1261 is in x4 and should not be in s4 [07:11:52] marostegui: yup. I didn't do it cause I didn't want to interfere too much with your stuff but yeah it should be removed [07:12:03] ok I will do it [07:12:58] Thanks <3 [07:23:03] FIRING: PuppetFailure: Puppet has failed on ms-be2089:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [07:34:35] ^-- objects3/sdf1 - unmounting and remounting the fs (which is what xfs_repair suggested) went OK, so I'll see how it goes. [08:07:48] RESOLVED: PuppetFailure: Puppet has failed on ms-be2089:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [08:19:50] ladsgroup@cumin1003:~$ cat x4_selects | awk -F"FROM " '{print $2}' | awk -F'WHERE' '{print $1}' | python3 -c "import re,sys; print(set(re.findall(r'\`([^\`]*)\`', sys.stdin.read())))" [08:19:50] {'imagelinks', 'langlinks', 'iwlinks', 'linktarget', 'pagelinks', 'categorylinks', 'existencelinks', 'templatelinks', 'externallinks', 'collation', 'page'} [08:19:50] ladsgroup@cumin1003:~$ wc -l x4_selects [08:19:50] 1465 x4_selects [08:20:09] captured a bit of queries, none try to read from tables that they shouldn't [08:20:40] I should try it the other way around too [08:28:57] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2196:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:29:11] s4 also is not showing any queries to links tables or collation [08:30:02] captured a couple thousand queries [08:33:37] Sorry Amir1 I was in a meeting [08:33:47] Thank you [08:33:52] We are going to start the split in 30 mins [08:34:18] oooh scary [08:55:19] zabe: let me know when you are around [08:56:19] marostegui: I am here:) [08:56:29] zabe: cool I will start in 4 minutes with the RO then! [09:00:42] zabe: starting in 1 min, I am double checking pt-heartbeat is running well on the x4 masters [09:00:59] sounds good [09:01:08] All good, so going for RO [09:01:14] I will !log stuff in the operations [09:01:19] channel, that is [09:53:48] FIRING: PuppetFailure: Puppet has failed on ms-be2089:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [10:20:49] sigh, looks like that disk is dead [11:00:45] marostegui Amir1 can you please review this update for wikireplicas users? https://etherpad.wikimedia.org/p/wikireplicas-x4 [11:02:15] if you expect the x4 replicas will take more than "several days" I'm happy to adjust the wording :) [11:02:50] also cc zabe ^ [11:03:27] dhinus: I think several days is fine, it won't take weeks no [11:15:36] Amir1 zabe does modules/mediawiki/files/mariadb/tables-catalog.yaml need patching? [11:28:02] sorry I had to go to a meeting [11:28:45] marostegui: that does need a patch but for documentation reasons, nothing is needed for the views or replication filters [11:29:06] Amir1: excellent thanks! [12:03:46] urandom: I'm doing firmware updates on aqs2011 and will then reimage it and subsequently aqs2012, which will be all the RAID aqs nodes done to bookworm by my end-of-day. The JBOD ones await post-reimage instructions :) [12:30:00] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter.service on db2196:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:31:43] ^i silence this, this is the host that crashed [13:07:33] Emperor: thanks! I’ll get the jbod worked out [13:07:57] urandom: [aqs2012 reimage just started] [13:54:03] FIRING: PuppetFailure: Puppet has failed on ms-be2089:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:13:32] I've opened T437293 for the failed disk and silenced that alert for 3 days [14:13:33] T437293: Disk (sdf) failed in ms-be2089 - https://phabricator.wikimedia.org/T437293