[08:20:06] @marostegui @cezmunsta I moved zarcillo to Trixie, please let me know if you see anything odd [08:20:14] ok thanks [11:25:34] federico3: any reason that db1176 is read-only? [11:26:11] uh, no [11:26:29] OK, I will fix that and see how much further the cb goes [11:54:39] marostegui: re "Like why it's never happened before" .. perhaps now that cumin hosts are on tixie could be another explanation, i.e. newer versions [11:55:48] so, any speedup could have meant a greater chance of seeing the connect/prepare states [11:56:46] yeah maybe [12:34:39] federico3: what did you see yesterday that made you think that something might have been cached during the cb run? [12:37:24] the cb seemed stateful in few places, let me review it [12:38:36] one of the suspicious bits was seeing that the expected position was *earlier* that the current actual position [12:39:26] ack [12:44:19] federico3: but that could be explained by the cookbook maybe getting stalled in between checks [12:45:25] FIRING: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db2250:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:45:34] fixing ^ [12:45:40] leftover from the clean up [12:47:16] but reviewing the cb it seems most functions wrapped by confirm_on_failure are idempotent so doing a retry should be ok, however in run() there's master_to_position: Any = confirm_on_failure(self.wait_master_to_position) that captures the position of the master and it's later on compared by enable_circular_replication and using self._validate_slave_status ... I wonder if the position should be compared using a disequality [12:50:24] RESOLVED: SystemdUnitFailed: wmf_auto_restart_prometheus-mysqld-exporter@s5.service on db2250:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:55:47] marostegui: "but that could be explained by the cookbook maybe getting stalled in between checks" I think so, also my MR passes in a callable instead of the result, so that shouldn't be an issue even it it was before /me crosses-fingers [12:56:15] let's see! [13:05:54] "**MASTER_FROM db1176.eqiad.wmnet should be read only**" <- is that expected given that the prepare completed? [13:06:35] that's the finalize? [13:06:42] Yes [13:06:54] then yes, because it assumes that eqiad is no longer the primary DC [13:07:16] so codfw master should be RW and eqiad already RO (which is done via the DC switchover cookbooks) [13:07:59] OK, I will do that manually for now and then repeat to make sure that it was just due to the final step that I skipped in the first run [13:10:12] yeah [13:27:29] So, I ran it c -> e and I still see the same, db2230 rw, db1176 ro [13:27:54] I will have a dig around [13:47:13] I've opened T438961 about the rclone problems; I think it probably makes sense to use upstream's 1.75.1 binary to copy the 4 currently-impacted objects from eqiad to codfw so that we don't end up deleting them entirely (which will become a risk once codfw is primary). [13:47:13] T438961: rclone cannot handle objects with 0x201B / SINGLE HIGH-REVERSED-9 QUOTATION MARK / ‛ in the name - https://phabricator.wikimedia.org/T438961 [17:34:15] FIRING: [70x] MysqlReplicationLagPtHeartbeat: MySQL instance db1156:9104 has too large replication lag (11m 21s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [18:22:48] FIRING: [114x] MysqlReplicationLagPtHeartbeat: MySQL instance db1156:9104 has too large replication lag (1h 1m 21s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [19:02:27] FIRING: [114x] MysqlReplicationLagPtHeartbeat: MySQL instance db1156:9104 has too large replication lag (1h 11m 21s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [19:07:57] FIRING: [49x] SystemdUnitFailed: ferm.service on aqs2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:20:25] marostegui: o/ [19:21:55] cezmunsta: thanks for coming online, so we lost power in codfw and as we have circular replication we had gtid disabled, which means the masters went down and came back with binlog issues , the classical duplicate entry [19:22:08] I just fixed s2 master by digging into binlogs and all that, but we have quite a few [19:22:37] federico3 is working on briniging the replicas back [19:22:57] thanks federico3 :) [19:23:36] I am going to start now trying to fix s3 codfw master, but I think federico3 has found some replicas also having issues, can you guys check those? [19:24:04] Do we have an easy reference list somewhere? [19:24:14] I'm tracking them in a task [19:25:13] Which task? [19:25:17] cezmunsta: we have this for now https://docs.google.com/document/d/14-BuH7ylcR2wEVJAawvhW04wnVlQiszrJQmtqvJmghU/edit?tab=t.0 [19:25:36] ack [19:25:41] ok I am going to start with s3, I will remove irc window as I need to focus on all the positions etc [19:29:11] federico3: is this the one? https://phabricator.wikimedia.org/T439031 I don't see any info on there, what have you done/not done so far? [19:31:58] yes, I'm adding the logs [19:34:43] s3 codfw master fixed [19:35:20] We need to enable GTID on the masters too, because if this happens again during the night or in the next few hours, they'll have the same issue [19:36:36] rolling back circular repl? [19:37:11] federico3: that was done earlier [19:38:39] starting s7 codfw master binlog/gtid archeology work to fix it [19:40:14] FIRING: [110x] MysqlReplicationLagPtHeartbeat: MySQL instance db1156:9104 has too large replication lag (1h 15m 4s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [19:44:22] FIRING: [52x] SystemdUnitFailed: ferm.service on aqs2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:44:39] s7 master fixed [19:45:51] FIRING: [110x] MysqlReplicationLagPtHeartbeat: MySQL instance db1156:9104 has too large replication lag (1h 15m 4s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [19:46:45] The DBs that show as down in orc - are they off? [19:47:49] yes [19:47:54] we need to bring them back and start replication [19:47:57] and see if they are broken [19:48:01] I am going to fix s8 master now [19:48:02] bye [19:56:14] FIRING: [160x] MysqlReplicationLagPtHeartbeat: MySQL instance db1156:9104 has too large replication lag (1h 15m 4s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [19:57:36] FIRING: [71x] SystemdUnitFailed: ferm.service on aqs2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [19:59:06] Working through the replicas that are down atm [20:00:25] cheers, s8 got fixed and i am with s1 now [20:01:09] replicas in s2 also look happy [20:01:53] FIRING: [72x] SystemdUnitFailed: ferm.service on aqs2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:02:34] federico3: are you doing down hosts too? [20:02:45] I just found you on one that I was checking [20:03:06] yes I've been starting replicas and updating the task [20:03:34] want to do some of the hosts? [20:04:11] FIRING: [73x] SystemdUnitFailed: ferm.service on aqs2001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:04:14] I only see replication issues, not the down hosts [20:04:51] going to start fixing x1 master [20:05:06] if someone can confirm that all sX masters are now fixed and I didn't leave anything back? [20:05:09] i will go to x1 [20:05:35] cezmunsta: I'm starting the process, then enabling repl and checking if it fails [20:06:06] Please check the logs before starting replication, to make sure that crash recovery looks OK [20:06:27] marostegui: s all look ok in orc on the primary [20:06:36] thanks :*** [20:06:40] that's good news [20:07:03] x and es6/7 do too [20:07:20] x1 is broken, I am fixing it now [20:07:31] FIRING: [79x] MysqlReplicationLagPtHeartbeat: MySQL instance db1158:9104 has too large replication lag (1h 17m 2s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [20:08:22] FIRING: [29x] SystemdUnitFailed: cassandra-b.service on aqs2005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:10:10] FIRING: [60x] MysqlReplicationLagPtHeartbeat: MySQL instance db1211:9104 has too large replication lag (51m 44s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [20:10:26] x1 is done, replicas there can also be done [20:12:11] FIRING: [29x] SystemdUnitFailed: cassandra-b.service on aqs2005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:12:25] starting with es6 [20:12:31] FIRING: [66x] MysqlReplicationLagPtHeartbeat: MySQL instance db1211:9104 has too large replication lag (51m 44s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [20:12:34] what about the replicas not on 3306, shall I skip them for now? [20:12:42] thoise are the backup sources [20:12:47] we can leave them for now yeah [20:13:31] es6 fixed, the master and the replicas are all catching up now [20:13:43] and test-s1? [20:14:09] i think we can leave that for the end [20:16:21] workin on the down ones in s2, currently the backup box [20:20:53] i've finished the masters and parsercache [20:21:06] ms1 looks broke, i will get it fixed [20:21:35] FIRING: [68x] MysqlReplicationLagPtHeartbeat: MySQL instance db1211:9104 has too large replication lag (51m 44s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [20:21:43] db2224 has a lot of warnings https://phabricator.wikimedia.org/P96519 [20:22:03] FIRING: [26x] SystemdUnitFailed: cassandra-b.service on aqs2005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:22:05] should I enable repl anyways? [20:22:13] yes, those are fine [20:22:18] FIRING: [68x] MysqlReplicationLagPtHeartbeat: MySQL instance db1211:9104 has too large replication lag (51m 44s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [20:22:30] ok what else is needed, I've finished the masters [20:22:35] I can take on more things [20:23:10] mX sections are donw, but those I can do tomorrow, they aren't used [20:23:14] FIRING: [68x] MysqlReplicationLagPtHeartbeat: MySQL instance db1211:9104 has too large replication lag (51m 44s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [20:23:27] s2 nearly caught up, just doing the 2 down in s1 [20:23:42] anyone doing x3 backup sources or should I? [20:24:00] actually normal replicas too, I will take x3 entirely [20:24:04] which host is that on? I have 1 of the backup servers still catching up [20:24:06] FIRING: [23x] SystemdUnitFailed: cassandra-b.service on aqs2005:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:24:08] ack [20:24:16] FIRING: [70x] MysqlReplicationLagPtHeartbeat: MySQL instance db1211:9104 has too large replication lag (51m 44s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [20:24:20] cezmunsta: db2200 is the backup for x3 [20:24:31] I'm following the list in the outage doc in numerical order [20:24:46] federico3: i think orchestrator will give you a faster view [20:24:58] marostegui: ah OK, not the one that I have (s2) [20:25:01] cezmunsta: do I do db2200 then? the backup source for x3? [20:25:48] @marostegui yes but it's quicker to do a for loop and skip masters etc [20:26:06] federico3: ok :) [20:26:19] ok i reached x1 db2231, should I start it? [20:26:22] marostegui: yes please if you are doing x3 [20:26:32] doing it! [20:26:52] I'm skipping m1 [20:27:06] s2 caught up [20:27:11] and m2 [20:27:12] yeah leave mX for tomorrow, I can do those [20:27:30] I will do x1, which looks unhappy [20:28:28] doing s1 backup atm [20:28:43] I will take care of x1 backup source too, as I am doing that section [20:28:55] ok, all done except: masters, m*, x* which I did not touch [20:29:01] yep [20:29:21] can I stop/start repl in s7/ [20:29:24] can I stop/start repl in s7? [20:29:32] yep [20:29:54] the backup sources are pending in s7, I can do them, I am finishing x1 [20:30:20] cezmunsta: you touching db2198 and db2200 for s7? [20:30:24] if not I will get them up [20:31:33] ok, replicas in s7 cleared errors [20:31:55] doing s7 backup sources [20:32:40] marostegui: I don't think so, am I in w? [20:32:54] too many panes :D [20:32:59] haha [20:33:16] RESOLVED: MysqlReplicationLagPtHeartbeat: MySQL instance db1267:9104 has too large replication lag (3h 0m 45s) - https://wikitech.wikimedia.org/wiki/MariaDB/troubleshooting#Depooling_a_replica - https://grafana.wikimedia.org/d/000000273/mysql?orgId=1&refresh=1m&var-job=All&var-server=db1267&var-port=9104 - https://alerts.wikimedia.org/?q=alertname%3DMysqlReplicationLagPtHeartbeat [20:33:22] fixing s8 backup source [20:34:21] so only pending backup sources for s5 and s4 [20:34:27] I will get s5 [20:34:59] anything else? [20:35:09] I also fixed test-s4 XD [20:35:54] fixed s5 and s4 backup sources [20:36:24] ok I think we are done? only pending misc which I will fix tomorrow as it is not used. federico3 you done from your side? cezmunsta ? [20:36:38] XD I was looking at test-s4 and decided to skip it [20:36:41] https://usercontent.irccloud-cdn.com/file/90Uqg8Ml/image.png [20:36:47] looking good [20:37:08] tets-s4 is lagging behind but I don't care, we can do it tomorrow [20:37:50] marostegui: a few left [20:38:19] cezmunsta: which ones? [20:38:22] I see all up in orc [20:38:26] just helping them catch up [20:38:42] ah, nah, don't worry cezmunsta they will [20:38:50] please go back to the sofa [20:39:02] I will make a small summary in mediawiki_security and also leave [20:39:05] you too federico3 [20:39:09] thank you SO MUCH for all the help [20:39:17] the m* hosts? nothing urgent? [20:39:23] tomorrow [20:39:25] they aren't used [20:39:47] I wouldn't mind having dinner :D [20:40:29] I am going to enable gtid on the masters [20:40:38] +1 [20:40:38] to avoid crash unsafe XD [20:40:53] please leave you both [20:40:55] thanks again so much [20:41:02] I will leave a summary in security once done with gtid [20:45:35] we've been lucky all the masters were recoverable without having to clone [21:49:12] I've started db2160 (backup source for mX hosts in codfw) so at least a backup can be taken (but replication is stopped as the masters are down) [21:49:49] federico3: yeah, but it was a bit of a pain, I had to dig in binlogs to make sure no transactions were skipped when fixing them, it's been a bit stressful [21:49:50] XD [21:50:30] mX backup sources started in codfw, I will leave masters down till tomorrow [21:54:39] yeah, ideally the time spent in circ repl without GTID should really be minimal 🤷 [21:55:50] Started db2185 too (zarcillo/orch replica in codfw)