[08:10:09] Rechecking the cookbooks for tomorrow and 00-downtime-db-readonly-checks findes 26 masters, as oppose to the expect 24 [08:15:59] slyngs: can you list them? [08:17:04] Maybe, I'll need to check how spicerack gets them. Hang on [08:17:24] Yeah just to see which ones are missing [08:17:28] or rather, added [08:19:35] So it might just be sudo cumin A:db-role-master and divided by two, minus something [08:19:45] buy one off-by-one error, get one free? [08:21:23] Okay, so spicerack finds: db[2157,2159,2161,2179,2191,2203-2205,2229,2241,2248].codfw.wmnet,db[1160,1162,1173,1184,1189,1193,1210,1220,1236,1258,1263].eqiad.wmnet,es[2037,2039].codfw.wmnet,es[1035,1037].eqiad.wmnet [08:21:33] slyngs: let me check [08:21:43] The query is sudo cumin "A:db-core and A:db-role-master" [08:24:22] slyngs: that list looks good [08:24:41] maybe it is x4 which is new (1 master in eqiad and 1 master in codfw) [08:26:07] The 24 is CORE_SECTIONS * CORE_DATACENTERS [08:26:29] So "s6", [08:26:29] "s5", [08:26:29] "s2", [08:26:29] "s7", [08:26:29] "s3", [08:26:30] "s8", [08:26:30] "s4", [08:26:31] "x4", [08:26:31] "s1", [08:26:32] "x1", [08:26:32] "x3", [08:26:33] "es6", [08:26:33] "es7", [08:26:59] Time two [08:27:08] slyngs: that is correct [08:27:10] 13 * 2 = 26 [08:27:12] What [08:27:21] Where is spicerack getting 24 [08:28:02] This is not a DB issue, I shall escalate to volans :-) [08:28:13] slyngs: ok let me know if i can help [08:28:55] https://gerrit.wikimedia.org/r/c/operations/software/spicerack/+/1337864 <- We probably need a new release [08:30:45] Thanks, I'll check on the spicerack release and let you know [08:30:51] ok! [09:03:15] marostegui: elukey is doing a spicerack release for us :-) [09:03:47] oh nice [09:03:50] elukey: thank you :* [09:06:40] <3 [09:29:51] federico3 cezmunsta lets coordinate here the es7 issue [09:30:03] ok [09:30:04] +1 [09:30:15] any error on the cookbook? [09:30:21] marostegui: yes: [09:30:41] Failed to run cookbooks.sre.switchdc.databases.prepare.PrepareSection.enable_circular_replication: MASTER_FROM es1035.eqiad.wmnet wrong SLAVE STATUS Read_Master_Log_Pos=1034005804, expected 1033818972 instead [09:31:02] I suspect the cb is caching an older value and not reading the current one when "retry" is issued perhaps [09:31:15] ^ eqiad showed as running and codfw showed as stopped at the point that it complained [09:32:03] federico3: did we see the issue with the Preparing state for es7 or was that just an earlier one? [09:32:08] ok let's disable writes in es7 for now [09:32:12] there's a little race when the cb starts repl and then checks for it: """wrong SLAVE STATUS Slave_IO_Running=Preparing, expected Yes instead""" [09:32:52] and this glitch then cascades into the cb having to go through the manual "retry" step and getting confused [09:32:54] ok es7 depooled [09:36:16] i've disabled circular replication [09:36:20] we need to check why codfw is lagging behind [09:37:28] so heartbeat isn't running on codfw [09:39:09] ok things are back again [09:39:15] let's try another run of the cookbook [09:39:24] writes are disabled in es7 [09:39:30] so let's try another run of the cookbook? [09:39:33] cezmunsta federico3 ^ [09:39:40] ack [09:40:06] https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1343935 a quick mitigation [09:41:12] There were some Sep 22 09:34:02 es1035 mysqld[4517]: 2026-09-22 9:34:02 154900019 [ERROR] Error reading packet from server: Lost connection to server during query (server_errno=2013) on the log [09:41:20] I wonder if there was some network issues? [09:41:43] yet it failed also during retries [09:42:13] i can access fine both severs from each other too... [09:42:15] dryrunning the change [09:42:18] federico3: how's the cookbook doing now? [09:42:19] right [09:42:33] can I run https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1343935 ? [09:42:50] sure [09:42:54] ok [09:43:21] remember writes are disabled in es7 [09:43:37] cb is running again [09:43:42] ok [09:44:18] @marostegui BTW you can connect to the tmux session on cumin1004 if you want [09:44:37] orc looks better now [09:44:37] things look good now in orchestrator [09:44:44] logs too [09:44:50] and it just ran ok [09:45:14] There are uncommitted dbctl changes in the alerts [09:45:19] .. from before [09:45:28] checking [09:45:41] https://grafana.wikimedia.org/d/fcdr7bv/mariadb-aggregated-replication-lag?from=now-15m&to=now&timezone=utc&refresh=30s is back to normal [09:45:42] what? so the depool script didn't commit the changes? [09:46:18] https://phabricator.wikimedia.org/P96490 [09:46:26] ok so es7 never had writes disabled... [09:46:29] let me clean that dbctl [09:47:04] yes I'm seeing the depooling uncommitted [09:47:25] i've cleaned dbctl [09:47:35] but check the above paste (later) because looks like it didn't commit things [09:47:44] logged [09:48:05] I suggest we use https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1343935 for the rest of the changes [09:48:09] +1 [09:48:29] I will create a task for the cookbook dbctl depooling thing [09:48:29] (hoping 5 s is enough) [09:48:48] federico3: I don't think it was related to the cookbook, but it doesn't hurt [09:49:19] no but the glitch in the cb creates a red herring [09:49:38] maybe also a task for doing the sleep properly (with retries, timeouts) [09:49:57] Creating one [09:50:47] can we move on to s6 ? [09:50:48] https://phabricator.wikimedia.org/T438831 [09:50:55] federico3: yes, keep going [09:51:03] new spicerack deployed on both cumin1004 and 2003 [09:51:59] https://phabricator.wikimedia.org/T438833 [09:54:46] cezmunsta: we can continue discussing in query unless we want to stay here [09:55:04] yep, return to DM is fine [09:55:13] s6 ran ok without the "Preparing" glitch [09:56:05] :) [10:13:49] The rest went through without issue [10:14:41] nice [10:14:57] I added this https://phabricator.wikimedia.org/T438831#12349386 to see if it helps on fiding the root cause [10:32:06] marostegui: The dc switchover cookbook now findes all the db masters, all good. [10:32:18] great thank you [12:59:03] marostegui: so, the issue with es7 earlier seems to be due to using "retry" when execution hit the Preparing state during PrepareSection.enable_circular_replication. The same issue hit x3, but that was during PrepareSection.master_to_restart_replication, so that retried starting pt-heartbeat rather than retrying CHANGE MASTER after replication had started [13:10:43] But what triggered that? [13:10:55] Like why it's never happened before [13:11:12] Luck [13:11:21] Connecting would also have caused it [13:39:20] cezmunsta: just read the summary on the task, very clear there, thanks for digging into that [15:18:37] Ah, "good", my rclone bug reproduces with s3 and the latest rclone [15:19:48] https://phabricator.wikimedia.org/P96495