[09:57:46] lunch [13:38:05] \o [13:48:45] o/ [14:39:59] o/ [15:04:25] still need 5mins before joining the wed-meeting... [18:27:28] gerrit looks down, -security is fixing [18:53:23] random thought: a full text query degredation could mean skipping interwiki (incl commons). But seems awkward for a user to only sometimes get it [21:00:14] out for school run. codfw is still red (sre monitoring) after codfw lost power [21:03:47] codfw's yellow again fwiw [21:03:51] just finished restarting all the impacted hosts [21:04:19] 200 shards left to chew through before green [21:55:12] fwiw, with all traffic on eqiad it looks not that different from all traffic on codfw. the clusters are behaving the same atleast [21:59:41] btw i havent restarted the non-directly-impacted hosts, and i suspect that will be necessary. still digging into the current state of things though, i'm in a bit of a catch 22 where we're in yellow and it'd be very easy to plunge it back into red (which may ultimately be necessary depending on how stuck things are but i'm not ready to make that determination yet) [22:00:09] codfw consumer-search is still crashlooping and bulk updates are getting rejected or stalled or smth i think [22:01:29] i was hoping that would fix itself when it stopped being red :S [22:01:41] i gotta run though, will try and check in a cpl hours [22:02:01] me too lol [22:02:05] (on the hope i mean) [22:18:07] lots of shard recoveries appear stalled. for example, `svwiki_titlesuggest_1790138661`, 3.8h and 0.0% [23:02:46] hmm, would be tempted to cancel the 0% that are stuck in init and see if it tries again [23:02:55] better question though is why are they stuck... [23:04:42] fwiw, chosing a random pair (2107, 2081) that is stuck in init for 2h the network seems fine, i can open a curl connection both directions. [23:09:45] i would randomly guess restarting nodes..2107 has the most recoveries. seems like a starting point [23:10:01] but indeed with the cluster yellow that's risky... [23:15:49] ebernhardson: oh, forgot to circle back here. i've already been doing that :D things are improving [23:16:23] ebernhardson: i was seeing stuff like `delaying recovery ... as it is not listed as assigned to target node` so i think it's some weirdness where the non-impacted hosts missed some routing updates. or like one way or another, their local state view doesn't seem to agree with the cluster state [23:19:36] we're back up to 97.58% active shard pct, so things are on the mend. still more restarts left to do ofc [23:45:45] ebernhardson: okay, I've got every host restarted except 2084, which is the elected chi master [23:46:02] ebernhardson: I don't see an alternative to restarting this and letting an election commence, but lmk if you've any thoughts