[06:54:29] hnowlan, godog, unless objections I'm going to depool drmrs for switches upgrade in 1h or less [06:54:56] for https://phabricator.wikimedia.org/T437984 [07:20:19] GitLab needs a short maintenance in around 15 minutes [07:51:33] GitLab maintenance done [08:06:49] fyi, bast6003 will be unreachable for ~20min the time its switch gets upgraded. We can't migrate the VM out of the rack as it's still on the old ganeti design [08:19:10] XioNoX: ack thank you [08:35:10] switch is back up, waiting for the interfaces to come up [08:36:51] interfaces up [08:56:04] o/ Is the "CoreOutboundSaturation network sre" alert something that is expected? [08:59:09] cezmunsta: see my message on _security, that seems like a monitoring bug, the interface is 40G and is doing ~10Gbps [08:59:43] XioNoX: thanks, yes I spotted that afterwards [09:01:41] rebooting asw1-b13 then I'll repool drmrs [09:15:45] switch is back up waiting for the interfaces [09:27:08] repooling drmrs [12:24:10] hey folks, I am going to disable mTLS for PKI api calls (more info in https://phabricator.wikimedia.org/T436809) [12:24:34] k [12:58:07] I am going to hold off to merge the next patch, https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344005 - it changes all the cfssl puppet execs that make api calls to pki* hosts, to avoid using mTLS. The main trouble is that as side-effect, it triggers a regeneration of a lot of certs in prod. Better to skip it when we'll be totally ok with codfw [12:58:20] sounds good [14:17:33] on rdb2013 we have an error with 6380 instance: Bad file format reading the append only file rdb2013-6380.aof.22039.incr.aof: make a backup of your AOF file, then use ./redis-check-aof --fix [14:17:38] is it already tracked somewhere? [14:19:26] yeah so it says AOF rdb2013-6380.aof.22039.incr.aof is not valid. Use the --fix option to try fixing it. [14:20:51] I think rdb2014 also has an issue - probably related to the power outage [14:21:31] I'll create a task [14:21:47] hnowlan: I think I've fixed it :D [14:21:52] oh nice! [14:22:10] yeah I'll log on sal what I've done [14:27:06] elukey: where did you see that message? not seeing similar for rdb2014 [14:27:58] I noticed FIRING: RedisInstanceDown: Redis instance down rdb2013:16380 redis_misc and then checked the redis logs [14:32:49] d'oh rdb2013 is the master for rdb2014, all good [18:45:01] jeena: The train blocker is mine. Sorry about that! Patch: https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1344753 [18:45:26] thanks duesen! Should I go ahead and deploy it? [18:45:44] please :) [18:45:47] I was about the schedule it for the deployment window, but if you just want to do it now, that will work too. I was about to close the lid, it's late here, and I need to pack for a trip tomorrow... [18:46:06] 👍 no need to wait for the window [19:05:12] error rate is decreasing, thanks again duesen [19:07:20] Sorry again for making a mess 😅 [19:07:47] it happens! [23:23:24] okay I finally have my laptop. where can I get as many wikipedia stickers as possible [23:27:57] I’m bringing a bunch to the summit I can share [23:28:27] a couple of them must be earned, though [23:28:50] like the coveted “I BROKE WIKIPEDIA .. AND THEN I FIXED IT” [23:31:07] I feel like there might be some perverse incentives here [23:31:37] lol doing it on purpose doesn’t count [23:31:52] breaking it on purpose doesn't count. fixing it on purpose is allowed [23:31:58] ^^