[13:39:25] hello traffic friends - I am around and would like to repool main-eqiad etcd for client traffic at some point this morning, by way of: [13:39:25] * https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321668 [13:39:25] * https://gerrit.wikimedia.org/r/c/operations/dns/+/1321667 [13:39:25] ... the former requiring pybal restarts in eqiad, and the latter (eventually) requiring liberica restarts in drmrs, esams, magru. let me know when would be a good time for this. [13:39:25] also, as usual, I'm happy to drive the restarts as long as someone is around in the event of surprises :) [13:39:44] swfrench-wmf: cjd91 can help you with that like yesterday :) [13:41:19] cjd91: sukhe: thank you and thank you! Chris, let me know when would be a good time for you. [13:45:44] swfrench-wmf: are you free now? [13:46:16] yes, now works if that works for you [13:47:28] so, to recap, this will require: [13:47:28] 1. merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321668 [13:47:28] 2. puppet-merge [13:47:29] 3. on each LVS host, starting with 1020 (secondary): [13:47:29] 3a. run-puppet-agent [13:47:29] 3b. restart pybal.service [13:49:28] ack and thank you. sorry about the silence there; I was reading the CRs [13:49:47] no worries at all! and thank you for doing so [13:50:54] at the risk of asking a silly question, 1321667 is to be merged after 1321668? [13:51:52] not a silly question, and no they can happen concurrently / independently. [13:53:07] gotcha; thanks for answering my question. I'm going to merge 1321668 now [13:53:45] 1321667 has the effect of switching the SRV records, so: [13:53:45] 1. for clients that frequently resolve them, they'll switch over smoothly over the course of ~ 5m [13:53:45] 2. for other clients that don't re-resolve (e.g., confd), I'll rolling restart them [13:53:58] sounds good re: 1321668 - thanks! [13:55:32] alright, I'll get started on the DNS side of things [13:55:42] 👍🏽 [13:59:16] https://www.irccloud.com/pastebin/LDm9Li3w/ [14:00:20] https://www.irccloud.com/pastebin/dKXu4sqO/ [14:00:37] I see etcd connections from 1020 on conf1007 [14:00:42] * swfrench-wmf thumbs up [14:01:58] moving on to lvs1019 [14:02:38] SGTM [14:05:04] https://www.irccloud.com/pastebin/bV7GjfWb/ [14:07:39] https://www.irccloud.com/pastebin/P8Jd9MQ1/ [14:07:51] ... and I see connections from 1019 on conf1007 [14:08:07] 👍🏽 [14:08:14] moving on to lvs1018 [14:10:53] https://www.irccloud.com/pastebin/cuvC6J3F/ [14:12:03] is it normal for clouddumps1001 to be disabled? https://www.irccloud.com/pastebin/2usslHUB/ [14:12:45] if it has been depooled, that's what I would expect, yeah - let's see [14:13:28] $ confctl select 'name=clouddumps1001.wikimedia.org' get [14:13:28] {"clouddumps1001.wikimedia.org": {"weight": 100, "pooled": "yes"}, "tags": "dc=eqiad,cluster=dumps,service=dumps-https"} [14:13:28] {"clouddumps1001.wikimedia.org": {"weight": 100, "pooled": "no"}, "tags": "dc=eqiad,cluster=dumps,service=dumps-nfs"} [14:13:28] {"clouddumps1001.wikimedia.org": {"weight": 100, "pooled": "yes"}, "tags": "dc=eqiad,cluster=dumps,service=dumps-rsync"} [14:13:52] so, it looks like (as I write this) it's depooled on the dumps-nfs service [14:15:05] thank you for looking. I didn't think to check if it was depooled 🤦🏽‍♂️ [14:16:07] no problem! this one's a bit weird, since (IIRC) the nomenclature in this output from pybal isn't quite what one thinks [14:16:27] i.e., enabled / disabled = pooled / not-pooled in the conftool sense [14:16:40] TIL [14:16:53] whereas "pooled" or not (in the output) is actually just whether the backend is included [14:17:35] I might be misremembering exactly, but that's a testament to it being confusing :D [14:18:07] thank you for the clarification [14:18:26] moving on to lvs1017 [14:19:07] * swfrench-wmf thumbs up [14:20:52] https://www.irccloud.com/pastebin/xXLmrCiu/ [14:21:22] https://www.irccloud.com/pastebin/ExOHbjZ3/ [14:21:48] nice! [14:21:49] icinga's not quite there yet, though [14:22:22] looks like the next check should fire momentarily [14:22:54] there we go [14:25:02] (another possibly silly question) is there anything I need to do for 1321667? [14:25:30] that's where the liberica restarts come in :) [14:26:00] gotcha [14:26:13] so, now that we've waited at least 5m (for the TTL on the SRV records to expire), we're ready to restart liberica in drmrs, esams, and magru [14:28:37] it's not as urgent as the pybal restarts (and since we're repooling *after* the maintenance, getting those libericas to reflect the SRV record change isn't blocking any work), but is still something we should do "soon" [14:28:54] which is to say, now is a fine pause point if you need a short brea [14:28:56] *break [14:30:45] ok. would you mind giving me 10 minutes before starting work on it? [14:31:15] sounds good! I'll be here [14:40:44] swfrench-wmf: thanks for waiting for me. I'm starting the cookbook for liberica in ulsfo now [14:41:01] no problem! [14:41:03] oh wait ... [14:41:08] no need to restart ulsfo [14:41:25] it's just drmrs, esams, and magru [14:41:54] yeah but let's not interrupt the cookbook now [14:41:57] let it finish [14:42:03] basically, the CDN sites in the "half" of the world that would normally read from eqiad [14:42:17] otherwise it will be in a weird state and that hasn't worked out nicely in the past (understandably) [14:42:21] +1 yeah don't interrupt anything if it's already started :) [14:42:42] well, damn. I'm sorry about that. will let you all know when it's done [14:43:09] no problem - it's liberica, so a restart is kinda a non-event :) [14:43:48] * cjd91 is relieved [14:44:02] ok. starting on esams now [14:44:16] * swfrench-wmf thumbs up [14:47:17] https://www.irccloud.com/pastebin/aAUv9mZP/ [14:47:33] doing magru next [14:48:43] sounds good - thanks! [14:52:43] https://www.irccloud.com/pastebin/EkIvighM/ [14:52:56] moving on to drmrs [14:58:32] https://www.irccloud.com/pastebin/Nly6RpMM/ [14:59:45] great, thank you very much for all your help this morning :) [15:00:31] sure thing! thanks for your patience :p [17:56:09] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12192715 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [19:00:54] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12192999 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (**FAIL**) - Removed from...