[08:44:22] marostegui, bjensen, today is magru router upgrade day, going to depool the site shortly - https://phabricator.wikimedia.org/T431750 [08:44:49] thanks [08:44:51] sounds good, thanks for the heads up :) [08:44:53] good luck! [08:53:31] btullis: can I deploy Ben Tullis: hadoop-test: Switch to using the new namenodes (818afe156d) ? [08:53:36] Yes please. [08:53:39] cool [08:53:58] Thx [09:02:30] hey folks, I'd need to reboot the idp and idm active hosts, my idea is to reboot them without failover to figure out what is the impact / annoyance if we do it, since in theory the downtime should be a couple of minutes [09:03:04] * marostegui holds to his oncall phone [09:07:20] rebooting now [09:07:28] (idp) [09:08:46] the host is already up [09:11:52] and idm1001 now [09:14:12] and done [09:15:06] nice! no impact at all? [09:20:50] yeah the reboots are super quick, so the failover is overkill imho [09:21:03] sometimes it is good to do it to excercise the procedure etc.. [09:23:57] absolutely, great stuff [10:07:44] all done with magru, giving it a few min then will repool [10:14:14] repooling magru [11:32:43] I am going to deploy proton soon via https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1319068, the new image picks up security updates for chromium etc.. [12:19:16] and done [13:04:05] FYI, I plan to pick up the remaining conf* host work in codfw shortly. I anticipate smoother sailing today, but there could of course be surprises lurking. cc: oncallers federico3, Raine [13:04:38] ack thanks [13:23:46] whoever improved the reimage script: thank you a lot! [14:45:59] federico3: Raine: db1245 may resurect today after vendor onsite maintenance, I think it is no longer in puppet but heads up in case it returns with noise. Ticket is https://phabricator.wikimedia.org/T431115 just FYI, just ignore it [14:46:38] ack, thank you jynus! [14:46:41] telling you because it may happen now or in a few hours time [14:46:48] idk [14:50:37] jynus: db1245 seems to still be in puppet, at least in site.pp and hiera hosts [14:50:57] cezmunsta: I don't mean on puppet code [14:51:12] I mean on puppetserver, it probably was cleandup up automatically [14:51:19] * cezmunsta nods [14:51:21] facts are purged after a while [14:52:27] but yeah, I don't have the dependencies in my mind to know if the prometheus alert will depend on gathered facts or zarcillo, in any case, there is a chance it comes back complaining despite all my efforts to disable notifications [14:53:09] or if the netbox failed state disables something, but you get the idea :-D [15:41:05] jynus: re: reimage, anything in particular that works better? [15:58:54] what am I doing wrong here? [15:58:58] sukhe@cumin1003:~$ sudo cumin "P:lvs::realserver::ipip%enabled=true and P:benthos%use_geoip=false" [15:59:01] doesn't work [15:59:02] individually they do [15:59:07] how can I combine them? [15:59:26] the actual query I want to run is [15:59:27] sudo cumin "P:lvs::realserver::ipip%enabled=true and P:firewall%provider=nftables" [16:00:35] individually they work... [16:01:45] topranks and I are trying to figure out if we have the necessary nftables support for IPIP and wanted to check if there are any IPIP hosts that have nftables set as the firewall provider [16:01:49] that's the full context :> [16:02:10] sukhe: try "P{P:lvs::realserver::ipip%enabled=true} and P{P:firewall%provider=nftables}" [16:02:44] aha [16:02:45] indeed [16:02:47] thanks taavi! [16:02:56] topranks: ^ so no other hosts other than the currently borked urldownloaders :) [16:03:00] urldownloader[1005-1006,2005-2006].wikimedia.org [16:04:31] sukhe: ok thanks [16:14:23] brett: These "Cacheable object with Set-Cookie found" logs from Varnish, when I query them in Logstash, is that sampled in any way? E.g. certain hosts/dcs or otherwise sampled? [16:18:16] Krinkle: AFAIK, no, it's just sent every time it occurs [16:20:42] k [16:36:32] I filed https://phabricator.wikimedia.org/T433518 Varnish: "Cacheable object with Set-Cookie found" on some rare action=raw responses [18:43:48] hopefully the last update from me today: all codfw conf* hosts have been reimaged and the cluster is once again serving etcd client read traffic. status summary is at the top of the task description in T428495. [18:43:48] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495