[07:01:32] good morning on-callers [07:01:43] I am going to merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/1338937 for pki.discovery.wmnet [07:03:54] ack [07:05:46] aaand lovely puppet on dns1004 caused a failure [07:06:56] Name 'pki.discovery.wmnet.': resolver plugin 'geoip' rejected resource name 'disc-pki' [07:07:12] following up on #traffic, I think I need to remove it from the dns repo [07:07:41] yeah the page is my fault sorry folks, you can ack, it is just the conf reload [07:08:06] I disabled puppet but not too fast apparently [07:08:09] sorry [07:08:31] elukey: thanks :) [07:10:28] no worries! [07:12:32] ok so reasoning out loud - I am working on pki.discovery.wmnet, the change basically reverts the service to service_setup [07:12:52] and I knew it would have removed the discovery records on dns*, but gdnsd of course complains: [07:12:53] gdnsd rightfully hates me gdnsd[2119199]: Name 'pki.discovery.wmnet.': resolver plugin 'geoip' rejected resource name 'disc-pki' [07:13:07] (sorry copy/paste from #traffic but you get it) [07:13:15] so maybe the safe thing is to revert, run puppet on dns1004/4004, patch the dns repo, deploy and then retry [07:13:47] the alternative is to patch the dns repo and deploy but I am not very comfortable doing it with gdnsd in this state [07:14:53] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1346373/1/hieradata/common/service.yaml [07:19:54] all right 1004 and 4004 are happy now [07:21:56] yep ty [07:24:39] basically https://gerrit.wikimedia.org/r/c/operations/dns/+/1346377 and then apply the service.yaml change again [07:37:01] elukey: looks good [07:43:14] vgutierrez: thankss retrying! [07:43:47] * elukey didn't get Valentin upset because he was not yet paying attention to IRC, lucky day :D [07:43:54] lol [07:48:34] vgutierrez: proceeding with https://gerrit.wikimedia.org/r/c/operations/puppet/+/1346379 ok? [07:49:01] yep [07:55:25] ook way better this time, gdnsd is happy [08:47:20] cezmunsta: to complete my morning I am going to restart pybal following https://wikitech.wikimedia.org/wiki/LVS#Configure_the_load_balancers, as FYI :D [08:47:57] * cezmunsta hides [08:55:54] 🍿 [09:24:52] <_joe_> moritzm: there's reports of idp-next being down, is that known? [09:25:22] <_joe_> indeed https://idp-test.wikimedia.org/ responds with a 404 [09:28:12] <_joe_> (I can't look rn, but it was reported on slack) [09:29:09] yes, I'm testing something on it. but why would anyone except Simon and myself use it? [09:29:45] ah,saw it. [09:29:59] growtbook-staging should use the main IDPs,I'll follow up there [09:31:15] <_joe_> thanks :) [09:31:26] <_joe_> I think a lot of people have pointed things to it :) [09:32:04] <_joe_> I mean in theory having a staging env is not a bad thing but... it usually gets designed [09:35:38] <_joe_> brouberol / btullis / atsukoito ^^ [09:36:45] airflow-test-k8s uses idp-test, I'll try avoid logging out ^^; [09:36:49] elukey: there seem to be Puppet errors relating to PKI, is that expected? [09:37:10] <_joe_> atsukoito: no what I meant is you should not point services to idp-test :) [09:37:28] <_joe_> where "you" is generically "DE SREs" [09:37:38] cezmunsta: if you are talking about the network-ca-related ones yes [09:38:11] ack [09:39:05] _joe_ moritzm: I think I'll stay with idp-test tho, because the whole purpose of that instance is to be test instance for integration, and yesterday during the upgrade it was instrumental that I was on test instance so it was not scary to tail logs [09:39:31] (i think brouberol did it this way for the same reason back then) [09:42:36] you can test all the things with using the main IDPs? we need the idp-test servers to test things central to the IDP service itself (new Java versions, webauthn, new Tomcat, new CAS, new settings) and while we have a few services there to test these are additional test instances of puppetboard next to the main puppetboard service [09:43:09] if random staging services from production use idp-test this defeats the purpose of having a tests service for the core IDP functions since we can't freely make changes [09:44:44] moritzm: (this is not urgent at all) is main IDP runs as a single active instance, or does it load-balanced around multiple? anyways, yeah, I see your point, we will move it to a prod instance [09:46:33] thanks for heads-up and explanation, I'll do it a bit later T439687 [09:46:34] T439687: migrate DPE services off idp-test - https://phabricator.wikimedia.org/T439687 [09:46:39] we have one IDP in each core site and they share the sessions via Redis. and the currently active one is targeted via the idp.w.o CNAME [09:47:20] atsukoito: thanks, no rush. but idp-test will remain down for some more, chasing some mystery TLS errors in Tomcat ATM which happen with the next LDAP server cluster [10:02:29] Thanks all for the idp-test nudge. Yes, I see your point. And if you are testing new versions or settings on idp-test, we can always temporarily switch our staging/pre-prod services to use it for short-term integration tests. [10:09:17] atsukoito: idp-test.w.o is back up [10:20:33] cezmunsta: just merged a change that should fix the puppet errors! [10:20:38] lvs state is ok, I reverted (sigh)