[02:06:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:06:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:14:15] FIRING: [2x] ProbeDown: Service idp1005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [07:14:24] Sorry, that's me [07:15:07] Should clear in a minute or so. Tomcat is back [07:19:15] RESOLVED: [2x] ProbeDown: Service idp1005:443 has failed probes (http_idp_wikimedia_org_ip4) - https://wikitech.wikimedia.org/wiki/CAS-SSO#Alerting - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [10:06:40] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:47:15] 10CAS-SSO, 06Infrastructure-Foundations: Enable selective MFA for Puppetboard - https://phabricator.wikimedia.org/T437879 (10SLyngshede-WMF) 03NEW [11:47:27] 10CAS-SSO, 06Infrastructure-Foundations: Enable selective MFA for Puppetboard - https://phabricator.wikimedia.org/T437879#12317286 (10SLyngshede-WMF) p:05Triage→03Medium [14:01:25] RESOLVED: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:42:25] FIRING: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:52:25] FIRING: [2x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:57:24] FIRING: [3x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:07:25] FIRING: [4x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:17:25] FIRING: [6x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:32:41] 10CFSSL-PKI, 06Infrastructure-Foundations, 13Patch-For-Review: Add a load balanced service in front of PKI - https://phabricator.wikimedia.org/T436809#12318566 (10elukey) I filed two changes to modify pki's settings in k8s and puppet (see above). After they are reviewed and deployed, we'll be able to proceed... [15:52:25] FIRING: [6x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:57:24] FIRING: [6x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:07:25] FIRING: [6x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:12:24] FIRING: [5x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:17:24] FIRING: [4x] SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:42:25] RESOLVED: SystemdUnitFailed: requestctl-credential-refresh.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:44:07] inflatador: you broke up acoustically; what use of Kerberos outside of what did you refer to? [16:59:04] heyo, the wiki.gives domain (which is a domain we own that is pointed to Acoustic's (FR's email & SMS marketing platform) servers to use as a shorturl domain in SMS sends) cert expired on Friday. It looks like a DigiCert cert. I can't find a reference in phab for who/how that cert was provisioned, but I think we (fr-tech) need ya'll's help with [16:59:04] renewing it (preferably asap! sorry!) [17:15:56] greg-g: responding on the task [17:18:13] thanks! [17:28:31] moritzm sorry, Luca already responded but I was asking if y'all had any Krb-dependent workloads in CODFW. DPE SRE doesn't so the thought is maybe we can get rid of krb2002 and just run 2 KRB VMs in EQIAD. I'll get a ticket started just so I can verify with my team as well [17:43:11] there is limited use for cuminunpriv, but that could easily live with the additional roundtrip via eqiad [17:52:50] ACK, I'm happy to follow y'all's lead as far as keeping a footprint in CODFW. I just realized that our team doesn't need one, so I thought it was worth asking [19:16:24] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [23:16:39] FIRING: SystemdUnitFailed: production-images-weekly-rebuild.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed