[02:06:26] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:01:10] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12348690 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=0b63e0fc-c67d-42da-a806-56ac1e145157) set by ayounsi@cumin1004 for 2:00:00 on 3 host(s) and their services with reas... [08:01:49] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12348691 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b0de275b-4fb3-4a76-b054-678595c0ceaf) set by ayounsi@cumin1004 for 2:00:00 on 13 host(s) and their services with rea... [08:38:33] 10netops, 06Infrastructure-Foundations: Junos file descriptors exhaustion - https://phabricator.wikimedia.org/T437984#12348809 (10ayounsi) [09:55:53] 07Puppet, 06Infrastructure-Foundations: Puppetmaster volatile data not synced to all puppet frontends for a month and a half (2024-04-27 to 2024-06-10) - https://phabricator.wikimedia.org/T367113#12349284 (10LSobanski) @CDanis are you still planning to work on this or should it be unassigned? [09:56:07] 07Puppet, 06Infrastructure-Foundations: Puppetmaster volatile data not synced to all puppet frontends for a month and a half (2024-04-27 to 2024-06-10) - https://phabricator.wikimedia.org/T367113#12349285 (10LSobanski) [10:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:09:11] 10netbox, 06Infrastructure-Foundations: https://netbox-exports.wikimedia.org/dns.git is very slow to clone (using dumb HTTP) - https://phabricator.wikimedia.org/T276403#12349349 (10LSobanski) a:05ssingh→03None [10:21:25] 10netops, 06Infrastructure-Foundations, 06SRE: HE Packet Loss Issues - https://phabricator.wikimedia.org/T438836 (10cmooney) 03NEW p:05Triage→03High [10:21:33] 10netops, 06Infrastructure-Foundations, 06SRE: HE Packet Loss Issues - https://phabricator.wikimedia.org/T438836#12349437 (10cmooney) [13:43:48] FIRING: PuppetFailure: Puppet has failed on netflow1003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:48:48] FIRING: [2x] PuppetFailure: Puppet has failed on netflow1003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:53:48] FIRING: [4x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [13:58:48] FIRING: [5x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:03:48] FIRING: [7x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:03:52] FIRING: PuppetFailure: Puppet has failed on cumin1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:05:24] 10CFSSL-PKI, 06Infrastructure-Foundations, 13Patch-For-Review: Add a load balanced service in front of PKI - https://phabricator.wikimedia.org/T436809#12350265 (10elukey) Today I had a very interesting chat with Valentin about this use case, and we ended up agreeing that dropping the mTLS complexity shouldn'... [14:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:08:48] FIRING: [7x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:13:48] FIRING: [7x] PuppetFailure: Puppet has failed on netflow1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:13:52] RESOLVED: PuppetFailure: Puppet has failed on cumin1004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:23:48] FIRING: [4x] PuppetFailure: Puppet has failed on netflow1003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:28:48] RESOLVED: [2x] PuppetFailure: Puppet has failed on netflow1003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [14:38:51] hi I/F, just a heads up we'll be skipping the k8s-ingress-aux-rw service from the switchover in line with it's previously marked excluded)) [14:53:06] however more so because we suspect that zarcillo might be reliant on it in some way)) [15:02:33] Hey IF, question on the `move-vlan` cookbook. I ran `sudo cookbook sre.hosts.move-vlan inplace cirrussearch1120` and it says `The host will lose connectivity until the switch port is updated.` . Will that happen automatically as part of the cookbook? The message above it suggests the switch port was already updated [15:13:14] inflatador: yeah it's handled by the cookbook, but there is a small window when there is a missmatch between the host's IP and the switch port vlan [15:14:13] XioNoX ACK, it was stuck for 10m or so, so I rebooted the host. It's up with the new IP but it can't ping its gateway [15:14:36] I'm in a meeting, did the cookbook run successfully? [15:15:11] ah, it just failed due to a timeout. I'll try it again [15:15:26] this is very low priority so don't feel like you need to stop anything [15:15:41] inflatador: try the `sre.network.configure-switch-interfaces` cookbook [15:16:35] XioNoX ACK, running it now [15:17:48] looks like it was a noop [15:18:06] `INFO:homer.transports.json_rpc:Empty diff for lsw1-d7-eqiad.mgmt.eqiad.wmnet, skipping device.` [15:18:16] inflatador: ok, I'll have a closer look, where did you run the cookbook? [15:18:26] (which cumin host?) [15:19:50] XioNoX it's in a cumin window under my user on cumin2003, feel free to su to `bking` or whatever you need. It's fine if the host is down for awhile so no need to hurry [15:20:08] I'll jsut look at the logs [15:20:10] thx [15:23:37] inflatador: I think I see what happened, the cookbook failed after updating the host's config, and it rolled back the netbox config [15:23:45] but the host is on the new IPs [15:24:16] Yeah, I'm logged in on host console and seeing the same thing [15:24:30] updating netbox manually [15:24:35] then will run the dns cookbook [15:47:01] inflatador: host is back up on the new ip [15:47:17] XioNoX nice, thanks for your help [16:05:32] inflatador: could you review and test-cookbook https://gerrit.wikimedia.org/r/c/operations/cookbooks/+/1344018 ? [16:59:32] XioNoX 👀 [17:31:17] inflatador: ❓ [17:32:25] XioNoX sorry, my silly way of saying that I'm looking at the patch/running test-cookbook now ;) [17:34:41] haha, cool! [17:38:55] `test-cookbook` worked great, thanks for the quick turnaround! Just added +1 [17:39:21] inflatador: did you have to type skip or abort? [17:40:00] XioNoX no, but I saw the prompt. I have one more host to do, I can run it again on that host if you like [17:42:21] sure yeah, it's for a race condition that doesn't always happen [17:59:09] the second run worked perfectly, we didn't have to skip/abort though [18:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:15:44] ok, thx! [22:06:41] FIRING: [3x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed