[00:06:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:16:03] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [00:16:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:17:28] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [00:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [00:22:14] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on sessionstore1005 - https://phabricator.wikimedia.org/T438583#12342929 (10Jclark-ctr) a:03Jclark-ctr [00:23:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:28:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:33:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.25% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [00:34:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 18.01% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:06:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:08:41] FIRING: [4x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://wikitech.wikimedia.org/wiki/Docker - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [01:10:10] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [01:11:09] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1343345 [01:11:09] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1343345 (owner: 10TrainBranchBot) [01:16:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:19:49] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1343345 (owner: 10TrainBranchBot) [01:36:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:46:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [01:54:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.59% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [01:58:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:00:04] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260920T0700) [02:00:04] Deploy window Automatic deployment of MediaWiki to pretrain wikis - see mw:Pretrain (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T0200) [02:03:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:06:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:06:38] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for fasw1-f5a-codfw.mgmt.codfw.wmnet is going to expire in 4d 14h 2m 7s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [02:12:13] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:16:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:17:13] FIRING: [3x] JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [02:20:48] FIRING: PuppetFailure: Puppet has failed on krb1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [02:24:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 20.18% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:29:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:30:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.97% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:35:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 21.35% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:36:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:36:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:41:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.52% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [02:46:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [02:54:53] PROBLEM - Host sretest2013 is DOWN: PING CRITICAL - Packet loss = 100% [03:06:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:09:41] 10ops-eqiad, 06DC-Ops: Alert for device ps1-a4-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438641 (10phaultfinder) 03NEW [03:16:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [03:30:48] RESOLVED: PuppetFailure: Puppet has failed on krb1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [03:33:20] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 21 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deployc" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343319 (https://phabricator.wikimedia.org/T438421) (owner: 10Tryvix1509) [04:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [04:28:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:36:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:46:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:55:29] (03PS1) 10Seanleong-wmde: Config change for launch of stopping sending LL notifications. [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343347 (https://phabricator.wikimedia.org/T438463) [04:58:26] FIRING: [6x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [05:03:26] FIRING: [6x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [05:06:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:10:10] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [05:14:19] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org [05:16:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:24:20] !log upgrade docker-report on build2004 to 0.0.20 T435314 [05:24:22] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [05:24:23] T435314: Migrate Docker reporting from build2002 to build2004 - https://phabricator.wikimedia.org/T435314 [05:30:25] (03CR) 10Marostegui: [C:03+1] Remove cumin1003 grant and drop it from mariadb root clients [puppet] - 10https://gerrit.wikimedia.org/r/1343057 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [05:31:25] FIRING: [10x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:36:25] FIRING: [10x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:46:25] FIRING: [10x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [05:49:57] (03CR) 10Giuseppe Lavagetto: [C:04-1] "I guess we need to discuss this offline, but I don't understand why you needed such complication when efficient caching is guaranteed by j" [docker-images/docker-pkg] - 10https://gerrit.wikimedia.org/r/1329694 (https://phabricator.wikimedia.org/T434958) (owner: 10Dduvall) [05:51:48] FIRING: PuppetFailure: Puppet has failed on krb1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [06:00:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.76% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:00:58] (03CR) 10Marostegui: mediabackups: Run mediabackup processes via systemd (034 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [06:01:48] RESOLVED: PuppetFailure: Puppet has failed on krb1002:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetFailure [06:05:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 24.38% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:06:38] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for fasw1-f5a-codfw.mgmt.codfw.wmnet is going to expire in 4d 10h 2m 7s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [06:17:13] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [06:29:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.73% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:32:56] (03CR) 10Filippo Giunchedi: [C:03+2] cloudnfs: report zero clients in nfsd-exporter when none are connected [puppet] - 10https://gerrit.wikimedia.org/r/1342541 (https://phabricator.wikimedia.org/T436252) (owner: 10Filippo Giunchedi) [06:34:06] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin1003 grant and drop it from mariadb root clients [puppet] - 10https://gerrit.wikimedia.org/r/1343057 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [06:36:25] FIRING: [10x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [06:37:26] (03CR) 10Muehlenhoff: [C:03+2] "I created https://phabricator.wikimedia.org/T438648 for the removal of the grant from production" [puppet] - 10https://gerrit.wikimedia.org/r/1343057 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [06:38:04] (03CR) 10Muehlenhoff: [C:03+2] Add explicit hostnames to LDAP cert SNIs [puppet] - 10https://gerrit.wikimedia.org/r/1342660 (https://phabricator.wikimedia.org/T331699) (owner: 10Muehlenhoff) [06:39:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.11% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:40:45] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 22.04% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [06:41:55] (03PS3) 10Filippo Giunchedi: wmcs: export prometheus metrics from wmcs-backup [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) [06:41:55] (03PS1) 10Filippo Giunchedi: backy2: run wmcs-backup metrics once an hour [puppet] - 10https://gerrit.wikimedia.org/r/1343348 (https://phabricator.wikimedia.org/T428893) [06:46:25] FIRING: [10x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:00:04] Amir1, urbanecm, and awight: I, the Bot under the Fountain, call upon thee, The Deployer, to do UTC morning backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T0700). [07:00:04] Ameisenigel and Hide_on_rosie: A patch you scheduled for UTC morning backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [07:00:56] !log jmm@cumin2003 DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003 [07:02:13] 06SRE, 10Data-Persistence-Backup, 10database-backups: Put db2201 back into backup production as a backup source - https://phabricator.wikimedia.org/T437411#12343147 (10Marostegui) This has worked fine - I will work for the s5 dump to be also done correctly before proceeding with the removal in db2250 (dumps... [07:03:26] (03CR) 10Brouberol: "Nicely done!" [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [07:03:55] (03PS1) 10Marostegui: installserver: Do not format db1282 [puppet] - 10https://gerrit.wikimedia.org/r/1343349 [07:04:46] Hi, I'm ready [07:06:25] FIRING: [10x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:07:16] I am ready for backport [07:10:26] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [07:11:00] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [07:12:31] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [07:13:27] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [07:18:31] (03PS2) 10Seanleong-wmde: Config change for launch of stopping sending LL notifications. [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343347 (https://phabricator.wikimedia.org/T438463) [07:18:45] PROBLEM - Host krb1002 is DOWN: PING CRITICAL - Packet loss = 100% [07:21:25] FIRING: [9x] SystemdUnitFailed: docker-reporter-kubernetes-aux_eqiad-images.service on build2004:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [07:23:34] Any deployers available for backport? [07:23:39] !log filippo@cumin1004 START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet [07:26:25] (03CR) 10CWilliams: mediabackups: Run mediabackup processes via systemd (034 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [07:27:02] (03PS3) 10CWilliams: mediabackups: Run mediabackup processes via systemd [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) [07:27:20] (03CR) 10Slyngshede: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1343309 (https://phabricator.wikimedia.org/T438627) (owner: 10Slyngshede) [07:27:33] (03PS4) 10CWilliams: mediabackups: Run mediabackup processes via systemd [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) [07:28:27] (03PS3) 10Seanleong-wmde: Config change for launch of stopping sending LL notifications. [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343347 (https://phabricator.wikimedia.org/T438463) [07:33:15] (03CR) 10Marostegui: mediabackups: Run mediabackup processes via systemd (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [07:33:16] (03PS2) 10Slyngshede: mw-api-ext: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333603 (https://phabricator.wikimedia.org/T433363) [07:35:57] 06SRE, 06Infrastructure-Foundations, 06Release-Engineering-Team (Radar): apt-get broken in docker-registry.wikimedia.org/bullseye:20260830 - https://phabricator.wikimedia.org/T437069#12343246 (10MoritzMuehlenhoff) 05Open→03Resolved a:03MoritzMuehlenhoff The combination of https://gerrit.wikimedia.o... [07:36:31] (03CR) 10Slyngshede: "I've adjusted it to match the run I used to estimate mw-web." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333603 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [07:37:26] !log filippo@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet [07:39:30] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12343284 (10MatthewVernon) @KOfori I can certainly co-ordinate the hardware swap. @Jclark-ctr are the parts ready to go now? If so, can we schedule this for early your-time,... [07:39:54] !log filippo@cumin1004 START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART [07:40:33] 10ops-eqiad, 06SRE, 06DC-Ops: krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12343296 (10MoritzMuehlenhoff) >>! In T435354#12338894, @VRiley-WMF wrote: > When I try to reboot it, it seems to want to default boot into a "SystemRescue" but, monitoring it and choosing the normal L... [07:43:02] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12343304 (10MatthewVernon) @Jclark-ctr relatedly, the system is currently complaining about errors on `/dev/sda` ; do we have a spare disk available, or is that going to be a... [07:45:01] !log filippo@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART [07:50:43] (03PS5) 10CWilliams: mediabackups: Run mediabackup processes via systemd [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) [07:53:03] !log ayounsi@cumin1004 START - Cookbook sre.network.tls for network device fasw1-f5b-codfw [07:53:11] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw [07:53:21] !log ayounsi@cumin1004 START - Cookbook sre.network.tls for network device fasw1-f5a-codfw [07:53:29] !log ayounsi@cumin1004 END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw [07:54:08] (03CR) 10CWilliams: mediabackups: Run mediabackup processes via systemd (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [07:56:23] FIRING: [9x] CertAlmostExpired: gNMI TLS certificate for fasw1-f5a-codfw.mgmt.codfw.wmnet is going to expire in 4d 8h 15m 7s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [07:57:30] !log add gnmic 0.49 to trixie-wikimedia - T438291 [07:57:32] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [07:57:33] T438291: Upgrade gNMIc to 0.49.0 - https://phabricator.wikimedia.org/T438291 [07:58:59] 06SRE, 06SRE Observability: Chanserv Feature Request - Add the manager oncall in the topic of -sre-private and -operations - https://phabricator.wikimedia.org/T438657 (10MLechvien-WMF) 03NEW [07:59:07] !log install gnmic 0.49 on all netflow hosts - T438291 [07:59:09] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:01:06] 06SRE, 06SRE Observability: Chanserv Feature Request - Add the manager oncall in the topic of -sre-private and -operations - https://phabricator.wikimedia.org/T438657#12343352 (10MLechvien-WMF) #sre_observability I tagged you because using SRE tag alone is not recommended, but I'm not sure if you're the right... [08:01:52] !log restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - T438291 [08:01:54] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [08:04:25] !log filippo@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision [08:05:00] !log filippo@cumin1004 START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [08:06:46] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 21 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343100 (https://phabricator.wikimedia.org/T403380) (owner: 10Ameisenigel) [08:08:08] (03CR) 10Marostegui: mediabackups: Run mediabackup processes via systemd (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [08:15:03] !log filippo@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [08:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [08:26:47] jouncebot: nowandnext [08:26:47] No deployments scheduled for the next 1 hour(s) and 33 minute(s) [08:26:47] In 1 hour(s) and 33 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1000) [08:27:18] (03CR) 10Urbanecm: [Growth] Remove unused config variables [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341161 (https://phabricator.wikimedia.org/T392944) (owner: 10Urbanecm) [08:27:22] (03PS2) 10Urbanecm: [Growth] Remove unused config variables [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341161 (https://phabricator.wikimedia.org/T392944) [08:27:29] (03CR) 10Urbanecm: [C:03+2] [Growth] Remove unused config variables [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341161 (https://phabricator.wikimedia.org/T392944) (owner: 10Urbanecm) [08:27:36] (03CR) 10TrainBranchBot: [C:03+2] "Approved by urbanecm@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341161 (https://phabricator.wikimedia.org/T392944) (owner: 10Urbanecm) [08:28:23] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 21 UTC afternoon backport window](https://wikitech.wikimedia.org/wiki/Deployments#deplo" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343319 (https://phabricator.wikimedia.org/T438421) (owner: 10Tryvix1509) [08:28:27] (03Merged) 10jenkins-bot: [Growth] Remove unused config variables [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1341161 (https://phabricator.wikimedia.org/T392944) (owner: 10Urbanecm) [08:28:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:29:03] !log urbanecm@deploy1003 Started scap sync-world: Backport for [[gerrit:1341161|[Growth] Remove unused config variables (T392944)]] [08:29:06] T392944: Enable the iterative way of refreshing LinkRecommendations for all wikis - https://phabricator.wikimedia.org/T392944 [08:44:17] (03CR) 10Volans: [C:03+1] "LGTM, couple of questions inline. Do you have a puppet compiler run by any chance?" [puppet] - 10https://gerrit.wikimedia.org/r/1342681 (https://phabricator.wikimedia.org/T301640) (owner: 10Majavah) [08:45:20] (03PS1) 10Klausman: dse-k8s/liftwing-studio: use latest CI/CD builds [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343459 [08:45:51] !log filippo@cumin1004 START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie [08:48:23] (03CR) 10Brouberol: [C:03+1] dse-k8s/liftwing-studio: use latest CI/CD builds [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343459 (owner: 10Klausman) [08:48:58] (03CR) 10Klausman: [C:03+2] dse-k8s/liftwing-studio: use latest CI/CD builds [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343459 (owner: 10Klausman) [08:50:44] (03CR) 10Marostegui: "recheck" [puppet] - 10https://gerrit.wikimedia.org/r/1343349 (owner: 10Marostegui) [08:51:42] (03Merged) 10jenkins-bot: dse-k8s/liftwing-studio: use latest CI/CD builds [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343459 (owner: 10Klausman) [08:56:06] (03PS1) 10Fabfur: external_clouds_vendors: do not fail hard on requestctl fetch [puppet] - 10https://gerrit.wikimedia.org/r/1343462 (https://phabricator.wikimedia.org/T438658) [08:57:12] (03CR) 10CI reject: [V:04-1] external_clouds_vendors: do not fail hard on requestctl fetch [puppet] - 10https://gerrit.wikimedia.org/r/1343462 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [08:57:21] (03CR) 10Marostegui: [C:03+2] installserver: Do not format db1282 [puppet] - 10https://gerrit.wikimedia.org/r/1343349 (owner: 10Marostegui) [08:57:35] moritzm: ok to merge your change? [08:57:46] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12343893 (10Jclark-ctr) @matthewvernon I won’t be able to grab the parts from the warehouse and get to the server before 13–14 UTC. Once I get onsite today, have the parts on... [08:58:12] !log ihurbain@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [08:58:36] marostegui: ah, sorry. please do [08:58:41] moritzm: doing [09:00:23] (03CR) 10Muehlenhoff: [C:03+2] kernel: Blocklist ipsec authentication headers [puppet] - 10https://gerrit.wikimedia.org/r/1343014 (owner: 10Muehlenhoff) [09:01:42] (03CR) 10Slyngshede: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343309 (https://phabricator.wikimedia.org/T438627) (owner: 10Slyngshede) [09:01:48] !log filippo@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage [09:01:57] !log urbanecm@deploy1003 Finished scap sync-world: Backport for [[gerrit:1341161|[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s) [09:02:00] T392944: Enable the iterative way of refreshing LinkRecommendations for all wikis - https://phabricator.wikimedia.org/T392944 [09:02:22] FIRING: GnmiInterfaceCountersDrop: ... [09:02:22] lsw1-b8-codfw is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=lsw1-b8-codfw:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [09:03:41] FIRING: [4x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://wikitech.wikimedia.org/wiki/Docker - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [09:04:49] !log ihurbain@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [09:04:50] !log ihurbain@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [09:07:19] !log filippo@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage [09:07:22] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [09:08:00] (03PS2) 10Muehlenhoff: kernel: Blocklist PPPOE [puppet] - 10https://gerrit.wikimedia.org/r/1343027 [09:08:50] (03CR) 10Muehlenhoff: [C:03+2] Fix missing dependencies [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342775 (https://phabricator.wikimedia.org/T434958) (owner: 10Dduvall) [09:08:53] (03CR) 10Muehlenhoff: [V:03+2 C:03+2] Fix missing dependencies [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1342775 (https://phabricator.wikimedia.org/T434958) (owner: 10Dduvall) [09:10:10] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [09:11:41] (03CR) 10Muehlenhoff: [C:03+2] kernel: Blocklist PPPOE [puppet] - 10https://gerrit.wikimedia.org/r/1343027 (owner: 10Muehlenhoff) [09:11:51] !log ihurbain@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [09:16:11] (03Abandoned) 10Elukey: sre.hosts.provision: add virtualization workload when virt is enabled [cookbooks] - 10https://gerrit.wikimedia.org/r/1333146 (https://phabricator.wikimedia.org/T435537) (owner: 10Elukey) [09:20:03] (03PS4) 10Blake: sidecars: Specify restartPolicy: Always. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) [09:20:40] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users, growthbook-customelevatedaccess for Anil - https://phabricator.wikimedia.org/T437611#12344040 (10MoritzMuehlenhoff) 05Resolved→03Open a:05AKanji-WMF→03brouberol This isn't properly resolved: @brouberol merged https://gerrit... [09:20:50] (03CR) 10Blake: "This is now ready for review!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) (owner: 10Blake) [09:22:23] !log bump space for prometheus k8s-dse in eqiad [09:22:24] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:26:40] (03PS1) 10Muehlenhoff: proton: Bump image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343475 [09:30:36] (03CR) 10Federico Ceratto: "Yes, that's independent from this PR." [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [09:35:14] (03CR) 10Marostegui: "ok, but please create a follow up for it" [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [09:35:20] !log installing chromium security updates [09:35:20] (03CR) 10Marostegui: [C:03+1] data-persistence: Alert on depooled hosts without silence [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [09:35:20] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [09:37:07] (03PS1) 10Fabfur: external_clouds_vendors: systemd timer only if conftool is set [puppet] - 10https://gerrit.wikimedia.org/r/1343476 (https://phabricator.wikimedia.org/T438658) [09:38:11] (03CR) 10Volans: [C:03+1] "Change looks good, PCC too seems ok." [puppet] - 10https://gerrit.wikimedia.org/r/1342636 (https://phabricator.wikimedia.org/T437219) (owner: 10Majavah) [09:38:49] (03PS6) 10Majavah: Add new role for cloudinfra database backups [puppet] - 10https://gerrit.wikimedia.org/r/1342681 (https://phabricator.wikimedia.org/T301640) [09:39:07] (03CR) 10Volans: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1342637 (owner: 10Majavah) [09:39:31] (03CR) 10Majavah: [V:03+1 C:03+2] P:wmcs: Make kubeadm etcd profile more generic [puppet] - 10https://gerrit.wikimedia.org/r/1342636 (https://phabricator.wikimedia.org/T437219) (owner: 10Majavah) [09:39:44] (03CR) 10Majavah: [C:03+2] P:wmcs::etcd: Remove wide etcd metrics firewall term [puppet] - 10https://gerrit.wikimedia.org/r/1342637 (owner: 10Majavah) [09:39:46] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (CORE_DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9455/co" [puppet] - 10https://gerrit.wikimedia.org/r/1342681 (https://phabricator.wikimedia.org/T301640) (owner: 10Majavah) [09:41:57] (03CR) 10Majavah: [V:03+1] "did a PCC on the database host. this is a new role so compilation would fail on the backup node, I can make+merge a separate patch adding " [puppet] - 10https://gerrit.wikimedia.org/r/1342681 (https://phabricator.wikimedia.org/T301640) (owner: 10Majavah) [09:43:49] (03PS2) 10Fabfur: external_clouds_vendors: do not fail hard on requestctl fetch [puppet] - 10https://gerrit.wikimedia.org/r/1343462 (https://phabricator.wikimedia.org/T438658) [09:43:50] (03PS2) 10Fabfur: external_clouds_vendors: systemd timer only if conftool is set [puppet] - 10https://gerrit.wikimedia.org/r/1343476 (https://phabricator.wikimedia.org/T438658) [09:43:59] 07sre-alert-triage, 06Data-Platform-SRE: Alert in need of triage: PuppetFailure (instance apifeatureusage1001:9100) - https://phabricator.wikimedia.org/T438701 (10LSobanski) 03NEW [09:44:37] 07sre-alert-triage, 06Infrastructure-Foundations: Alert in need of triage: DiskSpace (instance apt1002:9100) - https://phabricator.wikimedia.org/T438702 (10LSobanski) 03NEW [09:45:16] 07sre-alert-triage, 06Data-Platform-SRE: Alert in need of triage: PKICertificateExpiry (instance pki2002:9100) - https://phabricator.wikimedia.org/T438703 (10LSobanski) 03NEW [09:49:28] (03CR) 10Hnowlan: [V:03+1] "PCC SUCCESS (DIFF 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9456/console" [puppet] - 10https://gerrit.wikimedia.org/r/1333763 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [09:49:52] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for thumbor toolhub wikifeeds zotero codfw [puppet] - 10https://gerrit.wikimedia.org/r/1342185 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [09:50:38] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw [09:51:17] (03CR) 10Muehlenhoff: [C:03+2] proton: Bump image [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343475 (owner: 10Muehlenhoff) [09:53:03] (03PS2) 10Majavah: P:wmcs::etcd: Remove wide etcd metrics firewall term [puppet] - 10https://gerrit.wikimedia.org/r/1342637 [09:54:24] !log klausman@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply [09:54:42] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [09:55:34] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [09:55:34] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw [09:56:26] (03CR) 10Majavah: [C:03+2] P:wmcs::etcd: Remove wide etcd metrics firewall term [puppet] - 10https://gerrit.wikimedia.org/r/1342637 (owner: 10Majavah) [09:56:36] !log klausman@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply [09:56:39] (03PS2) 10Jelto: service::catalog: Set ipip for thumbor toolhub wikifeeds zotero eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1342186 (https://phabricator.wikimedia.org/T420436) [09:58:24] (03CR) 10Federico Ceratto: [C:03+2] data-persistence: Alert on depooled hosts without silence [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [09:59:03] FIRING: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [09:59:49] !log jmm@deploy1003 helmfile [staging] START helmfile.d/services/proton: apply [09:59:51] !log filippo@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie [10:00:04] Deploy window MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1000) [10:00:42] !log jmm@deploy1003 helmfile [staging] DONE helmfile.d/services/proton: apply [10:00:59] !log jmm@deploy1003 helmfile [staging] START helmfile.d/services/proton: apply [10:01:04] !log jmm@deploy1003 helmfile [staging] DONE helmfile.d/services/proton: apply [10:01:11] (03Merged) 10jenkins-bot: data-persistence: Alert on depooled hosts without silence [alerts] - 10https://gerrit.wikimedia.org/r/1337624 (https://phabricator.wikimedia.org/T436051) (owner: 10Federico Ceratto) [10:01:15] jouncebot: nowandnext [10:01:15] For the next 0 hour(s) and 58 minute(s): MediaWiki infrastructure (UTC mid-day) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1000) [10:01:15] In 2 hour(s) and 58 minute(s): UTC afternoon backport window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1300) [10:02:02] (03CR) 10Hnowlan: [V:03+1 C:03+2] kafka: migrate tls check to prometheus, per node [puppet] - 10https://gerrit.wikimedia.org/r/1333763 (https://phabricator.wikimedia.org/T407117) (owner: 10Hnowlan) [10:02:57] !log jmm@deploy1003 helmfile [codfw] START helmfile.d/services/proton: apply [10:03:20] 10ops-eqiad, 06SRE, 10Continuous-Integration-Infrastructure, 06DC-Ops, 06Release-Engineering-Team: Check and fix BIOS on new Supermicro cloudvirt* systems - https://phabricator.wikimedia.org/T435537#12344311 (10fgiunchedi) Today I re-run the provision cookbook for cloudvirt1077 to bring it back to the st... [10:03:51] (03CR) 10Volans: [C:03+1] "Nah, that's ok." [puppet] - 10https://gerrit.wikimedia.org/r/1342681 (https://phabricator.wikimedia.org/T301640) (owner: 10Majavah) [10:04:02] (03CR) 10Vgutierrez: external_clouds_vendors: do not fail hard on requestctl fetch (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1343462 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [10:04:03] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [10:04:11] !log jmm@deploy1003 helmfile [codfw] DONE helmfile.d/services/proton: apply [10:04:20] (03CR) 10Majavah: [V:03+1 C:03+2] Add new role for cloudinfra database backups [puppet] - 10https://gerrit.wikimedia.org/r/1342681 (https://phabricator.wikimedia.org/T301640) (owner: 10Majavah) [10:05:36] (03PS10) 10Arnaudb: mesh: add opt-in websocket support in configuration 1.17.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338751 (https://phabricator.wikimedia.org/T436657) [10:06:30] (03CR) 10Volans: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1343348 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [10:07:28] (03PS6) 10Tiziano Fogli: tcpircbot: replace newlines with spaces in PRIVMSG [puppet] - 10https://gerrit.wikimedia.org/r/1342704 (https://phabricator.wikimedia.org/T438269) [10:09:09] !log jmm@deploy1003 helmfile [eqiad] START helmfile.d/services/proton: apply [10:10:11] (03PS3) 10Fabfur: external_clouds_vendors: do not fail hard on requestctl fetch [puppet] - 10https://gerrit.wikimedia.org/r/1343462 (https://phabricator.wikimedia.org/T438658) [10:10:11] (03PS3) 10Fabfur: external_clouds_vendors: systemd timer only if conftool is set [puppet] - 10https://gerrit.wikimedia.org/r/1343476 (https://phabricator.wikimedia.org/T438658) [10:10:15] (03CR) 10Arnaudb: [C:04-1] "I ran a kind cluster test on this before writing the stack: with `mesh.upstream_timeout` at the production default (60s), an upgraded conn" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338751 (https://phabricator.wikimedia.org/T436657) (owner: 10Arnaudb) [10:11:39] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [10:11:53] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for thumbor toolhub wikifeeds zotero eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1342186 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [10:12:03] (03CR) 10David Caro: [C:03+1] "Just did a check and we are not really removing any language, so this is ok imo, people can upgrade to the newer version of their language" [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1342772 (https://phabricator.wikimedia.org/T400258) (owner: 10Majavah) [10:12:23] (03CR) 10Majavah: [C:03+2] Remove bullseye based images [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1342772 (https://phabricator.wikimedia.org/T400258) (owner: 10Majavah) [10:12:24] !log jmm@deploy1003 helmfile [eqiad] DONE helmfile.d/services/proton: apply [10:12:31] (03CR) 10Fabfur: external_clouds_vendors: do not fail hard on requestctl fetch (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1343462 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [10:13:33] (03Merged) 10jenkins-bot: Remove bullseye based images [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1342772 (https://phabricator.wikimedia.org/T400258) (owner: 10Majavah) [10:15:45] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.42% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [10:16:20] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [10:17:26] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [10:17:26] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad [10:17:28] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [10:19:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 23.11% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [10:19:59] !log zabe@deploy1003 mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704|T438704]]' [10:20:02] T438704: Request to move translatable page: Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment - https://phabricator.wikimedia.org/T438704 [10:21:01] !log zabe@deploy1003 mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704|T438704]]' [10:21:02] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet [10:21:42] !log zabe@deploy1003 mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704|T438704]]' [10:21:53] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users, growthbook-customelevatedaccess for Anil - https://phabricator.wikimedia.org/T437611#12344381 (10tappof) Okay, I've aligned the LDAP group membership, so `anilk` is now a member of `cn=wmf`. Thanks @MoritzMuehlenhoff [10:21:53] PROBLEM - Confd vcl based reload on cp6016 is CRITICAL: reload-vcl failed to run since 0h, 2 minutes. https://wikitech.wikimedia.org/wiki/Varnish [10:23:55] 06SRE, 10SRE-Access-Requests: Requesting access to Superset Dashboard for Hany EL Mokadem - https://phabricator.wikimedia.org/T438034#12344391 (10tappof) [10:29:17] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12344407 (10MatthewVernon) @Jclark-ctr OK, cool. I think let's try swapping the parts and rebooting and seeing if the disk is happy thereafter; if not we'll then need to try a... [10:30:13] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12344409 (10MatthewVernon) [the iDRAC thinks all the disks are fine, even though the OS thinks the drive is kaput] [10:36:11] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet [10:36:54] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet [10:38:02] !log create wbc_entity_usage table in x1 for all wikidata client wikis # T438499 [10:38:04] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [10:38:05] T438499: Create wbc_entity_usage on x1 for all wikidata client wikis - https://phabricator.wikimedia.org/T438499 [10:39:32] (03CR) 10Vgutierrez: [C:03+1] external_clouds_vendors: systemd timer only if conftool is set (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343476 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [10:41:40] (03PS1) 10Jelto: service::catalog: Set ipip for linkrecommendation* device-analytics rest-gateway [puppet] - 10https://gerrit.wikimedia.org/r/1343478 (https://phabricator.wikimedia.org/T420436) [10:41:43] (03PS1) 10Jelto: service::catalog: Set ipip for linkrecommendation* device-analytics rest-gateway [puppet] - 10https://gerrit.wikimedia.org/r/1343479 (https://phabricator.wikimedia.org/T420436) [10:46:38] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip for linkrecommendation* device-analytics rest-gateway [puppet] - 10https://gerrit.wikimedia.org/r/1343478 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [10:46:55] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip for linkrecommendation* device-analytics rest-gateway [puppet] - 10https://gerrit.wikimedia.org/r/1343479 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [10:47:27] 07sre-alert-triage, 06Infrastructure-Foundations: Alert in need of triage: DiskSpace (instance apt1002:9100) - https://phabricator.wikimedia.org/T438702#12344496 (10MoritzMuehlenhoff) The root cause is https://phabricator.wikimedia.org/T438715 [10:50:09] (03CR) 10Daniel Kertesz: external_clouds_vendors: do not fail hard on requestctl fetch (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343462 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [10:50:19] !log urbanecm@deploy1003 mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704|T438704]]' [10:50:22] T438704: Request to move translatable page: Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment - https://phabricator.wikimedia.org/T438704 [10:51:23] FIRING: [7x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:53:22] (03PS1) 10Tiziano Fogli: admin/data: grant access to hany (analytics_privatedata_users l1) [puppet] - 10https://gerrit.wikimedia.org/r/1343486 (https://phabricator.wikimedia.org/T438034) [10:54:09] 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to Superset Dashboard for Hany EL Mokadem - https://phabricator.wikimedia.org/T438034#12344511 (10tappof) [10:54:14] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet [10:54:27] (03Abandoned) 10Fabfur: external_clouds_vendors: do not fail hard on requestctl fetch [puppet] - 10https://gerrit.wikimedia.org/r/1343462 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [10:54:42] (03CR) 10CI reject: [V:04-1] admin/data: grant access to hany (analytics_privatedata_users l1) [puppet] - 10https://gerrit.wikimedia.org/r/1343486 (https://phabricator.wikimedia.org/T438034) (owner: 10Tiziano Fogli) [10:56:23] FIRING: [7x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:59:08] !incidents [10:59:08] 8354 (RESOLVED) ATSBackendErrorsHigh cache_text sre (releases.discovery.wmnet eqiad) [10:59:15] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet [11:01:58] (03Abandoned) 10Slyngshede: site.pp move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1331590 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [11:04:21] !log urbanecm@deploy1003 mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704|T438704]]' [11:04:24] T438704: Request to move translatable page: Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment - https://phabricator.wikimedia.org/T438704 [11:07:57] (03PS2) 10Tiziano Fogli: admin/data: grant access to hany (analytics_privatedata_users l1) [puppet] - 10https://gerrit.wikimedia.org/r/1343486 (https://phabricator.wikimedia.org/T438034) [11:09:28] (03PS5) 10Slyngshede: wmnet: update CNAME records for DB masters to codfw [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) [11:10:04] (03CR) 10Slyngshede: "S5 moved" [dns] - 10https://gerrit.wikimedia.org/r/1319815 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [11:12:20] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12344626 (10Certes) Me too. I almost always preview, and haven't tried submitting without a preview recently. Another user has reported an unexpected change of TA. That may share a... [11:12:34] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for an-worker1187.mgmt:22 - https://phabricator.wikimedia.org/T438607#12344627 (10Jclark-ctr) a:03Jclark-ctr [11:13:12] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet [11:13:57] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet [11:15:39] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for ms-fe1012.mgmt:22 - https://phabricator.wikimedia.org/T438609#12344636 (10Jclark-ctr) 05Open→03Resolved a:03Jclark-ctr reseated cable link returned [11:16:09] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users, growthbook-customelevatedaccess for Anil - https://phabricator.wikimedia.org/T437611#12344639 (10MoritzMuehlenhoff) 05Open→03Resolved a:05brouberol→03tappof >>! In T437611#12344381, @tappof wrote: > Okay, I've aligned t... [11:18:09] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for cirrussearch1089.mgmt:22 - https://phabricator.wikimedia.org/T438606#12344650 (10Jclark-ctr) 05Open→03Resolved a:03Jclark-ctr Reseated cable [11:19:28] (03CR) 10Btullis: [C:03+2] install_server: Add a UEFI partman recipe for cephosd servers [puppet] - 10https://gerrit.wikimedia.org/r/1343133 (https://phabricator.wikimedia.org/T438213) (owner: 10Btullis) [11:20:32] FIRING: CalicoKubeControllersDown: Calico Kubernetes Controllers not running - https://wikitech.wikimedia.org/wiki/Calico#Kube_Controllers - TODO - https://alerts.wikimedia.org/?q=alertname%3DCalicoKubeControllersDown [11:21:41] FIRING: SystemdUnitFailed: wmf_auto_restart_krb5-admin-server.service on krb2002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [11:22:09] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for an-worker1187.mgmt:22 - https://phabricator.wikimedia.org/T438607#12344660 (10Jclark-ctr) 05Open→03Resolved Reseated cable [11:22:31] 06SRE, 06Infrastructure-Foundations: Create nodejs 26 production images - https://phabricator.wikimedia.org/T437509#12344663 (10MoritzMuehlenhoff) @Jdforrester-WMF This is now available in our Docker registry: https://docker-registry.wikimedia.org/nodejs26-slim/tags/ Let me know if this works for you or if th... [11:24:08] 10ops-eqiad, 06SRE, 06DC-Ops: Unresponsive management for cirrussearch1108.mgmt:22 - https://phabricator.wikimedia.org/T438601#12344683 (10Jclark-ctr) 05Open→03Resolved a:03Jclark-ctr reseated cable [11:24:09] (03CR) 10Btullis: [C:03+2] dse-k8s: allow Pod egress to the second eqiad Pod range [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343097 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [11:24:31] 10ops-eqiad, 06SRE, 06DC-Ops: Alert for device ps1-a4-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438641#12344693 (10Jclark-ctr) a:03Jclark-ctr [11:25:32] RESOLVED: CalicoKubeControllersDown: Calico Kubernetes Controllers not running - https://wikitech.wikimedia.org/wiki/Calico#Kube_Controllers - TODO - https://alerts.wikimedia.org/?q=alertname%3DCalicoKubeControllersDown [11:26:54] 10ops-eqiad, 06SRE, 06DC-Ops: Alert for device ps1-a4-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438641#12344710 (10Jclark-ctr) ps1-a4-eqiad.mgmt.eqiad.wmnet #1: Sensor: Line, AA:L2, Current Value: 12.14 A (current) Thresholds: High: 12 [11:27:03] 10ops-eqiad, 06SRE, 06DC-Ops: Alert for device ps1-a4-eqiad.mgmt.eqiad.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438641#12344711 (10Jclark-ctr) 05Open→03Resolved [11:31:25] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply [11:31:39] !log jclark@cumin1004 START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART [11:31:49] !log jclark@cumin1004 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART [11:32:12] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply [11:32:13] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply [11:33:00] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply [11:33:01] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply [11:33:29] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply [11:33:30] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply [11:34:11] PROBLEM - Host dse-k8s-worker2004 is DOWN: PING CRITICAL - Packet loss = 100% [11:34:21] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply [11:34:23] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply [11:34:32] !log jclark@cumin1004 START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie [11:34:43] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12344732 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jclark@cumin1004 for host ml-serve1016.eqiad.wmnet with OS trixie [11:34:55] (03PS1) 10KartikMistry: machinetranslation: staging: Update to 2026-09-21-112314-production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343502 (https://phabricator.wikimedia.org/T437213) [11:35:02] FIRING: KubernetesCalicoDown: dse-k8s-worker2004.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=codfw%20prometheus%2Fk8s-dse&var-instance=dse-k8s-worker2004.codfw.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [11:35:30] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply [11:35:31] (03Merged) 10jenkins-bot: dse-k8s: allow Pod egress to the second eqiad Pod range [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343097 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [11:35:32] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply [11:35:50] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply [11:35:51] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply [11:36:00] FIRING: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker2004 has a BGP session which is not in the 'established' state. [11:36:35] (03PS4) 10Muehlenhoff: Inline profile::docker::reporter::credentials and use the same Hiera option [puppet] - 10https://gerrit.wikimedia.org/r/1342656 (https://phabricator.wikimedia.org/T435314) [11:37:03] RECOVERY - Host dse-k8s-worker2004 is UP: PING OK - Packet loss = 0%, RTA = 31.75 ms [11:37:07] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply [11:37:08] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply [11:37:38] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply [11:37:39] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply [11:38:17] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply [11:38:18] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply [11:38:52] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply [11:38:53] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply [11:39:40] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply [11:39:41] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply [11:40:02] RESOLVED: KubernetesCalicoDown: dse-k8s-worker2004.codfw.wmnet is not running calico-node Pod - https://wikitech.wikimedia.org/wiki/Calico#Operations - https://grafana.wikimedia.org/d/G8zPL7-Wz/?var-dc=codfw%20prometheus%2Fk8s-dse&var-instance=dse-k8s-worker2004.codfw.wmnet - https://alerts.wikimedia.org/?q=alertname%3DKubernetesCalicoDown [11:40:14] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet [11:40:19] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply [11:40:20] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply [11:40:53] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply [11:40:54] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply [11:41:00] RESOLVED: [2x] NodeBGPSessionStatusNotEstablished: Kubernetes node dse-k8s-worker2004 has a BGP session which is not in the 'established' state. [11:41:30] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply [11:41:31] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply [11:42:12] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply [11:42:14] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply [11:42:47] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply [11:42:49] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply [11:43:37] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply [11:43:38] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply [11:44:06] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply [11:44:07] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply [11:44:31] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply [11:44:32] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply [11:44:56] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply [11:44:57] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply [11:45:13] !log btullis@cumin1004 START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet [11:45:31] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply [11:45:32] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply [11:45:38] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply [11:47:29] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: Q1:rack/setup/install 10 new ceph nodes - https://phabricator.wikimedia.org/T438213#12344812 (10BTullis) a:05BTullis→03None [11:53:34] (03PS5) 10Blake: sidecars: enable a restartPolicy: Always option. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) [11:55:17] (03PS6) 10Blake: sidecars: create a restartPolicy: Always option. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333837 (https://phabricator.wikimedia.org/T417800) [11:56:02] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1342656 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [12:05:25] !log btullis@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet [12:14:19] (03CR) 10Arendpieter: "recheck" [software/bitu] - 10https://gerrit.wikimedia.org/r/1341185 (https://phabricator.wikimedia.org/T392350) (owner: 10Arendpieter) [12:17:37] (03PS1) 10JavierMonton: stream: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343531 (https://phabricator.wikimedia.org/T438460) [12:17:44] FIRING: KubernetesDeploymentUnavailableReplicas: ... [12:17:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [12:17:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [12:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [12:22:38] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for linkrecommendation* device-analytics rest-gateway [puppet] - 10https://gerrit.wikimedia.org/r/1343478 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [12:23:06] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw [12:23:38] (03CR) 10Vgutierrez: "ATS metric names look good:" [puppet] - 10https://gerrit.wikimedia.org/r/1343309 (https://phabricator.wikimedia.org/T438627) (owner: 10Slyngshede) [12:25:21] (03PS1) 10Brouberol: cloudnative-pg: only alerts when no WALs were archived AND WAL archiving were attempted [alerts] - 10https://gerrit.wikimedia.org/r/1343532 [12:27:27] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [12:27:44] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [12:27:46] brouberol@deploy1003: Failed to log message to wiki. Somebody should check the error logs. [12:28:14] (03CR) 10Btullis: [C:03+1] "Great, thanks. Makes good sense." [alerts] - 10https://gerrit.wikimedia.org/r/1343532 (owner: 10Brouberol) [12:28:26] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs [12:28:26] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw [12:28:29] (03CR) 10Brouberol: [C:03+2] cloudnative-pg: only alerts when no WALs were archived AND WAL archiving were attempted [alerts] - 10https://gerrit.wikimedia.org/r/1343532 (owner: 10Brouberol) [12:28:32] (03CR) 10A-pizzata: [C:03+1] stream: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343531 (https://phabricator.wikimedia.org/T438460) (owner: 10JavierMonton) [12:28:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:29:19] (03CR) 10JavierMonton: [C:03+2] stream: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343531 (https://phabricator.wikimedia.org/T438460) (owner: 10JavierMonton) [12:29:41] !log brouberol@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [12:30:20] (03PS9) 10CWilliams: mediabackups: Run mediabackup processes via systemd [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) [12:30:47] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for AHan-WMF - https://phabricator.wikimedia.org/T438538#12344936 (10tappof) [12:30:48] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [12:31:28] (03CR) 10Volans: "The idea looks good, some suggestions on the code inline. I've skipped reviewing the tests for now." [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [12:32:04] FIRING: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [12:32:25] (03Merged) 10jenkins-bot: stream: webrequest-pageview [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343531 (https://phabricator.wikimedia.org/T438460) (owner: 10JavierMonton) [12:34:23] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [12:35:04] (03CR) 10JMeybohm: [C:03+1] "Optional: I think it might make sense to design this in a way that assumes `InterfaceNamePrefix` with `cali` prefix will be our new defaul" [puppet] - 10https://gerrit.wikimedia.org/r/1343023 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [12:35:23] (03CR) 10JMeybohm: [C:03+1] hieradata: Detect local Pod traffic by interface name on dse-k8s-eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343024 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [12:35:24] (03PS2) 10Jelto: service::catalog: Set ipip for linkrecommendation* device-analytics rest-gateway [puppet] - 10https://gerrit.wikimedia.org/r/1343479 (https://phabricator.wikimedia.org/T420436) [12:36:02] !log delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX) [12:36:03] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [12:36:23] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [12:37:04] RESOLVED: MediaWikiEditFailures: Elevated MediaWiki edit failures (session_loss) for cluster - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000208/edit-count?orgId=1&viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiEditFailures [12:37:09] (03CR) 10Jelto: [C:03+2] service::catalog: Set ipip for linkrecommendation* device-analytics rest-gateway [puppet] - 10https://gerrit.wikimedia.org/r/1343479 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [12:39:15] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users for AHan-WMF - https://phabricator.wikimedia.org/T438538#12344984 (10tappof) out-of-band verification in progress [12:44:22] (03CR) 10Tiziano Fogli: tcpircbot: replace newlines with spaces in PRIVMSG (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1342704 (https://phabricator.wikimedia.org/T438269) (owner: 10Tiziano Fogli) [12:46:09] (03CR) 10Marostegui: [C:03+1] "Thanks for getting this ready! PCC also looks good" [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [12:46:42] (03CR) 10Giuseppe Lavagetto: [C:04-1] "I think it makes sense to guard the inclusion of this in the class that includes it instead." [puppet] - 10https://gerrit.wikimedia.org/r/1343476 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [12:46:49] FIRING: HelmReleaseBadStatus: Helm release wdqs-next/scholarly-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:47:26] (03PS6) 10Arnaudb: mesh: document the idle timeout key the templates actually read [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338960 (https://phabricator.wikimedia.org/T436657) [12:47:26] (03CR) 10Arnaudb: "indeed, it is now a separate change :-)" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338960 (https://phabricator.wikimedia.org/T436657) (owner: 10Arnaudb) [12:47:44] RESOLVED: KubernetesDeploymentUnavailableReplicas: ... [12:47:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [12:47:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [12:48:01] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [12:48:10] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad [12:48:31] !log btullis@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [12:50:26] (03PS1) 10Giuseppe Lavagetto: puppetserver::volatile: only fetch cloud vendors if CA is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1343536 (https://phabricator.wikimedia.org/T438658) [12:50:28] (03PS1) 10Giuseppe Lavagetto: external_clouds_vendors: stop generating the datafile [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) [12:50:59] (03PS10) 10CWilliams: mediabackups: Run mediabackup processes via systemd [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) [12:51:19] !log jelto@cumin1004 START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [12:51:38] (03CR) 10CI reject: [V:04-1] external_clouds_vendors: stop generating the datafile [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [12:52:11] (03CR) 10CWilliams: mediabackups: Run mediabackup processes via systemd (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [12:53:46] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs [12:53:46] !log jelto@cumin1004 END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad [12:53:48] (03CR) 10CWilliams: [C:03+2] mediabackups: Run mediabackup processes via systemd [puppet] - 10https://gerrit.wikimedia.org/r/1343056 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [12:54:21] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [12:54:43] !log jclark@cumin1004 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie [12:54:50] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12345068 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jclark@cumin1004 for host ml-serve1016.eqiad.wmnet with OS trixie executed with errors: - m... [12:54:53] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [12:56:49] RESOLVED: HelmReleaseBadStatus: Helm release wdqs-next/scholarly-external on k8s-dse@codfw in state pending-rollback - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [12:57:50] !log brouberol@cumin1003 START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org [12:59:17] !log filippo@cumin1004 START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet [13:00:05] Lucas_WMDE, urbanecm, and TheresNoTime: I, the Bot under the Fountain, call upon thee, The Deployer, to do UTC afternoon backport window deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1300). [13:00:05] codenamenoreste, mfossati, Dreamy_Jazz, and Hide_on_rosie: A patch you scheduled for UTC afternoon backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [13:00:12] o/ [13:00:12] \o [13:00:13] I’m in a meeting and can’t deploy, sorry [13:00:30] (you can try pinging me at :45 if no other deployer showed up until then) [13:00:35] I can self-deploy mine [13:01:06] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12345097 (10Jclark-ctr) @JHathaway I’m having issues with this new device and am unable to image it. It is not seeing the two PCI network cards, only the onboard copper po... [13:01:50] !log brouberol@cumin1003 END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org [13:01:56] codenamenoreste: after you [13:02:11] Hi, ready for backport. [13:03:07] I don't see codenamenoreste in this channel, so I guess I can go ahead [13:03:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:03:41] FIRING: [4x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://wikitech.wikimedia.org/wiki/Docker - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:04:25] * TheresNoTime can do deployments if needed, however if the self-deployers could go first that'd be great [13:04:35] mfossati: go ahead imho [13:04:40] cool [13:05:03] (03CR) 10TrainBranchBot: [C:03+2] "Approved by mfossati@deploy1003 using scap backport" [extensions/MultimediaViewer] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1343122 (https://phabricator.wikimedia.org/T437076) (owner: 10Marco Fossati) [13:05:24] (03CR) 10Tiziano Fogli: [C:03+1] prometheus: report metrics from automount units [puppet] - 10https://gerrit.wikimedia.org/r/1342971 (https://phabricator.wikimedia.org/T438062) (owner: 10Filippo Giunchedi) [13:06:52] (03PS1) 10Jelto: service::catalog: Set ipip for shellbox* in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1343539 (https://phabricator.wikimedia.org/T420436) [13:06:55] (03PS1) 10Jelto: service::catalog: Set ipip for shellbox* in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343540 (https://phabricator.wikimedia.org/T420436) [13:07:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [13:07:22] RESOLVED: [2x] GnmiInterfaceCountersDrop: asw1-b3-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [13:07:25] (03Merged) 10jenkins-bot: Let AA measure eligible readers w/o beta opt-in [extensions/MultimediaViewer] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1343122 (https://phabricator.wikimedia.org/T437076) (owner: 10Marco Fossati) [13:07:42] !log mfossati@deploy1003 Started scap sync-world: Backport for [[gerrit:1343122|Let AA measure eligible readers w/o beta opt-in (T437076)]] [13:07:46] T437076: Predicting power for five-arm test - https://phabricator.wikimedia.org/T437076 [13:08:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [13:08:19] FIRING: [3x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:10:03] (03CR) 10Filippo Giunchedi: [C:03+2] prometheus: report metrics from automount units [puppet] - 10https://gerrit.wikimedia.org/r/1342971 (https://phabricator.wikimedia.org/T438062) (owner: 10Filippo Giunchedi) [13:10:10] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:10:45] !log filippo@cumin1004 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet [13:12:07] (03CR) 10Elukey: [C:03+1] Inline profile::docker::reporter::credentials and use the same Hiera option [puppet] - 10https://gerrit.wikimedia.org/r/1342656 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [13:14:04] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12345155 (10cmooney) @Jclark-ctr just to confirm I notice the task description says this should be 25G? I didn't set it up like that and I can see the SFP in lsw1-e5-eqia... [13:14:08] !log mfossati@deploy1003 mfossati: Backport for [[gerrit:1343122|Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:14:11] T437076: Predicting power for five-arm test - https://phabricator.wikimedia.org/T437076 [13:14:44] checking ... [13:15:13] !log mfossati@deploy1003 mfossati: Continuing with deployment [13:16:53] (03PS1) 10Zabe: Move wbc_entity_usage to x1 for mediawikiwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343542 (https://phabricator.wikimedia.org/T438716) [13:18:30] 06SRE, 06Infrastructure-Foundations, 07LDAP: Migrate the r/w LDAP servers to Trixie and MDB storage - https://phabricator.wikimedia.org/T331699#12345215 (10MoritzMuehlenhoff) The current openLDAP 2.4 installation uses GNUTLS, while OpenLDAP 2.6 in Trixie uses OpenSSL. There were various TLS errors, which mad... [13:19:01] PROBLEM - Druid historical on an-druid1007 is CRITICAL: PROCS CRITICAL: 0 processes with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [13:19:07] (03CR) 10Jasmine: [C:03+1] "LGTM, thanks!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333603 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [13:19:26] (03PS1) 10Fabfur: conftool: do not block on requestctl fetching errors [puppet] - 10https://gerrit.wikimedia.org/r/1343543 (https://phabricator.wikimedia.org/T438658) [13:20:26] (03PS1) 10Majavah: hieradata: Add new project-proxy hosts to cache_hosts [puppet] - 10https://gerrit.wikimedia.org/r/1343544 (https://phabricator.wikimedia.org/T438333) [13:20:29] (03PS1) 10Majavah: P:wcms::cloudvps_meta: Publish data about infra service IPs [puppet] - 10https://gerrit.wikimedia.org/r/1343545 [13:20:33] (03CR) 10Fabfur: "Abandoned in favor of I0da664d8db5b4acb140f82ffbb0c787fb3fe2501" [puppet] - 10https://gerrit.wikimedia.org/r/1343476 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [13:20:39] (03Abandoned) 10Fabfur: external_clouds_vendors: systemd timer only if conftool is set [puppet] - 10https://gerrit.wikimedia.org/r/1343476 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [13:22:01] !log mfossati@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343122|Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s) [13:22:05] T437076: Predicting power for five-arm test - https://phabricator.wikimedia.org/T437076 [13:22:26] done! [13:23:03] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [13:23:26] FIRING: [4x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://wikitech.wikimedia.org/wiki/Docker - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:23:30] Dreamy_Jazz: you [13:23:36] * you're up next [13:23:47] self-deploy I assume too? [13:24:04] (03CR) 10Majavah: [V:03+1] "PCC SUCCESS (): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9457/console" [puppet] - 10https://gerrit.wikimedia.org/r/1343545 (owner: 10Majavah) [13:25:04] (03CR) 10Muehlenhoff: [C:03+2] Inline profile::docker::reporter::credentials and use the same Hiera option [puppet] - 10https://gerrit.wikimedia.org/r/1342656 (https://phabricator.wikimedia.org/T435314) (owner: 10Muehlenhoff) [13:25:49] Yeah, in meeting so others can go first [13:26:04] (03PS1) 10Majavah: P:wmcs::etcd: Add profile to allow backing up etcd cluster data [puppet] - 10https://gerrit.wikimedia.org/r/1343547 (https://phabricator.wikimedia.org/T438731) [13:26:12] (03PS1) 10Zabe: Set temporary virtual domain to x1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343548 (https://phabricator.wikimedia.org/T438716) [13:26:34] Hide_on_rosie: ready? [13:27:34] ready [13:27:47] (03CR) 10TrainBranchBot: [C:03+2] "Approved by samtar@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343319 (https://phabricator.wikimedia.org/T438421) (owner: 10Tryvix1509) [13:28:04] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip for shellbox* in codfw [puppet] - 10https://gerrit.wikimedia.org/r/1343539 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [13:28:19] FIRING: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:28:27] (03CR) 10JMeybohm: [C:03+1] service::catalog: Set ipip for shellbox* in eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343540 (https://phabricator.wikimedia.org/T420436) (owner: 10Jelto) [13:28:46] (03Merged) 10jenkins-bot: arywiki: Create patroller and autopatrolled user groups [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343319 (https://phabricator.wikimedia.org/T438421) (owner: 10Tryvix1509) [13:29:00] !log samtar@deploy1003 Started scap sync-world: Backport for [[gerrit:1343319|arywiki: Create patroller and autopatrolled user groups (T438421)]] [13:29:03] T438421: Request patroller and autopatrolled user groups on Moroccan Arabic Wikipedia (arywiki) - https://phabricator.wikimedia.org/T438421 [13:29:42] (03CR) 10JMeybohm: [C:03+1] mesh: document the idle timeout key the templates actually read [deployment-charts] - 10https://gerrit.wikimedia.org/r/1338960 (https://phabricator.wikimedia.org/T436657) (owner: 10Arnaudb) [13:30:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 859.3ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [13:30:31] (03PS4) 10Zabe: Move wbc_entity_usage to x1 for mediawikiwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343542 (https://phabricator.wikimedia.org/T438716) [13:32:01] RECOVERY - Druid historical on an-druid1007 is OK: PROCS OK: 1 process with command name java, args org.apache.druid.cli.Main server historical https://wikitech.wikimedia.org/wiki/Analytics/Systems/Druid [13:32:11] (Out of meeting, happy to self-deploy) [13:33:14] !log samtar@deploy1003 samtar, tryvix1509: Backport for [[gerrit:1343319|arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:33:26] FIRING: [4x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://wikitech.wikimedia.org/wiki/Docker - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [13:33:35] !log brouberol@deploy1003 helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [13:33:55] Hide_on_rosie: live on mwdebug for testing [13:34:55] FIRING: [3x] CoreRouterInterfaceDown: Core router interface down - cr1-eqiad:et-0/0/0 (DISABLED) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [13:35:28] (03PS2) 10Majavah: P:wmcs::etcd: Add profile to allow backing up etcd cluster data [puppet] - 10https://gerrit.wikimedia.org/r/1343547 (https://phabricator.wikimedia.org/T438731) [13:35:41] check OK [13:36:07] !log samtar@deploy1003 samtar, tryvix1509: Continuing with deployment [13:38:19] RESOLVED: [2x] HelmReleaseBadStatus: Helm release wdqs-next/main-external on k8s-dse@codfw in state pending-upgrade - https://wikitech.wikimedia.org/wiki/Kubernetes/Deployments#Rolling_back_in_an_emergency - https://grafana.wikimedia.org/d/UT4GtK3nz?var-site=codfw&var-cluster=k8s-dse&var-namespace=wdqs-next - https://alerts.wikimedia.org/?q=alertname%3DHelmReleaseBadStatus [13:40:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 876ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [13:40:33] (03CR) 10Klausman: cookbooks/idm: Add user-cleanup cookbook for DPE SRE hosts (031 comment) [cookbooks] - 10https://gerrit.wikimedia.org/r/1342235 (https://phabricator.wikimedia.org/T437615) (owner: 10Klausman) [13:40:40] !log samtar@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343319|arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s) [13:40:43] T438421: Request patroller and autopatrolled user groups on Moroccan Arabic Wikipedia (arywiki) - https://phabricator.wikimedia.org/T438421 [13:40:45] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 807.3ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [13:40:49] Hide_on_rosie: deploy done :) [13:40:53] Dreamy_Jazz: all yours [13:40:55] Thanks! [13:40:59] (03PS1) 10Santiago Faci: Test Kitchen UI: Deploy v2.0.0 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343552 (https://phabricator.wikimedia.org/T421814) [13:41:07] Thanks [13:41:29] !log cmooney@cumin1004 START - Cookbook sre.dns.netbox [13:41:32] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply [13:41:39] !log javiermonton@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply [13:41:49] (03CR) 10TrainBranchBot: [C:03+2] "Approved by dreamyjazz@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337580 (https://phabricator.wikimedia.org/T434045) (owner: 10Novem Linguae) [13:42:21] (03Abandoned) 10Zabe: Set temporary virtual domain to x1 [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343548 (https://phabricator.wikimedia.org/T438716) (owner: 10Zabe) [13:42:50] (03Merged) 10jenkins-bot: nlwiki: enable SecurePoll local elections [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1337580 (https://phabricator.wikimedia.org/T434045) (owner: 10Novem Linguae) [13:43:01] (03PS6) 10Zabe: Move wbc_entity_usage to x1 for mediawikiwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343542 (https://phabricator.wikimedia.org/T438716) [13:43:02] !log dreamyjazz@deploy1003 Started scap sync-world: Backport for [[gerrit:1337580|nlwiki: enable SecurePoll local elections (T434045)]] [13:43:06] T434045: Enable SecurePoll on nl.wikipedia.org - https://phabricator.wikimedia.org/T434045 [13:43:26] !log reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # T438716 [13:43:29] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [13:43:30] T438716: Move wbc_entity_usage to x1 for mediawikiwiki - https://phabricator.wikimedia.org/T438716 [13:44:00] (03CR) 10Btullis: [C:03+2] network: Add the second Pod range for dse-k8s-eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343090 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [13:44:16] Dreamy_Jazz: could you ping me once you are done? [13:44:24] Sure will do [13:45:29] !log cmooney@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004" [13:45:32] !log cmooney@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004" [13:45:33] !log cmooney@cumin1004 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [13:47:35] !log dreamyjazz@deploy1003 dreamyjazz, novemlinguae: Backport for [[gerrit:1337580|nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [13:47:40] 10ops-eqiad, 06SRE, 10Data-Persistence-Backup, 06DC-Ops: db1245 crashed - https://phabricator.wikimedia.org/T431115#12345358 (10Marostegui) >>! In T431115#12330903, @VRiley-WMF wrote: > Ran a reprovisioning on this server while we wait for Dell to come back (working out a warrenty issue with them on this s... [13:48:48] Checking... [13:48:52] (03CR) 10CDanis: [C:03+1] hieradata: Detect local Pod traffic by interface name on dse-k8s-eqiad [puppet] - 10https://gerrit.wikimedia.org/r/1343024 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [13:49:06] (03CR) 10Majavah: [C:03+1] hieradata: organize wmcs cluster into separate cloud clusters (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1338132 (https://phabricator.wikimedia.org/T437272) (owner: 10Filippo Giunchedi) [13:49:33] (03CR) 10Muehlenhoff: [C:03+2] Remove cluster::management role from cumin1003 [puppet] - 10https://gerrit.wikimedia.org/r/1343055 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [13:50:34] (03CR) 10CDanis: [C:03+1] k8s: Let kube-proxy detect local Pod traffic by interface name [puppet] - 10https://gerrit.wikimedia.org/r/1343023 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [13:51:05] !log dreamyjazz@deploy1003 dreamyjazz, novemlinguae: Continuing with deployment [13:51:13] (03CR) 10Ayounsi: "I had a look at the task and a quick look at the CRs, overall lgtm (and great idea!) but let me know if I should have a closer look." [puppet] - 10https://gerrit.wikimedia.org/r/1343023 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [13:54:14] (03PS1) 10CWilliams: medibackups: Enable systemd units [puppet] - 10https://gerrit.wikimedia.org/r/1343554 [13:54:54] (03CR) 10CWilliams: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343554 (owner: 10CWilliams) [13:55:19] (03CR) 10CI reject: [V:04-1] medibackups: Enable systemd units [puppet] - 10https://gerrit.wikimedia.org/r/1343554 (owner: 10CWilliams) [13:55:32] !log dreamyjazz@deploy1003 Finished scap sync-world: Backport for [[gerrit:1337580|nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s) [13:55:36] T434045: Enable SecurePoll on nl.wikipedia.org - https://phabricator.wikimedia.org/T434045 [13:55:45] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-web releases routed via main (k8s) 853ms - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [13:55:52] Just checking for any errors per the doc for enabling SecurePoll [13:56:53] zabe: Over to you [13:56:56] Thanks! [13:57:01] (03CR) 10Zabe: [C:03+2] Move wbc_entity_usage to x1 for mediawikiwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343542 (https://phabricator.wikimedia.org/T438716) (owner: 10Zabe) [13:57:04] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [13:58:05] (03Merged) 10jenkins-bot: Move wbc_entity_usage to x1 for mediawikiwiki [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343542 (https://phabricator.wikimedia.org/T438716) (owner: 10Zabe) [13:59:35] (03PS1) 10Zabe: Set db explicitly to false for virtual-wikibase-entityusage [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343556 [13:59:49] (03CR) 10Zabe: [C:03+2] Set db explicitly to false for virtual-wikibase-entityusage [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343556 (owner: 10Zabe) [14:00:12] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:00:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:00:46] (03Merged) 10jenkins-bot: Set db explicitly to false for virtual-wikibase-entityusage [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343556 (owner: 10Zabe) [14:01:10] !log zabe@deploy1003 Started scap sync-world: Backport for [[gerrit:1343542|Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556|Set db explicitly to false for virtual-wikibase-entityusage]] [14:01:14] T438716: Move wbc_entity_usage to x1 for mediawikiwiki - https://phabricator.wikimedia.org/T438716 [14:01:22] FIRING: GnmiInterfaceCountersDrop: ... [14:01:22] asw1-b4-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?orgId=1&from=now-24h&to=now&var-site=%24__all&var-instance=asw1-b4-magru:9804&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [14:02:12] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:02:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:02:33] (03CR) 10Thibaut Le Page: [C:03+1] toolforge: add the logs-cli package [puppet] - 10https://gerrit.wikimedia.org/r/1341719 (https://phabricator.wikimedia.org/T432572) (owner: 10David Caro) [14:03:04] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [14:03:08] (03CR) 10David Caro: [C:03+2] toolforge: add the logs-cli package [puppet] - 10https://gerrit.wikimedia.org/r/1341719 (https://phabricator.wikimedia.org/T432572) (owner: 10David Caro) [14:04:39] Error: UPGRADE FAILED: release next failed, and has been rolled back due to atomic being set: cannot patch "mediawiki-next-tls-proxy-certs" with kind Certificate: Internal error occurred: failed calling webhook "webhook.cert-manager.io": failed to call webhook: Post "https://cert-manager-webhook.cert-manager.svc:443/validate?timeout=30s": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [14:05:08] !log zabe@deploy1003 Started scap sync-world: Backport for [[gerrit:1343542|Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556|Set db explicitly to false for virtual-wikibase-entityusage]] [14:05:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:05:41] zabe: lmk when you're done, I have a beta cluster config to deploy :-) [14:06:01] (03PS2) 10CWilliams: medibackups: Enable systemd units [puppet] - 10https://gerrit.wikimedia.org/r/1343554 (https://phabricator.wikimedia.org/T438012) [14:06:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:06:22] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b4-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [14:06:25] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:06:39] (03CR) 10Marostegui: [C:03+1] medibackups: Enable systemd units [puppet] - 10https://gerrit.wikimedia.org/r/1343554 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [14:06:58] (03PS3) 10CWilliams: medibackups: Enable systemd units [puppet] - 10https://gerrit.wikimedia.org/r/1343554 (https://phabricator.wikimedia.org/T438012) [14:07:16] zabe: nvm—I will do this tomorrow [14:07:23] !log cdobbins@cumin1004 START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie [14:07:24] ok:) [14:07:57] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Tuesday, September 22 UTC morning backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploy" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343347 (https://phabricator.wikimedia.org/T438463) (owner: 10Seanleong-wmde) [14:08:19] !log zabe@deploy1003 zabe: Backport for [[gerrit:1343542|Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556|Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [14:08:22] T438716: Move wbc_entity_usage to x1 for mediawikiwiki - https://phabricator.wikimedia.org/T438716 [14:08:52] (03CR) 10Awight: [C:03+1] "Great, thanks for gluing this in!" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343347 (https://phabricator.wikimedia.org/T438463) (owner: 10Seanleong-wmde) [14:08:55] !log zabe@deploy1003 zabe: Continuing with deployment [14:10:12] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:10:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:11:43] (03CR) 10Fabfur: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1343536 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [14:12:12] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:12:12] (03CR) 10CDanis: conftool: do not block on requestctl fetching errors (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343543 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [14:12:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:13:08] (03CR) 10Fabfur: "thanks, definitely better" [puppet] - 10https://gerrit.wikimedia.org/r/1343543 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [14:13:17] !log zabe@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343542|Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556|Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s) [14:13:44] FIRING: KubernetesDeploymentUnavailableReplicas: ... [14:13:44] Deployment mw-web.eqiad.main in mw-web at eqiad has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=eqiad&var-cluster=k8s&var-namespace=mw-web&var-deployment=mw-web.eqiad.main - ... [14:13:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [14:13:48] (03CR) 10Arnaudb: "Thanks for the highlight @jmeybohm@wikimedia.org, @dpogorzelski@wikimedia.org, fwiw I tested the hypothesis directly: rendered our sidecar" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342612 (https://phabricator.wikimedia.org/T437706) (owner: 10Dpogorzelski) [14:14:29] 06SRE, 06Infrastructure-Foundations, 07Epic, 07Kubernetes: aux-k8s: eqiad expansion, codfw creation, & future hopes and dreams - https://phabricator.wikimedia.org/T378742#12345494 (10CDanis) [14:14:36] (03CR) 10CDanis: [C:03+1] conftool: do not block on requestctl fetching errors [puppet] - 10https://gerrit.wikimedia.org/r/1343543 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [14:14:38] (03PS1) 10VolkerE: Revert "Enable Reading Recommendations experiment on test wiki" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343560 [14:15:00] (03PS2) 10Fabfur: conftool: do not block on requestctl fetching errors [puppet] - 10https://gerrit.wikimedia.org/r/1343543 (https://phabricator.wikimedia.org/T438658) [14:15:12] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:15:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:15:38] (03PS2) 10VolkerE: Revert "Enable Reading Recommendations experiment on test wiki" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343560 (https://phabricator.wikimedia.org/T437339) [14:16:13] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:16:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:16:19] (03CR) 10Muehlenhoff: [C:03+1] "Looks good, thanks!" [puppet] - 10https://gerrit.wikimedia.org/r/1343156 (owner: 10JHathaway) [14:17:28] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:19:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 15.74% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [14:20:44] (03CR) 10Fabfur: [C:03+2] conftool: do not block on requestctl fetching errors [puppet] - 10https://gerrit.wikimedia.org/r/1343543 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [14:20:50] !log elukey@cumin1004 START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART [14:20:54] ACKNOWLEDGEMENT - Dell PowerEdge or Supermicro Broadcom RAID Controller on an-worker1194 is CRITICAL: communication: 0 OK : controller: 1 Needs Attention : physical_disk: 1 Failed : virtual_disk: 1 OfLn : bbu: 0 OK : enclosure: 0 OK : CLI Version = 007.1910.0000.0000 Oct 08, 2021 nagiosadmin RAID handler auto-ack: https://phabricator.wikimedia.org/T438744 https://wikitech.wikimedia.org/wiki/PERCCli%23Monitoring [14:21:00] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on an-worker1194 - https://phabricator.wikimedia.org/T438744 (10ops-monitoring-bot) 03NEW [14:21:08] (03PS3) 10Fabfur: conftool: do not block on requestctl fetching errors [puppet] - 10https://gerrit.wikimedia.org/r/1343543 (https://phabricator.wikimedia.org/T438658) [14:21:25] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:21:26] (03CR) 10Jforrester: "Did we want to deploy this today?" [extensions/PersonalDashboard] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1343126 (https://phabricator.wikimedia.org/T438387) (owner: 10Zabe) [14:22:16] 10ops-eqiad, 06SRE, 06Data-Platform-SRE, 06DC-Ops: QFY2627 :rack/setup/install ml-serve1016 - https://phabricator.wikimedia.org/T438161#12345535 (10elukey) While reprovisioning for another issue I noticed: ` BIOS: IPv4PXESupport is set to Enabled, while we want Disabled BIOS: HTTPBootPolicy is set to Boot... [14:25:10] (03CR) 10Fabfur: [C:03+2] conftool: do not block on requestctl fetching errors [puppet] - 10https://gerrit.wikimedia.org/r/1343543 (https://phabricator.wikimedia.org/T438658) (owner: 10Fabfur) [14:25:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1013.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:25:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:25:38] (03PS4) 10CWilliams: medibackups: Enable systemd units [puppet] - 10https://gerrit.wikimedia.org/r/1343554 (https://phabricator.wikimedia.org/T438012) [14:26:13] !log elukey@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART [14:26:25] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [14:27:02] RECOVERY - Confd vcl based reload on cp6016 is OK: reload-vcl successfully ran 0h, 0 minutes ago. https://wikitech.wikimedia.org/wiki/Varnish [14:27:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:27:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:29:21] !incidents [14:29:22] 8354 (RESOLVED) ATSBackendErrorsHigh cache_text sre (releases.discovery.wmnet eqiad) [14:30:05] Deploy window Test Kitchen Experiment Deployment Window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1430) [14:30:27] (03CR) 10Fabfur: [C:03+2] puppetserver::volatile: only fetch cloud vendors if CA is enabled [puppet] - 10https://gerrit.wikimedia.org/r/1343536 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [14:33:36] (03PS2) 10Giuseppe Lavagetto: external_clouds_vendors: stop generating the datafile [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) [14:33:49] (03CR) 10CWilliams: [C:03+2] medibackups: Enable systemd units [puppet] - 10https://gerrit.wikimedia.org/r/1343554 (https://phabricator.wikimedia.org/T438012) (owner: 10CWilliams) [14:34:15] !log cdobbins@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage [14:34:54] PROBLEM - Check whether ferm is active by checking the default input chain on wikikube-worker1071 is CRITICAL: ERROR ferm input drop default policy not set, ferm might not have been started correctly https://wikitech.wikimedia.org/wiki/Monitoring/check_ferm [14:35:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:35:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:35:45] (03CR) 10CI reject: [V:04-1] external_clouds_vendors: stop generating the datafile [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [14:36:18] !log elukey@cumin1004 START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie [14:37:00] (03PS3) 10Giuseppe Lavagetto: external_clouds_vendors: stop generating the datafile [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) [14:37:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:37:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:38:44] FIRING: [2x] KubernetesDeploymentUnavailableReplicas: Deployment eventstreams-production in eventstreams at codfw has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [14:39:41] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage [14:39:53] (03PS1) 10Majavah: O:wmcs::toolforge: Add role for etcd backups [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) [14:40:14] (03PS3) 10Majavah: P:wmcs::etcd: Add profile to allow backing up etcd cluster data [puppet] - 10https://gerrit.wikimedia.org/r/1343547 (https://phabricator.wikimedia.org/T438731) [14:40:14] (03PS2) 10Majavah: O:wmcs::toolforge: Add role for etcd backups [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) [14:41:22] (03CR) 10CI reject: [V:04-1] O:wmcs::toolforge: Add role for etcd backups [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) (owner: 10Majavah) [14:41:38] 06SRE, 10LDAP-Access-Requests: Grant Access to ciadmin for Vaughn Walters - https://phabricator.wikimedia.org/T433615#12345676 (10hashar) >>! In T433615#12179913, @Stashbot wrote: > {nav icon=file, name=Mentioned in SAL (#wikimedia-releng), href=https://sal.toolforge.org/log/PHF_yJ8B1kByGTxAo_v3} [2026-08-... [14:41:52] (03CR) 10CI reject: [V:04-1] O:wmcs::toolforge: Add role for etcd backups [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) (owner: 10Majavah) [14:42:16] (03PS3) 10Majavah: O:wmcs::toolforge: Add role for etcd backups [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) [14:43:00] (03CR) 10LSobanski: "Provisionally approved in IF team meeting once verified that this is required." [puppet] - 10https://gerrit.wikimedia.org/r/1342807 (https://phabricator.wikimedia.org/T268199) (owner: 10Dzahn) [14:45:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:46:13] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1020.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:47:10] !log elukey@cumin1004 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie [14:47:22] (03CR) 10Volans: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) (owner: 10Majavah) [14:48:47] (03PS1) 10Filippo Giunchedi: node-exporter: fix literal backslash escaping for unit name exclusion [puppet] - 10https://gerrit.wikimedia.org/r/1343571 (https://phabricator.wikimedia.org/T438062) [14:50:22] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12345732 (10Pigsonthewing) I've been seeing the same, intermittently for several days, not only on en.Wikipedia, but also on en.Wikisource. [14:51:13] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:51:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:51:18] (03PS4) 10Majavah: O:wmcs::toolforge: Add role for etcd backups [puppet] - 10https://gerrit.wikimedia.org/r/1343566 (https://phabricator.wikimedia.org/T339934) [14:52:13] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:52:14] (03CR) 10Slyngshede: [C:03+2] mw-api-ext: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333603 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [14:53:26] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: Automate PDU Deployment Process - https://phabricator.wikimedia.org/T403173#12345749 (10ayounsi) a:03ayounsi [14:54:13] (03PS2) 10Filippo Giunchedi: node-exporter: fix literal backslash escaping for unit name exclusion [puppet] - 10https://gerrit.wikimedia.org/r/1343571 (https://phabricator.wikimedia.org/T438062) [14:54:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:54:41] (03CR) 10Filippo Giunchedi: "I'm sure there's a better way™" [puppet] - 10https://gerrit.wikimedia.org/r/1343571 (https://phabricator.wikimedia.org/T438062) (owner: 10Filippo Giunchedi) [14:55:00] (03Merged) 10jenkins-bot: mw-api-ext: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1333603 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [14:55:01] PROBLEM - Postgres Replication Lag on puppetdb2003 is CRITICAL: POSTGRES_HOT_STANDBY_DELAY CRITICAL: DB puppetdb (host:localhost) 68093760 and 11 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [14:55:13] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:56:01] RECOVERY - Postgres Replication Lag on puppetdb2003 is OK: POSTGRES_HOT_STANDBY_DELAY OK: DB puppetdb (host:localhost) 2426952 and 0 seconds https://wikitech.wikimedia.org/wiki/Postgres%23Monitoring [14:56:04] 07sre-alert-triage, 06Infrastructure-Foundations: Alert in need of triage: DiskSpace (instance apt1002:9100) - https://phabricator.wikimedia.org/T438702#12345755 (10LSobanski) 05Open→03Resolved a:03LSobanski [14:56:23] FIRING: [7x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [14:57:05] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [14:57:13] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [14:57:13] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:57:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:57:24] 06SRE, 10observability, 06Traffic, 13Patch-For-Review: HAProxy metrics go down on config reload - https://phabricator.wikimedia.org/T343000#12345760 (10dkertesz) I did some local testing to validate the approach of my [[ https://gerrit.wikimedia.org/r/c/operations/puppet/+/1343039 | CR ]] and it seems to b... [14:59:05] (03CR) 10Filippo Giunchedi: "This is what systemd reports executing with debug on and this patch https://phabricator.wikimedia.org/P96482" [puppet] - 10https://gerrit.wikimedia.org/r/1343571 (https://phabricator.wikimedia.org/T438062) (owner: 10Filippo Giunchedi) [15:01:02] (03CR) 10Fabfur: "LGTM just a small fix" [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [15:01:36] !log slyngshede@deploy1003 helmfile [codfw] START helmfile.d/services/mw-api-ext: apply [15:01:55] !log slyngshede@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply [15:03:18] (03CR) 10Slyngshede: [C:03+2] mw-web: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332725 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [15:03:49] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie [15:05:06] (03CR) 10Tiziano Fogli: [C:03+1] node-exporter: fix literal backslash escaping for unit name exclusion [puppet] - 10https://gerrit.wikimedia.org/r/1343571 (https://phabricator.wikimedia.org/T438062) (owner: 10Filippo Giunchedi) [15:05:43] (03Merged) 10jenkins-bot: mw-web: upsize for single-DC serving [deployment-charts] - 10https://gerrit.wikimedia.org/r/1332725 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [15:06:33] (03CR) 10Filippo Giunchedi: [C:03+2] node-exporter: fix literal backslash escaping for unit name exclusion [puppet] - 10https://gerrit.wikimedia.org/r/1343571 (https://phabricator.wikimedia.org/T438062) (owner: 10Filippo Giunchedi) [15:07:12] !log slyngshede@deploy1003 helmfile [codfw] START helmfile.d/services/mw-web: apply [15:07:38] !log slyngshede@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-web: apply [15:09:16] I need to use the 8.30am portals window to do a non-portals deploy. Any issues with that? I'm confirming with Jan he doesn't need to use it. [15:11:08] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [15:13:09] (03PS1) 10Jdlrobson: Fixes: '.action_context' should be string [extensions/WikimediaEvents] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1343579 (https://phabricator.wikimedia.org/T437122) [15:16:31] !log cdobbins@cumin1004 conftool action : set/pooled=yes; selector: name=ncredir3006.* [15:18:21] jouncebot: nowandnext [15:18:22] No deployments scheduled for the next 0 hour(s) and 11 minute(s) [15:18:22] In 0 hour(s) and 11 minute(s): Wikimedia Portals Update (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1530) [15:18:31] !log elukey@cumin1004 START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie [15:18:44] FIRING: [2x] KubernetesDeploymentUnavailableReplicas: Deployment eventstreams-production in eventstreams at codfw has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [15:19:27] 06SRE, 06ServiceOps, 07Datacenter-Switchover: Services without a service IP cannot automatically be switched by the switchdc cookbook - https://phabricator.wikimedia.org/T285707#12345915 (10MLechvien-WMF) a:03Blake [15:19:29] !log elukey@puppetserver1001 conftool action : set/pooled=false; selector: name=registry2004.* [15:20:53] (03PS4) 10Filippo Giunchedi: wmcs: export prometheus metrics from wmcs-backup [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) [15:20:53] (03PS2) 10Filippo Giunchedi: backy2: run wmcs-backup metrics once an hour [puppet] - 10https://gerrit.wikimedia.org/r/1343348 (https://phabricator.wikimedia.org/T428893) [15:21:10] (03CR) 10Filippo Giunchedi: "Thank you for the review volans !" [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [15:24:07] (03CR) 10CI reject: [V:04-1] wmcs: export prometheus metrics from wmcs-backup [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) (owner: 10Filippo Giunchedi) [15:24:59] PROBLEM - Host sessionstore1005 is DOWN: PING CRITICAL - Packet loss = 100% [15:25:00] 10ops-eqiad, 06SRE, 06DC-Ops: krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12345939 (10VRiley-WMF) Very strange as there is no usb stick in this machine. I will say I did get an email back from support saying the following... Gracefully shut down the system from the operatin... [15:25:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:26:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:28:26] FIRING: [5x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:29:13] ok starting shortly. [15:30:05] jan_drewniak: Your horoscope predicts another Wikimedia Portals Update deploy. May Zuul be (nice) with you. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1530). [15:30:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:30:27] (03CR) 10TrainBranchBot: [C:03+2] "Approved by jdlrobson@deploy1003 using scap backport" [extensions/WikimediaEvents] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1343579 (https://phabricator.wikimedia.org/T437122) (owner: 10Jdlrobson) [15:30:35] (03PS5) 10Filippo Giunchedi: wmcs: export prometheus metrics from wmcs-backup [puppet] - 10https://gerrit.wikimedia.org/r/1341187 (https://phabricator.wikimedia.org/T428893) [15:30:35] (03PS3) 10Filippo Giunchedi: backy2: run wmcs-backup metrics once an hour [puppet] - 10https://gerrit.wikimedia.org/r/1343348 (https://phabricator.wikimedia.org/T428893) [15:31:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1013.eqiad.wmnet, wdqs1020.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:32:02] (03Merged) 10jenkins-bot: Fixes: '.action_context' should be string [extensions/WikimediaEvents] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1343579 (https://phabricator.wikimedia.org/T437122) (owner: 10Jdlrobson) [15:32:13] FIRING: [2x] JobUnavailable: Reduced availability for job docker-registry in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:32:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:32:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:32:19] !log jdlrobson@deploy1003 Started scap sync-world: Backport for [[gerrit:1343579|Fixes: '.action_context' should be string (T437122)]] [15:32:22] T437122: Launch experiment for donor consent - https://phabricator.wikimedia.org/T437122 [15:32:50] !log slyngshede@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply [15:33:02] !log slyngshede@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply [15:35:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:35:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:36:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:36:24] !log jdlrobson@deploy1003 jdlrobson: Backport for [[gerrit:1343579|Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [15:37:04] (03CR) 10JHathaway: [C:03+2] ssh server: fix match indentation [puppet] - 10https://gerrit.wikimedia.org/r/1343156 (owner: 10JHathaway) [15:37:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:38:09] !log elukey@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage [15:39:33] !log cdobbins@cumin1004 START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie [15:39:59] !log jdlrobson@deploy1003 jdlrobson: Continuing with deployment [15:42:46] !log elukey@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage [15:44:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-web releases routed via main at eqiad: 19.24% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-web&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [15:44:59] !log jdlrobson@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343579|Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s) [15:45:02] T437122: Launch experiment for donor consent - https://phabricator.wikimedia.org/T437122 [15:48:50] (03CR) 10Cklimas: [C:03+2] wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343140 (owner: 10PipelineBot) [15:49:16] !log fceratto@deploy1003 helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' . [15:50:56] (03Abandoned) 10Jgiannelos: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1342812 (owner: 10PipelineBot) [15:51:17] (03Merged) 10jenkins-bot: wikifeeds: pipeline bot promote [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343140 (owner: 10PipelineBot) [15:53:44] RESOLVED: KubernetesDeploymentUnavailableReplicas: ... [15:53:44] Deployment eventstreams-production in eventstreams at codfw has persistently unavailable replicas - https://wikitech.wikimedia.org/wiki/Kubernetes/Troubleshooting#Troubleshooting_a_deployment - https://grafana.wikimedia.org/d/a260da06-259a-4ee4-9540-5cab01a246c8/kubernetes-deployment-details?var-site=codfw&var-cluster=k8s&var-namespace=eventstreams&var-deployment=eventstreams-production - ... [15:53:44] https://alerts.wikimedia.org/?q=alertname%3DKubernetesDeploymentUnavailableReplicas [15:54:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [15:54:51] !log cklimas@deploy1003 helmfile [staging] START helmfile.d/services/wikifeeds: apply [15:55:10] !log cklimas@deploy1003 helmfile [staging] DONE helmfile.d/services/wikifeeds: apply [15:55:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [15:56:02] (done) [16:00:04] !log elukey@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie [16:00:05] !log cklimas@deploy1003 helmfile [eqiad] START helmfile.d/services/wikifeeds: apply [16:00:13] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:00:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:00:37] !log cklimas@deploy1003 helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply [16:00:54] !log cklimas@deploy1003 helmfile [codfw] START helmfile.d/services/wikifeeds: apply [16:01:24] !log cklimas@deploy1003 helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply [16:01:41] (03PS1) 10Andrew Bogott: Nova policy.yaml: disable instance pause and suspend [puppet] - 10https://gerrit.wikimedia.org/r/1343587 (https://phabricator.wikimedia.org/T431307) [16:02:13] FIRING: [2x] JobUnavailable: Reduced availability for job docker-registry in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:02:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:03:13] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:08:11] (03CR) 10JMeybohm: [C:03+1] profile::k8s::deployment_server: add python3-docker-report [puppet] - 10https://gerrit.wikimedia.org/r/1341899 (https://phabricator.wikimedia.org/T437297) (owner: 10Elukey) [16:09:16] !log cdobbins@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage [16:09:22] (03PS1) 10Hnowlan: orchestrator: allow healthchecking over HTTP [puppet] - 10https://gerrit.wikimedia.org/r/1343590 (https://phabricator.wikimedia.org/T407329) [16:10:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:10:49] !log cmooney@cumin1004 START - Cookbook sre.dns.netbox [16:11:51] (03CR) 10JMeybohm: [C:03+1] golang: add trixie-based golang-1.26 image [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1341880 (https://phabricator.wikimedia.org/T423851) (owner: 10Blake) [16:12:13] FIRING: [4x] JobUnavailable: Reduced availability for job docker-registry in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:12:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:13:06] (03CR) 10Hnowlan: [V:03+1] "PCC SUCCESS (NOOP 1): https://integration.wikimedia.org/ci/job/operations-puppet-catalog-compiler/label=puppet7-compiler-node/9460/console" [puppet] - 10https://gerrit.wikimedia.org/r/1343590 (https://phabricator.wikimedia.org/T407329) (owner: 10Hnowlan) [16:13:19] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage [16:13:42] PROBLEM - Host wikikube-worker1364 is DOWN: PING CRITICAL - Packet loss = 100% [16:15:10] RECOVERY - Host wikikube-worker1364 is UP: PING OK - Packet loss = 0%, RTA = 0.34 ms [16:15:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:15:54] !log cmooney@cumin1004 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004" [16:15:58] !log cmooney@cumin1004 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004" [16:15:58] !log cmooney@cumin1004 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [16:15:58] 10ops-codfw, 06SRE, 06DC-Ops: Alert for device ps1-a2-codfw.mgmt.codfw.wmnet - PDU sensor over limit - https://phabricator.wikimedia.org/T438282#12346287 (10Jhancock.wm) 05Open→03Resolved [16:16:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:17:13] FIRING: [4x] JobUnavailable: Reduced availability for job docker-registry in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:19:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:19:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1011.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:20:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:20:48] (03PS2) 10Andrew Bogott: Nova policy.yaml: disable instance pause and suspend [puppet] - 10https://gerrit.wikimedia.org/r/1343587 (https://phabricator.wikimedia.org/T431307) [16:21:01] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343587 (https://phabricator.wikimedia.org/T431307) (owner: 10Andrew Bogott) [16:21:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [16:22:10] !log jclark@cumin1004 START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [16:23:13] (03PS3) 10Andrew Bogott: Nova policy.yaml: disable instance pause and suspend [puppet] - 10https://gerrit.wikimedia.org/r/1343587 (https://phabricator.wikimedia.org/T431307) [16:23:36] !log jclark@cumin1004 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [16:24:28] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343587 (https://phabricator.wikimedia.org/T431307) (owner: 10Andrew Bogott) [16:27:06] !log jclark@cumin1004 START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [16:27:30] !log jclark@cumin1004 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [16:27:59] (03CR) 10Hnowlan: [C:03+2] hadoop: add hdfs alert for HA status [alerts] - 10https://gerrit.wikimedia.org/r/1304769 (https://phabricator.wikimedia.org/T407138) (owner: 10Hnowlan) [16:28:21] (03CR) 10Brouberol: [C:03+1] kafka: converge topic config from hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [16:30:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:30:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1020.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:31:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:31:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:32:08] !log jclark@cumin1004 START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [16:34:05] (03CR) 10Giuseppe Lavagetto: external_clouds_vendors: stop generating the datafile (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [16:34:09] (03Merged) 10jenkins-bot: hadoop: add hdfs alert for HA status [alerts] - 10https://gerrit.wikimedia.org/r/1304769 (https://phabricator.wikimedia.org/T407138) (owner: 10Hnowlan) [16:35:14] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs2013.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:35:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:35:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1019.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:36:14] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:36:25] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie [16:37:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:37:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:40:14] PROBLEM - PyBal backends health check on lvs2014 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs2013.codfw.wmnet, wdqs2007.codfw.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:40:54] (03PS1) 10Ebernhardson: opensearch semantic test: Update to opensearch 3.8.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343599 (https://phabricator.wikimedia.org/T438058) [16:41:14] RECOVERY - PyBal backends health check on lvs2014 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:41:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1014.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:41:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:42:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:42:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [16:43:56] (03CR) 10Ebernhardson: [C:03+2] opensearch semantic test: Update to opensearch 3.8.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343599 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [16:44:08] !log tappof@cumin1004 DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006 [16:44:15] !log jclark@cumin1004 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [16:46:24] (03Merged) 10jenkins-bot: opensearch semantic test: Update to opensearch 3.8.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343599 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [16:46:42] !log cdobbins@cumin1004 conftool action : set/pooled=yes; selector: name=ncredir6002.* [16:48:31] (03PS6) 10Tiziano Fogli: kafka-logging1006: disable icinga notifications [puppet] - 10https://gerrit.wikimedia.org/r/1342257 (https://phabricator.wikimedia.org/T432444) [16:48:31] (03PS10) 10Tiziano Fogli: kafka-logging: remove kafka-logging1003 [puppet] - 10https://gerrit.wikimedia.org/r/1329299 (https://phabricator.wikimedia.org/T432444) [16:48:31] (03PS11) 10Tiziano Fogli: kafka-logging: bring up kafka-logging1006 with node id 1006 [puppet] - 10https://gerrit.wikimedia.org/r/1329302 (https://phabricator.wikimedia.org/T432444) [16:48:32] (03PS5) 10Tiziano Fogli: kafka-logging100[78]: disable icinga notifications [puppet] - 10https://gerrit.wikimedia.org/r/1342553 (https://phabricator.wikimedia.org/T432444) [16:48:33] (03PS15) 10Tiziano Fogli: kafka-logging: add kafka-logging100[7-8] to eqiad cluster [puppet] - 10https://gerrit.wikimedia.org/r/1327496 (https://phabricator.wikimedia.org/T432444) [16:48:34] (03PS3) 10Tiziano Fogli: kafka-logging: remove kafka-logging100[12] [puppet] - 10https://gerrit.wikimedia.org/r/1342557 (https://phabricator.wikimedia.org/T432444) [16:49:41] (03CR) 10Tiziano Fogli: [C:03+2] kafka-logging1006: disable icinga notifications [puppet] - 10https://gerrit.wikimedia.org/r/1342257 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [16:54:03] !log cdobbins@cumin1004 START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie [16:56:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [16:56:47] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [16:56:53] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [16:57:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:00:04] 06SRE, 06Infrastructure-Foundations, 06Release-Engineering-Team (Radar): apt-get broken in docker-registry.wikimedia.org/bullseye:20260830 - https://phabricator.wikimedia.org/T437069#12346546 (10dancy) Thanks @MoritzMuehlenhoff ! [17:00:05] Deploy window MediaWiki infrastructure (UTC late) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1700) [17:00:05] ryankemper: I, the Bot under the Fountain, call upon thee, The Deployer, to do Wikidata Query Service weekly deploy deploy. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T1700). [17:00:12] !log jclark@cumin1004 START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm [17:00:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:00:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:00:20] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12346547 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jclark@cumin1004 for host sessionstore1005.eqiad.wmnet with OS bookworm [17:01:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:01:25] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12346558 (10MatthewVernon) FTR: system is refusing to boot after the h/w swap, so trying a reimage. [17:02:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:04:04] PROBLEM - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is CRITICAL: CRITICAL: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:04:07] (03PS4) 10Andrew Bogott: Nova policy.yaml: disable instance pause, lock and suspend [puppet] - 10https://gerrit.wikimedia.org/r/1343587 (https://phabricator.wikimedia.org/T431307) [17:04:53] (03PS1) 10CDobbins: sretest2013: remove manual references to this host [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) [17:05:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:05:15] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:05:47] (03CR) 10Tiziano Fogli: [C:03+2] kafka-logging: remove kafka-logging1003 [puppet] - 10https://gerrit.wikimedia.org/r/1329299 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [17:05:53] (03CR) 10Tiziano Fogli: [C:03+2] kafka-logging: bring up kafka-logging1006 with node id 1006 [puppet] - 10https://gerrit.wikimedia.org/r/1329302 (https://phabricator.wikimedia.org/T432444) (owner: 10Tiziano Fogli) [17:06:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:06:15] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:06:23] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [17:07:43] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-09-18 - 2026-10-09), 07Essential-Work: Q1:rack/setup/install cirrussearch11[26-30] - https://phabricator.wikimedia.org/T436285#12346592 (10RobH) [17:09:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:10:15] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:12:15] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:12:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:12:43] !log jclark@cumin1004 START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [17:13:03] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=logging-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [17:14:36] KafkaUnderReplicatedPartitions is expected due to host swap [17:14:36] ^^ it's me [17:16:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:16:15] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:17:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:17:15] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:17:55] 06SRE, 10SRE-Access-Requests: Requesting access to nda for SEgt-WMF - https://phabricator.wikimedia.org/T438767 (10SEgt-WMF) 03NEW [17:18:52] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users (Level 3) for SEgt-WMF - https://phabricator.wikimedia.org/T438767#12346701 (10Dzahn) [17:23:26] FIRING: [5x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [17:25:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:26:01] !log jclark@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage [17:26:15] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:27:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:27:15] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:27:53] jclark@cumin1004 provision (PID 4060469) is awaiting input [17:29:55] !log jclark@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage [17:30:00] !log jclark@cumin1004 END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [17:33:43] (03CR) 10Ssingh: "For adding sretest2013, see references to another cp host such as cp2044 and check all the places in the Puppet repo on where that needs t" [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) (owner: 10CDobbins) [17:34:13] (03CR) 10Ssingh: "My bad, for adding *cp2059*, see references to another cp host..." [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) (owner: 10CDobbins) [17:34:54] RECOVERY - Check whether ferm is active by checking the default input chain on wikikube-worker1071 is OK: OK ferm input default policy is set https://wikitech.wikimedia.org/wiki/Monitoring/check_ferm [17:34:55] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [17:36:22] 06SRE, 10SRE-Access-Requests: Requesting access to analytics-privatedata-users (Level 3) for SEgt-WMF - https://phabricator.wikimedia.org/T438767#12346814 (10FRomeo_WMF) I approve this request [17:36:39] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12346815 (10MatthewVernon) [reimage paused at the "wait, where's the UEFI partition" point, so we re-provisioned to legacy-boot and tried a --no-pxe reimage, which booted the... [17:40:03] !log jclark@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm [17:40:10] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12346822 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jclark@cumin1004 for host sessionstore1005.eqiad.wmnet with OS bookworm completed: - sessionsto... [17:43:26] FIRING: [4x] ProbeDown: Service registry1004:5001 has failed probes (http_docker_registry_health_ip4) - https://wikitech.wikimedia.org/wiki/Docker - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [17:44:02] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12346850 (10Jclark-ctr) I was able to replace the parts. I’m also having issues uploading the license. waiting on response from Dell [17:44:58] 10ops-eqiad, 06SRE, 06DC-Ops: Degraded RAID on sessionstore1005 - https://phabricator.wikimedia.org/T438583#12346853 (10Jclark-ctr) 05Open→03Resolved [17:48:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [17:49:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:50:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [17:50:40] !log cdobbins@cumin1004 START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage [17:51:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:52:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [17:53:24] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:54:04] RECOVERY - Check unit status of httpbb_kubernetes_mw-api-ext-next_hourly on cumin2003 is OK: OK: Status of the systemd unit httpbb_kubernetes_mw-api-ext-next_hourly https://wikitech.wikimedia.org/wiki/Monitoring/systemd_unit_state [17:54:22] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage [17:54:43] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12346915 (10RoySmith) I've been seeing this a lot on [[:en:Special:Checkuser]]. Chatting with some other CUs on Discord, I estimated I'm getting it on about 50% of the checks I run.... [17:55:03] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12346917 (10MatthewVernon) We've swopped the disk; I blanked the replacment, and have endeavoured to bring it into service thus: ` sudo sfdisk -d /dev/sdc >/tmp/partitiontable... [18:04:26] (03PS1) 10Michael Große: Growth: Set minimum registration date for A/B test [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343611 (https://phabricator.wikimedia.org/T432126) [18:06:37] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b4-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [18:06:50] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12346967 (10RoySmith) There's also a thread currently at https://en.wikipedia.org/wiki/Wikipedia:Village_pump_(technical)#Loss_of_session_data which suggests this is a widespread pro... [18:13:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:16:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [18:17:41] (03PS1) 10Ebernhardson: opensearch semantic test: Drop config for insights plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343612 [18:18:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [18:19:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [18:20:10] (03CR) 10Ebernhardson: [C:03+2] opensearch semantic test: Drop config for insights plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343612 (owner: 10Ebernhardson) [18:22:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [18:22:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [18:22:21] (03CR) 10Btullis: "Agreed." [puppet] - 10https://gerrit.wikimedia.org/r/1343023 (https://phabricator.wikimedia.org/T429773) (owner: 10Btullis) [18:22:43] (03Merged) 10jenkins-bot: opensearch semantic test: Drop config for insights plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343612 (owner: 10Ebernhardson) [18:23:44] (03CR) 10CDobbins: sretest2013: remove manual references to this host (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) (owner: 10CDobbins) [18:24:28] !log cdobbins@cumin1004 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie [18:24:30] (03PS2) 10CDobbins: sretest2013: remove manual references to this host [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) [18:25:42] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [18:25:47] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [18:26:17] (03CR) 10Fabfur: external_clouds_vendors: stop generating the datafile (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1343537 (https://phabricator.wikimedia.org/T438658) (owner: 10Giuseppe Lavagetto) [18:26:25] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [18:27:23] !log mvernon@cumin1004 START - Cookbook sre.discovery.service-route check sessionstore: maintenance [18:27:23] !log mvernon@cumin1004 END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance [18:30:16] !log mvernon@cumin1004 START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired [18:30:37] !log repool eqiad sessionstore T437915 [18:30:39] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [18:30:40] T437915: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915 [18:32:07] !log cdobbins@cumin1004 conftool action : set/pooled=yes; selector: name=ncredir5003.* [18:35:19] !log mvernon@cumin1004 END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired [18:37:45] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12347156 (10asilvering) I've been having this problem on some 1/5 of my reasonably substantial edits. I have not yet had it on a talk page reply using the reply tool. I've also been... [18:40:50] (03PS1) 10Gmodena: wdqs: add wikidata prefixes config [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343616 (https://phabricator.wikimedia.org/T438476) [18:41:19] (03CR) 10Clare Ming: Test Kitchen UI: Deploy v2.0.0 release to staging (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343552 (https://phabricator.wikimedia.org/T421814) (owner: 10Santiago Faci) [18:45:08] (03CR) 10Catrope: "Why can't this use MW core's copy of vue-router?" [extensions/PersonalDashboard] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1343126 (https://phabricator.wikimedia.org/T438387) (owner: 10Zabe) [18:46:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [18:46:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [18:46:59] 06SRE, 10LDAP-Access-Requests: Grant Access to ldap/wmf for "Nick Battista" - https://phabricator.wikimedia.org/T438782 (10NBattista-WMF) 03NEW [18:47:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [18:47:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [18:49:40] (03PS1) 10Ebernhardson: opensearch semantic test: Enable the opensearch security plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343619 [18:50:08] 10ops-eqiad, 06SRE, 10Cassandra, 06DC-Ops: sessionstore1005.eqiad.wmnet is down (again) - https://phabricator.wikimedia.org/T437915#12347179 (10MatthewVernon) p:05Unbreak!→03High [downgrading to High for now, since it looks to be handling the load OK having been pooled] [18:50:08] (03PS2) 10Ebernhardson: opensearch semantic test: Enable the opensearch security plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343619 [18:52:10] (03PS2) 10Sbisson: Keep Article Guidance on where it is on today [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342285 (https://phabricator.wikimedia.org/T433293) [18:52:56] (03CR) 10ScheduleDeploymentBot: "Scheduled for deployment in the [Monday, September 21 UTC late backport window](https://wikitech.wikimedia.org/wiki/Deployments#deploycal-" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1342285 (https://phabricator.wikimedia.org/T433293) (owner: 10Sbisson) [18:53:36] (03PS2) 10Santiago Faci: Test Kitchen UI: Deploy v2.0.0 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343552 (https://phabricator.wikimedia.org/T421814) [18:54:30] (03CR) 10Santiago Faci: Test Kitchen UI: Deploy v2.0.0 release to staging (031 comment) [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343552 (https://phabricator.wikimedia.org/T421814) (owner: 10Santiago Faci) [18:55:35] (03CR) 10Ebernhardson: [C:03+2] opensearch semantic test: Enable the opensearch security plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343619 (owner: 10Ebernhardson) [18:55:38] (03CR) 10Clare Ming: [C:03+2] Test Kitchen UI: Deploy v2.0.0 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343552 (https://phabricator.wikimedia.org/T421814) (owner: 10Santiago Faci) [18:57:52] (03Merged) 10jenkins-bot: opensearch semantic test: Enable the opensearch security plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343619 (owner: 10Ebernhardson) [18:58:12] (03Merged) 10jenkins-bot: Test Kitchen UI: Deploy v2.0.0 release to staging [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343552 (https://phabricator.wikimedia.org/T421814) (owner: 10Santiago Faci) [18:59:23] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [18:59:27] !log ebernhardson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply [19:00:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [19:00:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [19:02:03] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply [19:02:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:02:29] !log sfaci@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply [19:03:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:05:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [19:06:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [19:07:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:07:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:10:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [19:10:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [19:11:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:11:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:11:24] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12347232 (10Vanamonde93) I complained about this on discord on 19 September: since then it's affected, I would estimate, 25% of my actions, including CU actions. I cannot remember th... [19:23:14] (03PS3) 10CDobbins: sretest2013: remove manual references to this host [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) [19:25:05] (03PS2) 10Dzahn: site: add zuul1005 as a zuul executor [puppet] - 10https://gerrit.wikimedia.org/r/1319547 (https://phabricator.wikimedia.org/T427353) [19:25:36] (03CR) 10Dzahn: [C:03+2] "after a meeting with Tyler - we clarified the executor should be physical" [puppet] - 10https://gerrit.wikimedia.org/r/1319547 (https://phabricator.wikimedia.org/T427353) (owner: 10Dzahn) [19:26:23] (03CR) 10Ssingh: "Patch looks good, thanks. Before I +1, can you check with DC-Ops on the precursor steps to this? Essentially, we are changing the hostname" [puppet] - 10https://gerrit.wikimedia.org/r/1343602 (https://phabricator.wikimedia.org/T436691) (owner: 10CDobbins) [19:26:51] FIRING: [3x] ATSBackendErrorsHigh: ATS: elevated 5xx error ratio from eventgate-logging-external.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [19:27:07] 👋 [19:27:11] o/ [19:27:59] tappof and I have a kafka-logging eqiad maintenance in flight, possibly external services needs a refresh? [19:28:36] refresh in what sense? [19:29:25] it's hitting more than one DC, so yeah, seems eventgate is the issue [19:30:04] eventgate-logging-external, I mean :) [19:30:42] I do see this diff for admin_ng in eqiad, herron is that what you mean? https://www.irccloud.com/pastebin/9UITJtUp/ [19:31:00] and a matching one in codfw [19:31:10] rzl yes exactly [19:31:11] yes rzl [19:31:41] eventgate should hit the remaining 4 brokers no problem though so this is unexpected, but most likely thats the issue [19:32:21] sounds like a good documentation TODO for the maintenance you're working on :) do you want to deploy that diff then? [19:32:44] or I can, if you'd rather, but I don't have any context so I'm trusting your review (e.g. is it correct to delete the old IP and add the new one at the same time) [19:32:51] I think moreso we need to figure out why that's acting as a spof, its in the checklist :( [19:32:55] yes please do if you're already there! [19:33:30] do you happen to have context on the kerberos-kdc diff too? otherwise I need to find out if that's safe to roll out [19:33:43] negative [19:33:45] no rzl [19:33:54] okay, let me see what I can find [19:34:11] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12347291 (10FlammablePizza) I have also had this error appearing occasionally over the last few days. Never had it happen before this. [19:34:21] (03PS2) 10Btullis: admin_ng: pin the ceph-csi charts to the versions in production [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343130 (https://phabricator.wikimedia.org/T407166) [19:34:22] (03PS4) 10Btullis: ceph-csi-rbd: rebase the chart on upstream v3.14.2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343123 (https://phabricator.wikimedia.org/T407166) [19:34:22] (03PS4) 10Btullis: ceph-csi-cephfs: rebase the chart on upstream v3.14.2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343124 (https://phabricator.wikimedia.org/T407166) [19:34:22] (03PS1) 10Btullis: admin_ng: move dse-k8s-codfw to ceph-csi-rbd 0.2.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343624 (https://phabricator.wikimedia.org/T407166) [19:34:24] (03PS1) 10Btullis: admin_ng: move dse-k8s-codfw to ceph-csi-cephfs 0.2.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343625 (https://phabricator.wikimedia.org/T407166) [19:36:37] > sounds like a good documentation TODO for the maintenance you're working on :) [19:36:43] Yeah, it's a step on the checklist, but the previous one took longer than expected due to the Kafka ACLs and some mess with the firewall used by the insetup role and the production one. Anyway thanks rzl! [19:38:13] got it -- sorry both, I didn't mean for that to sound smug <3 [19:38:41] No, you didn't sound smug... I was just making conversation :) [19:38:53] still chasing the other diff, the IP address points to krb1002 and it looks like it's been worked on lately https://phabricator.wikimedia.org/T435354 [19:39:14] I just want to make sure it wasn't recently *removed* in which case I'd be adding it back [19:41:31] half-considering a targeted local edit to just update the kafka-logging-eqiad part and leave the kerberos-kdc part dirty, which would get us out of this situation faster, but it's distasteful and a little error-prone [19:41:42] I'm curious why eventgate isn't falling back to another broker, it should be aware of 5. we took 1003 out, added 1006, but the cluster as a whole is ok [19:43:54] hmm yeah, in theory the diff reflects the current state and just hasn't been applied yet...famous last words [19:45:19] (03CR) 10CI reject: [V:04-1] admin_ng: move dse-k8s-codfw to ceph-csi-rbd 0.2.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343624 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [19:45:44] (03CR) 10CI reject: [V:04-1] admin_ng: move dse-k8s-codfw to ceph-csi-cephfs 0.2.0 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343625 (https://phabricator.wikimedia.org/T407166) (owner: 10Btullis) [19:47:04] https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/production/modules/profile/manifests/kubernetes/deployment_server/global_config.pp#252 [19:47:29] right okay, I didn't realize that's how we populate kerberos in external_services. neat [19:48:53] so it's directly a puppetdb query on the host's role, so https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329636 was the change that triggered the diff [19:49:44] nice find [19:50:13] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12347412 (10Daimona) [19:50:29] inflatador: if you're around, I'm going to deploy a latent change adding krb1002 to external_services, yell if you know of a reason I shouldn't do that please :) [19:51:14] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [19:51:14] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1012.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [19:51:24] (or, not that patch, that patch added krb1004. same line though. one sec) [19:51:51] RESOLVED: [3x] ATSBackendErrorsHigh: ATS: elevated 5xx error ratio from eventgate-logging-external.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [19:52:14] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:52:14] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [19:53:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:53:48] okay, the more I dig (https://gerrit.wikimedia.org/r/1328688, https://phabricator.wikimedia.org/T435854) the more it looks like krb1002 is removed from external-services on purpose, and I don't want to re-add it without a positive ack from someone who knows what's going on there [19:55:31] the correct way to resolve that diff is to either apply it if appropriate, or remove the puppet role from the host if appropriate, or potentially add some stuff to https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/production/modules/profile/manifests/kubernetes/deployment_server/global_config.pp#252 to just exclude it from that wmflib::role::ips() call? [19:56:11] rzl: and meanwhile it seems the eventgate issue resolved? [19:56:28] (we were just recently in serviceops talking about the danger of unresolved admin_ng diffs, I'll take this story as a useful example) [19:56:42] yeah apparently. I also don't know why it only paged in drmrs [19:57:21] herron, tappof: what's the state on your end? [19:58:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [19:58:54] tappof pointed out the alert clearing seemed to correlate with a broker restart on 1006 (we needed to add disk) curious if the eventgate alert will come back [20:00:05] RoanKattouw, urbanecm, TheresNoTime, kindrobot, and cjming: May I have your attention please! UTC late backport window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T2000) [20:00:05] Ameisenigel and stephanebisson: A patch you scheduled for UTC late backport window is about to be deployed. Please be around during the process. Note: If you break AND fix the wikis, you will be rewarded with a sticker. [20:03:39] okay -- if it's not currently urgent I'll go hands-off for some other work, but I'll see if we can get an answer on what to do with that krb1002 diff in external-services -- if not inflatador, maybe ryankemper happens to be online? [20:04:08] thanks rzl [20:04:22] souns good rzl thanks, if it pages again we'll erm page you :) [20:04:29] it also occurred to me, just speculation, if the host was offline for two weeks it would have dropped out of puppetdb and so also from external-services -- the change that precipitated the diff might have just been the machine turning back on! [20:04:35] Ready for backport, if any dev is around [20:05:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [20:05:19] if that's true, *and* if it's not desired, then that's an argument against using wmflib::role::ips for that sort of thing [20:06:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [20:06:24] (03PS1) 10Ebernhardson: opensearch semantic test: Update image to 3.8.0-2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343633 (https://phabricator.wikimedia.org/T438058) [20:11:28] Ameisenigel: i can deploy for you - 1 sec [20:11:52] (03CR) 10Ebernhardson: [C:03+2] opensearch semantic test: Update image to 3.8.0-2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343633 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [20:12:34] (03PS2) 10Ameisenigel: Disable wgMFCustomSiteModules on German Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343100 (https://phabricator.wikimedia.org/T403380) [20:13:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [20:14:14] (03Merged) 10jenkins-bot: opensearch semantic test: Update image to 3.8.0-2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343633 (https://phabricator.wikimedia.org/T438058) (owner: 10Ebernhardson) [20:15:15] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [20:16:15] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [20:17:13] FIRING: JobUnavailable: Reduced availability for job redis_arclamp in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [20:18:11] (03CR) 10TrainBranchBot: [C:03+2] "Approved by cjming@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343100 (https://phabricator.wikimedia.org/T403380) (owner: 10Ameisenigel) [20:19:21] (03Merged) 10jenkins-bot: Disable wgMFCustomSiteModules on German Wikipedia [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343100 (https://phabricator.wikimedia.org/T403380) (owner: 10Ameisenigel) [20:19:36] !log cjming@deploy1003 Started scap sync-world: Backport for [[gerrit:1343100|Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] [20:19:39] T403380: Set wgMFCustomSiteModules to false for dewiki - https://phabricator.wikimedia.org/T403380 [20:21:47] FIRING: RedisInstanceDown: Redis instance down arclamp2001:9121 redis_arclamp - https://wikitech.wikimedia.org/wiki/Redis#Cluster_redis_arclamp - https://grafana.wikimedia.org/d/000000174/redis?orgId=1&var-site=codfw&var-job=redis_arclamp&var-instance=arclamp2001:9121 - https://alerts.wikimedia.org/?q=alertname%3DRedisInstanceDown [20:23:36] !log cjming@deploy1003 ameisenigel, cjming: Backport for [[gerrit:1343100|Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [20:24:22] Ameisenigel: on test servers - ok to sync? [20:25:26] !log ihurbain@deploy1003 helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [20:25:55] !log ihurbain@deploy1003 helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [20:25:56] !log ihurbain@deploy1003 helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [20:26:27] !log ihurbain@deploy1003 helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [20:30:22] Yes, ready for sync [20:30:37] !log cjming@deploy1003 ameisenigel, cjming: Continuing with deployment [20:35:32] !log cjming@deploy1003 Finished scap sync-world: Backport for [[gerrit:1343100|Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s) [20:35:35] T403380: Set wgMFCustomSiteModules to false for dewiki - https://phabricator.wikimedia.org/T403380 [20:35:51] Ameisenigel: should be live! [20:36:25] stephanebisson: are you around? do you need a deployer or prefer to self-service? [20:37:14] Thanks! [20:37:20] yw! [20:58:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:00:05] alexsanford, Reedy, sbassett, Maryum, and manfredi: Time to snap out of that daydream and deploy Weekly Security deployment window. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T2100). [21:03:01] hi preparing to deploy a security patch [21:03:09] looks like the other deploys have wrapped up [21:06:38] FIRING: [8x] CertAlmostExpired: gNMI TLS certificate for lsw1-c3-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [21:13:18] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=logging-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [21:14:10] (03PS1) 10Lerickson: Remove a default value that never took effect anyway. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343643 (https://phabricator.wikimedia.org/T438789) [21:16:51] FIRING: ATSBackendErrorsHigh: ATS: elevated 5xx error ratio from eventgate-logging-external.discovery.wmnet in eqiad #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://grafana.wikimedia.org/d/1T_4O08Wk/ats-backends-origin-servers-overview?orgId=1&viewPanel=12&var-site=eqiad&var-cluster=text&var-origin=eventgate-logging-external.discovery.wmnet - https://alerts.wikimedia.org/?q=alertname%3DATSBackendError [21:17:22] hmmm, [21:17:58] hmmm hopefully not related to the current security deploy [21:18:17] maryum: unlikely, it came up earlier too [21:18:23] okay phew [21:18:37] !log Deployed security fix for T437708 [21:18:38] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:18:43] if you're patching anything eventgate-related I'd love to know about it :) otherwise I think you're in the clear but let me know before you start anything new [21:19:28] rzl so we could stop the new broker or update external services [21:20:28] whats your preference? stopping the broker puts kafka in a degraded state but fine for short term ~24h [21:20:29] how disruptive is it to stop the new broker? [21:20:35] ah jinx, got ti [21:20:41] haha [21:21:07] I'm leaning that way but I just pinged dpe sre to see if we can get their eyes on the krb diff, let me see if that works [21:21:51] FIRING: [3x] ATSBackendErrorsHigh: ATS: elevated 5xx error ratio from eventgate-logging-external.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [21:22:11] ok [21:22:32] herron: https://grafana.wikimedia.org/goto/sb5rss?orgId=default I'm curious, do the start and stop times of the 5xxs mean anything to you? [21:22:48] very step-function on and off [21:23:43] ATS gets 5xxs when the broker is running rzl [21:24:10] yeah I think the 1006 took on some partitions this writes to but cant reach, so if we stop the broker it will fall back to a reachable host [21:24:17] probably because it gets elected as the leader for some partitions. [21:24:42] ah got it, I was trying to figure out why it was clean the rest of the time if it hadn't been stopped, but leader election makes sense to me [21:25:24] because kafka-logging1006 is a new host that's replacing kafka-logging1003 [21:25:40] We're putting it into "production" [21:26:21] okay inflatador is taking a look -- if it's easy to handle the krb1002, we'll go ahead and do that [21:26:34] if you have any prep work before stopping the broker, you may want to get ready to, just in case we end up going that route [21:27:14] honestly, if it's not a huge deal, maybe consider doing both just to mitigate the impact? stop the broker now, and if resolving krb1002 is quick, we can just start it again. thoughts? [21:29:01] (03PS1) 10Bking: Kerberos: remove krb1002 from Kerberos role [puppet] - 10https://gerrit.wikimedia.org/r/1343645 (https://phabricator.wikimedia.org/T438229) [21:29:57] (03CR) 10Bking: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1343645 (https://phabricator.wikimedia.org/T438229) (owner: 10Bking) [21:30:54] herron, I don't think it would be a problem to stop the broker for a while. We're replicating the last partition .. Do you have any thoughts? [21:32:14] tappof yeah sounds fine to me as well [21:32:42] ack herron just stopped rzl [21:33:07] (03CR) 10RLazarus: [C:03+1] "Not a Kerberos knower, but the syntax etc LGTM :)" [puppet] - 10https://gerrit.wikimedia.org/r/1343645 (https://phabricator.wikimedia.org/T438229) (owner: 10Bking) [21:33:37] (03CR) 10Bking: [C:03+2] Kerberos: remove krb1002 from Kerberos role [puppet] - 10https://gerrit.wikimedia.org/r/1343645 (https://phabricator.wikimedia.org/T438229) (owner: 10Bking) [21:34:38] herron tappof: okay cool! ^ that should be all we needed to resolve the other diff, but once we run puppet on the deploy host I'll know for sure [21:34:55] then I'll let you know once external-services is applied and you'll be in business [21:35:10] FIRING: CoreRouterInterfaceDown: Core router interface down - cr2-magru:xe-0/1/3 (reserved for moved telxius transport to eqiad) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr2-magru:9804 - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [21:35:28] ack rzl and thanks [21:35:57] running now, takes a while [21:36:51] RESOLVED: [3x] ATSBackendErrorsHigh: ATS: elevated 5xx error ratio from eventgate-logging-external.discovery.wmnet in drmrs #page - https://wikitech.wikimedia.org/wiki/Apache_Traffic_Server#Debugging - https://alerts.wikimedia.org/?q=alertname%3DATSBackendErrorsHigh [21:37:37] 10ops-eqiad, 06SRE, 06DC-Ops, 13Patch-For-Review: krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12347797 (10bking) Hey all, just a heads-up that we are going to replace this host with a VM, so DPE SRE no longer needs it for Kerberos. If y'all want to reimage or make any othe... [21:38:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:43:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:43:47] no, right, running puppet on the deployment host didn't do it because puppetdb hadn't picked up the change yet [21:43:53] sigh [21:44:56] trying to remember where puppet needs to run first, if it's just the puppetservers or not [21:46:04] probably puppetdb*. or after 30 minutes I can give up trying to remember, lol [21:46:11] let's try one and then the other [21:48:03] RESOLVED: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=logging-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [21:53:26] FIRING: [6x] ProbeDown: Service kafka-logging1006:9093 has failed probes (tcp_kafka_broker_tls_kafka_logging1006_ip4) - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [21:53:39] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [21:56:08] (03CR) 10Jforrester: "Because your team has vetoed that ahead of the Security review in T436156, AIUI." [extensions/PersonalDashboard] (wmf/1.47.0-wmf.20) - 10https://gerrit.wikimedia.org/r/1343126 (https://phabricator.wikimedia.org/T438387) (owner: 10Zabe) [21:59:03] FIRING: KafkaUnderReplicatedPartitions: Under replicated partitions for Kafka cluster logging-eqiad in eqiad - https://wikitech.wikimedia.org/wiki/Kafka/Administration - https://grafana.wikimedia.org/d/000000027/kafka?orgId=1&var-datasource=eqiad%20prometheus/ops&var-kafka_cluster=logging-eqiad - https://alerts.wikimedia.org/?q=alertname%3DKafkaUnderReplicatedPartitions [22:02:13] FIRING: [2x] JobUnavailable: Reduced availability for job atlas_exporter in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [22:04:51] okay, the kerberos-kdc IP diff is still present, and I'll follow up and learn what I'm missing there about puppetdb. in the meantime b.tullis says the diff is also safe to apply, so I'm just going to go ahead and do that and we can figure the rest out later [22:05:28] (and all this goes into serviceops's ongoing discussion of latent external-services diffs) [22:05:43] !log rzl@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [22:06:37] FIRING: [2x] GnmiInterfaceCountersDrop: asw1-b4-magru is exporting less than half the gNMI interface counters it had 24h ago - https://wikitech.wikimedia.org/wiki/Network_monitoring#GnmiInterfaceCountersDrop - https://alerts.wikimedia.org/?q=alertname%3DGnmiInterfaceCountersDrop [22:06:43] !log rzl@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [22:07:43] !log rzl@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [22:08:21] !log rzl@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [22:08:57] !log rzl@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [22:09:55] !log rzl@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [22:10:46] !log rzl@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [22:11:25] !log rzl@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [22:11:39] herron (and t.appof, if you're still here this late): external-services applied [22:12:04] awesome! thanks rzl [22:13:54] for wikikube anyway. other clusters liable to still have those diffs and more -- do you know which ones you need? [22:15:01] if it's just eventgate-logging-external then it's just wikikube and you're all set [22:17:30] (03PS1) 10Dzahn: zuul: update hieradata to reflect main node names [puppet] - 10https://gerrit.wikimedia.org/r/1343649 [22:21:50] (03CR) 10Dzahn: [C:03+2] add discovery record for codesearch [dns] - 10https://gerrit.wikimedia.org/r/1342805 (https://phabricator.wikimedia.org/T268199) (owner: 10Dzahn) [22:22:14] (03PS3) 10Dzahn: add discovery record for codesearch [dns] - 10https://gerrit.wikimedia.org/r/1342805 (https://phabricator.wikimedia.org/T268199) [22:22:30] (03PS2) 10Btullis: admin_ng: move dse-k8s-codfw to ceph-csi-rbd v3.14.2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343624 (https://phabricator.wikimedia.org/T407166) [22:22:34] (03PS2) 10Btullis: admin_ng: move dse-k8s-codfw to ceph-csi-cephfs v3.14.2 [deployment-charts] - 10https://gerrit.wikimedia.org/r/1343625 (https://phabricator.wikimedia.org/T407166) [22:24:18] (03PS1) 10Btullis: dse-k8s: Delegate reverse DNS for the second eqiad Pod range [dns] - 10https://gerrit.wikimedia.org/r/1343650 (https://phabricator.wikimedia.org/T430658) [22:26:41] FIRING: [2x] SystemdUnitFailed: prometheus-node-textfile-export_service_type.service on cumin1003:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [22:31:49] (03CR) 10Dzahn: "can we please have this back" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337931 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [22:33:44] (03CR) 10Dzahn: "comment was in reference to https://phabricator.wikimedia.org/T438777 (and https://phabricator.wikimedia.org/T268199)" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1337931 (https://phabricator.wikimedia.org/T416452) (owner: 10Muehlenhoff) [22:38:12] (03CR) 10Jdlrobson: [C:03+1] "please backport when you can" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1343560 (https://phabricator.wikimedia.org/T437339) (owner: 10VolkerE) [22:47:16] 06SRE, 10MediaWiki-Page-editing: increase in session_fail_preview messages on enwiki - https://phabricator.wikimedia.org/T423206#12347993 (10Hamishcn) Ditto on meta. [22:51:32] (03CR) 10Dzahn: "recheck" [dns] - 10https://gerrit.wikimedia.org/r/1342805 (https://phabricator.wikimedia.org/T268199) (owner: 10Dzahn) [22:57:02] (03CR) 10RLazarus: [C:03+2] hieradata: Remove profile::swift::proxy::private_container_list [puppet] - 10https://gerrit.wikimedia.org/r/1343143 (https://phabricator.wikimedia.org/T334488) (owner: 10RLazarus) [22:59:45] (03CR) 10RLazarus: [C:03+1] "Today was a little bit of a mess, so I didn't get to this, apologies. Tomorrow of course is also the services component of the DC switchov" [puppet] - 10https://gerrit.wikimedia.org/r/1339694 (https://phabricator.wikimedia.org/T437403) (owner: 10Zabe) [23:00:04] Deploy window Readers deployment window (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260921T2300) [23:40:57] (03PS1) 10TrainBranchBot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1343657 [23:40:57] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1343657 (owner: 10TrainBranchBot) [23:49:47] (03Merged) 10jenkins-bot: Branch commit for wmf/branch_cut_pretest [core] (wmf/branch_cut_pretest) - 10https://gerrit.wikimedia.org/r/1343657 (owner: 10TrainBranchBot)