[00:06:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:08:31] (03PS2) 10Ladsgroup: Revert^2 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330710 [00:09:29] (03CR) 10Scott French: [C:03+1] "Thanks, Bryan! We'll get this merged when we're ready to go on this step." [puppet] - 10https://gerrit.wikimedia.org/r/1329690 (https://phabricator.wikimedia.org/T436178) (owner: 10BryanDavis) [00:15:11] (03CR) 10Ladsgroup: "This is the value of file type extensions allowed in our production:" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330710 (owner: 10Ladsgroup) [00:17:32] (03CR) 10Ladsgroup: [C:03+2] Revert^2 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330710 (owner: 10Ladsgroup) [00:20:01] (03CR) 10Scott French: [C:03+1] "Thanks, Luca - sounds reasonable to me!" [puppet] - 10https://gerrit.wikimedia.org/r/1330254 (https://phabricator.wikimedia.org/T428022) (owner: 10Elukey) [00:20:03] (03Merged) 10jenkins-bot: Revert^2 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330710 (owner: 10Ladsgroup) [00:20:43] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv4 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940409 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [00:23:39] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [00:24:00] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [00:27:00] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [00:27:37] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [00:31:58] (03PS6) 10Scott French: api-gateway: Drop support for debug_hosts [deployment-charts] - 10https://gerrit.wikimedia.org/r/1305236 (https://phabricator.wikimedia.org/T433752) [00:31:58] (03PS5) 10Scott French: api-gateway: Remove stale test assertion and noop Lua code [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311874 (https://phabricator.wikimedia.org/T433752) [00:31:58] (03PS5) 10Scott French: api-gateway: Drop support for php_engine_routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311875 (https://phabricator.wikimedia.org/T433752) [00:31:59] (03PS4) 10Scott French: api-gateway: Basic cluster specifier support and Lua plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) [00:32:00] (03PS6) 10Scott French: api-gateway: Support x-wikimedia-debug routing [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311964 (https://phabricator.wikimedia.org/T433752) [00:32:01] (03PS3) 10Scott French: api-gateway: Support Host-based diversion in the mw-api plugin [deployment-charts] - 10https://gerrit.wikimedia.org/r/1322140 (https://phabricator.wikimedia.org/T433752) [00:35:26] (03PS1) 10Ladsgroup: Revert^3 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330726 [00:35:38] (03PS2) 10Ladsgroup: Revert^3 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330726 [00:35:43] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv4 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940409 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [00:36:07] (03PS3) 10Ladsgroup: Revert^3 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330726 [00:36:11] (03CR) 10Ladsgroup: [C:03+2] Revert^3 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330726 (owner: 10Ladsgroup) [00:36:58] (03CR) 10Scott French: "Thank you so much for the reviews!" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1311963 (https://phabricator.wikimedia.org/T433752) (owner: 10Scott French) [00:38:27] (03Merged) 10jenkins-bot: Revert^3 "[hardening] thumbor: Flip from allow-unless-denied to vice versa" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330726 (owner: 10Ladsgroup) [00:39:56] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [00:40:05] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [00:43:33] (03PS1) 10Ladsgroup: thumbor: Bump the chart version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330729 [00:44:39] FIRING: CoreBGPDown: Core BGP session down between cr1-magru and cr1-eqiad (195.200.68.136) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=magru&var-device=cr1-magru:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [00:45:12] (03CR) 10Ladsgroup: [C:03+2] thumbor: Bump the chart version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330729 (owner: 10Ladsgroup) [00:47:27] (03Merged) 10jenkins-bot: thumbor: Bump the chart version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330729 (owner: 10Ladsgroup) [00:49:39] RESOLVED: CoreBGPDown: Core BGP session down between cr1-magru and cr1-eqiad (195.200.68.136) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=magru&var-device=cr1-magru:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr1-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [00:52:12] !log ladsgroup@deploy1003 helmfile [staging] START helmfile.d/services/thumbor: apply [00:53:18] !log ladsgroup@deploy1003 helmfile [staging] DONE helmfile.d/services/thumbor: apply [00:53:28] !log ladsgroup@deploy1003 helmfile [codfw] START helmfile.d/services/thumbor: apply [00:53:56] !log ladsgroup@deploy1003 helmfile [codfw] DONE helmfile.d/services/thumbor: apply [00:54:06] !log ladsgroup@deploy1003 helmfile [eqiad] START helmfile.d/services/thumbor: apply [00:55:33] !log ladsgroup@deploy1003 helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [01:06:50] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [01:08:42] (03CR) 10RLazarus: [C:03+1] P:services_proxy::envoy: Drop support for split and introduce splits (032 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1328247 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [01:08:51] (03CR) 10RLazarus: [C:03+1] "Still a great catch, and glad to see the noop in PCC." [puppet] - 10https://gerrit.wikimedia.org/r/1330689 (https://phabricator.wikimedia.org/T427666) (owner: 10Scott French) [01:11:31] (03PS1) 10TrainBranchBot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1330755 [01:11:31] (03CR) 10TrainBranchBot: [C:03+2] Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1330755 (owner: 10TrainBranchBot) [01:16:47] 06SRE, 06ServiceOps, 10ServiceOps-Upgrades-Hardware: mc20[56-73] implementation tracking - https://phabricator.wikimedia.org/T436273#12263944 (10RLazarus) p:05Triage→03Medium [01:17:45] (03Merged) 10jenkins-bot: Branch commit for wmf/next [core] (wmf/next) - 10https://gerrit.wikimedia.org/r/1330755 (owner: 10TrainBranchBot) [01:37:53] (03PS3) 10Ryan Kemper: Enable presto join spill config [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [01:47:09] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [01:55:22] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [01:58:39] FIRING: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [01:58:54] FIRING: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:00:41] !log mwpresync@deploy1003 Started scap build-images: Publishing wmf/next image [02:08:34] !log mwpresync@deploy1003 Finished scap build-images: Publishing wmf/next image (duration: 07m 53s) [02:24:43] FIRING: [2x] RipeAtlasAnchorUnreachable: ipv4 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940409 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [02:24:47] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [02:25:47] !log `ryankemper@pcc-db1002:~$ sudo -u jenkins-deploy puppetdb-populate --host an-test-coord1002.eqiad.wmnet` (refreshed stale catalog) [02:25:48] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [02:26:42] (03CR) 10Subramanya Sastry: [C:03+1] Produnto: Add IPv6 range for GitLab [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329728 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [02:27:02] (03CR) 10Subramanya Sastry: [C:03+1] Grant produnto-update to groups that have editprotected [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328752 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [02:28:39] RESOLVED: CoreBGPDown: Core BGP session down between cr2-codfw and cr2-eqiad (208.80.154.216) - group Confed_eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#BGP_status - https://grafana.wikimedia.org/d/ed8da087-4bcb-407d-9596-d158b8145d45/bgp-neighbors-detail?orgId=1&var-site=codfw&var-device=cr2-codfw:9804&var-bgp_group=Confed_eqiad&var-bgp_neighbor=cr2-eqiad - https://alerts.wikimedia.org/?q=alertname%3DCoreBGPDown [02:28:54] RESOLVED: [2x] CoreRouterInterfaceDown: Core router interface down - cr2-codfw:et-0/1/4 (Transport: cr2-eqiad:et-1/1/5 (Lumen, 449169461)) - https://wikitech.wikimedia.org/wiki/Network_monitoring#Router_interface_down - https://alerts.wikimedia.org/?q=alertname%3DCoreRouterInterfaceDown [02:29:43] RESOLVED: [2x] RipeAtlasAnchorUnreachable: ipv4 ping to esams RIPE Atlas anchor: failures over threshold for measurement 59940409 - https://wikitech.wikimedia.org/wiki/Network_monitoring#Atlas_alerts - https://grafana.wikimedia.org/d/K1qm1j-Wz/ripe-atlas?orgId=1 - https://alerts.wikimedia.org/?q=alertname%3DRipeAtlasAnchorUnreachable [02:44:27] (03PS2) 10Andrea Denisse: feat: on-call analytics tool correlating ICS shifts with klaxon incidents [software/klaxon] - 10https://gerrit.wikimedia.org/r/1314087 (owner: 10CDanis) [02:44:28] (03PS2) 10Andrea Denisse: [WIP DNM] ICS import maybe-improvements [software/klaxon] - 10https://gerrit.wikimedia.org/r/1314088 (owner: 10CDanis) [02:45:59] (03CR) 10CI reject: [V:04-1] feat: on-call analytics tool correlating ICS shifts with klaxon incidents [software/klaxon] - 10https://gerrit.wikimedia.org/r/1314087 (owner: 10CDanis) [02:46:00] (03CR) 10CI reject: [V:04-1] [WIP DNM] ICS import maybe-improvements [software/klaxon] - 10https://gerrit.wikimedia.org/r/1314088 (owner: 10CDanis) [02:56:20] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328752 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [02:56:20] (03CR) 10TrainBranchBot: [C:03+2] "Approved by tstarling@deploy1003 using scap backport" [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329728 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [02:57:22] (03Merged) 10jenkins-bot: Grant produnto-update to groups that have editprotected [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1328752 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [02:57:26] (03Merged) 10jenkins-bot: Produnto: Add IPv6 range for GitLab [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1329728 (https://phabricator.wikimedia.org/T421436) (owner: 10Tim Starling) [02:58:06] !log tstarling@deploy1003 Started scap sync-world: Backport for [[gerrit:1328752|Grant produnto-update to groups that have editprotected (T421436)]], [[gerrit:1329728|Produnto: Add IPv6 range for GitLab (T421436)]] [02:58:09] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [03:02:18] !log tstarling@deploy1003 tstarling: Backport for [[gerrit:1328752|Grant produnto-update to groups that have editprotected (T421436)]], [[gerrit:1329728|Produnto: Add IPv6 range for GitLab (T421436)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [03:02:19] (03CR) 10Ryan Kemper: [C:03+2] Enable presto join spill config [puppet] - 10https://gerrit.wikimedia.org/r/1330287 (https://phabricator.wikimedia.org/T435862) (owner: 10Joal) [03:05:26] !log tstarling@deploy1003 tstarling: Continuing with deployment [03:10:50] !log tstarling@deploy1003 Finished scap sync-world: Backport for [[gerrit:1328752|Grant produnto-update to groups that have editprotected (T421436)]], [[gerrit:1329728|Produnto: Add IPv6 range for GitLab (T421436)]] (duration: 12m 44s) [03:10:53] T421436: Deploy Produnto extension to production - https://phabricator.wikimedia.org/T421436 [03:38:21] !log ryankemper@cumin2003 START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons. [03:40:41] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons. [03:42:56] !log T435862 Verified `spill_enabled=true` and `join_spill_enabled=true` on the Presto test cluster after restarting its coordinator and worker; both nodes are active and distributed query `20260828_033922_00002_u9k78` succeeded [03:42:58] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [03:42:59] T435862: With spill_enabled = true, Presto fails queries with CROSS JOIN - https://phabricator.wikimedia.org/T435862 [03:58:51] !log T435862 Restarted both production Presto coordinators (standby first, then active) after enabling JOIN spilling; the active coordinator reports `spill_enabled=true` and `join_spill_enabled=true` and all 15 workers rejoined [03:58:53] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [03:58:54] T435862: With spill_enabled = true, Presto fails queries with CROSS JOIN - https://phabricator.wikimedia.org/T435862 [03:59:47] !log ryankemper@cumin2003 START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons. [04:06:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:08:48] PROBLEM - Backup freshness on backup1014 is CRITICAL: All failures: 1 (krb1004), Fresh: 141 jobs https://wikitech.wikimedia.org/wiki/Bacula%23Monitoring [04:17:29] (03CR) 10Ryan Kemper: kafka: converge topic config from hieradata (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [04:19:28] (03PS2) 10Andrea Denisse: Create views to correlate shift data [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330790 (https://phabricator.wikimedia.org/T434730) [04:19:28] (03CR) 10Andrea Denisse: "I just noted I only submitted the test on my previous patch. Some part of the DB are dirty, we can polish up the DB schema once it's time " [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330790 (https://phabricator.wikimedia.org/T434730) (owner: 10Andrea Denisse) [04:28:04] (03PS3) 10Andrea Denisse: Create views to correlate shift data [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330790 (https://phabricator.wikimedia.org/T434730) [04:31:52] !log ryankemper@cumin2003 END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons. [04:34:41] !log T435862 Completed the rolling restart of all 15 production Presto workers; verified `true/true` spill settings across coordinators and workers, full runtime membership, the formerly failing query, and Superset dashboard 757. This is done [04:34:43] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [04:34:44] T435862: With spill_enabled = true, Presto fails queries with CROSS JOIN - https://phabricator.wikimedia.org/T435862 [04:36:48] (03PS3) 10Ryan Kemper: kafka: converge topic config from hieradata [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) [04:38:21] 10ops-eqiad, 06SRE, 06DC-Ops: Lumen eqiad<->codfw transport outage (Aug 2026). - https://phabricator.wikimedia.org/T435810#12264108 (10cmooney) Lumen did some work overnight it seems: ` 2026-08-28 02:33:43 GMT - Lumen observed a higher-level issue in Columbus, OH on 8/24 that bounced 15times but has been cle... [04:39:48] 06SRE, 10LDAP-Access-Requests, 06WMF-NDA-Requests: Request for LDAP NDA Access for CentralNotice Metrics for Yahya - https://phabricator.wikimedia.org/T436294#12264109 (10Aklapper) @Yahya Did you manually add #WMF-NDA-Requests to this task, or was that already part of some task template which you used? [04:39:59] 06SRE, 10LDAP-Access-Requests: Request for LDAP NDA Access for CentralNotice Metrics for Yahya - https://phabricator.wikimedia.org/T436294#12264110 (10Aklapper) [04:58:56] 06SRE, 10LDAP-Access-Requests: Request for LDAP NDA Access for CentralNotice Metrics for Yahya - https://phabricator.wikimedia.org/T436294#12264134 (10Yahya) It was part of a template mentioned here: [[https://wikitech.wikimedia.org/wiki/Volunteer_NDA|wikitech:Volunteer NDA]] [05:05:18] (03CR) 10Ryan Kemper: "check puppet" [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [05:05:27] (03CR) 10Ryan Kemper: kafka: converge topic config from hieradata (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1318976 (https://phabricator.wikimedia.org/T276088) (owner: 10Ryan Kemper) [05:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [05:08:48] RECOVERY - Backup freshness on backup1014 is OK: Fresh: 142 jobs https://wikitech.wikimedia.org/wiki/Bacula%23Monitoring [05:55:37] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [06:00:04] Deploy window MediaWiki infrastructure (UTC early) (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260828T0600) [06:19:45] 06SRE, 10LDAP-Access-Requests: Request for LDAP NDA Access for CentralNotice Metrics for Yahya - https://phabricator.wikimedia.org/T436294#12264209 (10Aklapper) Ah, thanks! But are you requesting LDAP group membership to access CentralNotice Metrics (not sure where these are located), and/or access to Phab tas... [06:19:55] (03CR) 10Slyngshede: [C:03+1] upload: Drop profile::cache::upload::upload_webp_hits_threshold [puppet] - 10https://gerrit.wikimedia.org/r/1327656 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [06:22:14] (03CR) 10Slyngshede: [C:03+1] "Presumably we might want to try to go even lower later, based on the numbers." [puppet] - 10https://gerrit.wikimedia.org/r/1327657 (https://phabricator.wikimedia.org/T431150) (owner: 10Ladsgroup) [06:24:29] (03CR) 10Muehlenhoff: [C:03+1] "Looks good" [software/bitu] - 10https://gerrit.wikimedia.org/r/1320185 (owner: 10Perryprog) [06:25:10] (03PS2) 10Muehlenhoff: service::node: Use the LVS-balanced URL downloaders [puppet] - 10https://gerrit.wikimedia.org/r/1328178 (https://phabricator.wikimedia.org/T429175) [06:26:30] !log jmm@cumin2003 START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org [06:27:55] (03CR) 10Samwilson: [C:03+1] "Works in my local testing." [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270150 (https://phabricator.wikimedia.org/T290345) (owner: 10TheDJ) [06:30:37] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org [07:00:04] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260828T0700) [07:16:34] (03CR) 10Muehlenhoff: "Yeah, I've simply ignored this. I think we can decom the old bullseye nodes end of next week (barring any issues with the new trixie nodes" [cookbooks] - 10https://gerrit.wikimedia.org/r/1329523 (https://phabricator.wikimedia.org/T429175) (owner: 10Muehlenhoff) [07:17:38] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12264298 (10MoritzMuehlenhoff) [07:47:08] 06SRE, 10decommission-hardware: decommission cumin2002.codfw.wmnet - https://phabricator.wikimedia.org/T436330 (10MoritzMuehlenhoff) 03NEW [07:47:21] 06SRE, 10decommission-hardware: decommission cumin2002.codfw.wmnet - https://phabricator.wikimedia.org/T436330#12264412 (10MoritzMuehlenhoff) [07:49:46] 06SRE, 10decommission-hardware: decommission cumin2002.codfw.wmnet - https://phabricator.wikimedia.org/T436330#12264414 (10MoritzMuehlenhoff) [07:49:46] !log jmm@cumin2003 START - Cookbook sre.hosts.decommission for hosts cumin2002.codfw.wmnet [07:49:46] 06SRE, 06Infrastructure-Foundations: Upgrade Cumin hosts to Trixie - https://phabricator.wikimedia.org/T427897#12264413 (10MoritzMuehlenhoff) [07:49:57] (03PS2) 10Dpogorzelski: kserve: 0.20 upstream alignment [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327557 (https://phabricator.wikimedia.org/T433973) [07:49:57] (03PS3) 10Dpogorzelski: kserve: add LLMInferenceService support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329506 (https://phabricator.wikimedia.org/T433973) [07:49:58] (03PS1) 10Dpogorzelski: kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) [07:50:52] (03PS2) 10Dpogorzelski: kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) [07:51:55] (03CR) 10Slyngshede: [C:03+2] Use os.environ.get throughout docker_settings.py [software/bitu] - 10https://gerrit.wikimedia.org/r/1320185 (owner: 10Perryprog) [07:52:34] (03PS1) 10Tiziano Fogli: prometheus: lower the zombie series detection threshold [alerts] - 10https://gerrit.wikimedia.org/r/1331385 [07:52:54] jmm@cumin2003 decommission (PID 1334898) is awaiting input [07:54:38] (03Merged) 10jenkins-bot: Use os.environ.get throughout docker_settings.py [software/bitu] - 10https://gerrit.wikimedia.org/r/1320185 (owner: 10Perryprog) [07:55:51] (03CR) 10Daniel Kinzler: mediawiki-vhost: Redirect /api/ to /w/rest.php (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1329607 (https://phabricator.wikimedia.org/T433547) (owner: 10Clément Goubert) [07:56:26] (03PS1) 10Muehlenhoff: Replace grant for cumin2002 with cumin2003 [puppet] - 10https://gerrit.wikimedia.org/r/1331411 (https://phabricator.wikimedia.org/T427884) [07:57:44] (03PS2) 10Muehlenhoff: Replace grant for cumin2002 with cumin2003 [puppet] - 10https://gerrit.wikimedia.org/r/1331411 (https://phabricator.wikimedia.org/T427884) [07:59:36] (03CR) 10CI reject: [V:04-1] kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [08:00:14] (03CR) 10CI reject: [V:04-1] kserve: add LLMInferenceService support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329506 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [08:01:01] (03CR) 10CI reject: [V:04-1] kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [08:02:12] (03CR) 10Jelto: [V:03+1 C:03+1] "lgtm, verified build locally" [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1330459 (https://phabricator.wikimedia.org/T427069) (owner: 10JMeybohm) [08:06:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:18:34] (03PS1) 10Slyngshede: switchdc: remove parsoid [cookbooks] - 10https://gerrit.wikimedia.org/r/1331446 (https://phabricator.wikimedia.org/T433363) [08:18:45] jmm@cumin2003 decommission (PID 1334898) is awaiting input [08:22:49] (03CR) 10Slyngshede: "mw-parsoid now no longer used to serve anything prod-related" [cookbooks] - 10https://gerrit.wikimedia.org/r/1331446 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [08:28:49] !log jmm@cumin2003 START - Cookbook sre.dns.netbox [08:34:18] jmm@cumin2003 decommission (PID 1334898) is awaiting input [08:35:54] (03CR) 10Jelto: [V:03+1 C:03+1] "lgtm, verified build on `build2004` and binary locally" [debs/kubernetes] (v1.34) - 10https://gerrit.wikimedia.org/r/1330382 (https://phabricator.wikimedia.org/T427069) (owner: 10JMeybohm) [08:43:10] (03CR) 10CWilliams: "Is there a reason that this uses REDACTED, whereas production.sql.erb uses `<%= @cumin_pass %>`?" [puppet] - 10https://gerrit.wikimedia.org/r/1331411 (https://phabricator.wikimedia.org/T427884) (owner: 10Muehlenhoff) [08:43:23] (03PS1) 10Muehlenhoff: Rename puppetmaster::puppetdb to puppetdb [puppet] - 10https://gerrit.wikimedia.org/r/1331470 [08:58:35] !log jmm@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin2002.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [08:59:02] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin2002.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003" [08:59:03] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [08:59:04] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin2002.codfw.wmnet [08:59:11] 06SRE, 10decommission-hardware: decommission cumin2002.codfw.wmnet - https://phabricator.wikimedia.org/T436330#12264597 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by jmm@cumin2003 for hosts: `cumin2002.codfw.wmnet` - cumin2002.codfw.wmnet (**PASS**) - Downtimed host on Icinga/Alertmanag... [08:59:51] (03CR) 10Muehlenhoff: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1331470 (owner: 10Muehlenhoff) [09:04:45] (03PS1) 10Filippo Giunchedi: cloudnfs: add textfile exporter for extra nfsd metrics [puppet] - 10https://gerrit.wikimedia.org/r/1331487 (https://phabricator.wikimedia.org/T436252) [09:05:23] (03CR) 10CI reject: [V:04-1] cloudnfs: add textfile exporter for extra nfsd metrics [puppet] - 10https://gerrit.wikimedia.org/r/1331487 (https://phabricator.wikimedia.org/T436252) (owner: 10Filippo Giunchedi) [09:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [09:07:17] (03PS2) 10Filippo Giunchedi: cloudnfs: add textfile exporter for extra nfsd metrics [puppet] - 10https://gerrit.wikimedia.org/r/1331487 (https://phabricator.wikimedia.org/T436252) [09:12:48] (03PS3) 10Dpogorzelski: kserve: 0.20 upstream alignment [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327557 (https://phabricator.wikimedia.org/T433973) [09:12:48] (03PS4) 10Dpogorzelski: kserve: add LLMInferenceService support [deployment-charts] - 10https://gerrit.wikimedia.org/r/1329506 (https://phabricator.wikimedia.org/T433973) [09:12:49] (03PS3) 10Dpogorzelski: kserve: expose LLMInferenceService via istio [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331389 (https://phabricator.wikimedia.org/T433973) [09:30:16] (03PS1) 10Muehlenhoff: Remove obsolete config for cumin2002 [puppet] - 10https://gerrit.wikimedia.org/r/1331494 [09:35:01] (03CR) 10Muehlenhoff: [C:03+2] Move weekly build of production images to build2004 [puppet] - 10https://gerrit.wikimedia.org/r/1329554 (https://phabricator.wikimedia.org/T417389) (owner: 10Muehlenhoff) [09:35:47] 06SRE, 10decommission-hardware: decommission cumin2002.codfw.wmnet - https://phabricator.wikimedia.org/T436330#12264676 (10MoritzMuehlenhoff) [09:37:49] (03CR) 10Clément Goubert: mediawiki-vhost: Redirect /api/ to /w/rest.php (033 comments) [puppet] - 10https://gerrit.wikimedia.org/r/1329607 (https://phabricator.wikimedia.org/T433547) (owner: 10Clément Goubert) [09:39:05] (03PS2) 10Clément Goubert: mediawiki: Redirect /api/ to /w/rest.php [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330505 (https://phabricator.wikimedia.org/T433547) [09:42:30] (03PS1) 10Lucas Werkmeister (WMDE): wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331503 (https://phabricator.wikimedia.org/T435632) [09:44:22] (03PS4) 10Clément Goubert: mediawiki-vhost: Redirect /api/ to /w/rest.php [puppet] - 10https://gerrit.wikimedia.org/r/1329607 (https://phabricator.wikimedia.org/T433547) [09:47:35] (03CR) 10Clément Goubert: [C:03+1] switchdc: remove parsoid [cookbooks] - 10https://gerrit.wikimedia.org/r/1331446 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [09:55:37] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [10:04:51] (03PS1) 10Muehlenhoff: Remove cumin2002 from site.pp/preseed [puppet] - 10https://gerrit.wikimedia.org/r/1331514 [10:07:44] (03CR) 10Clément Goubert: "Just tested the latest changes on `mw-experimental` with the right `wgRestPath` and it seems to work just fine. This should be mergeable e" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330505 (https://phabricator.wikimedia.org/T433547) (owner: 10Clément Goubert) [10:11:05] 06SRE, 06Infrastructure-Foundations: Integrate Bookworm 12.15 point update - https://phabricator.wikimedia.org/T434631#12264850 (10MoritzMuehlenhoff) [10:11:54] (03PS2) 10Atsuko: dse-k8s-eqiad: provision the airflow instance [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327553 (https://phabricator.wikimedia.org/T416709) [10:14:06] (03Abandoned) 10Muehlenhoff: kerberos: exclude krb1002 to allow reimage and move-vlan [puppet] - 10https://gerrit.wikimedia.org/r/1294952 (https://phabricator.wikimedia.org/T421706) (owner: 10Elukey) [10:14:55] (03CR) 10Muehlenhoff: [C:03+2] Remove cumin2002 from site.pp/preseed [puppet] - 10https://gerrit.wikimedia.org/r/1331514 (owner: 10Muehlenhoff) [10:15:28] 06SRE, 10decommission-hardware: decommission cumin2002.codfw.wmnet - https://phabricator.wikimedia.org/T436330#12264864 (10MoritzMuehlenhoff) [10:15:45] 10ops-codfw, 06SRE, 06DC-Ops, 10decommission-hardware: decommission cumin2002.codfw.wmnet - https://phabricator.wikimedia.org/T436330#12264865 (10MoritzMuehlenhoff) [10:16:53] (03PS1) 10Joal: Update X-is-browser webrequest_sample turnilo conf [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331524 (https://phabricator.wikimedia.org/T435152) [10:18:09] (03PS1) 10Muehlenhoff: Add VM for cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1331525 (https://phabricator.wikimedia.org/T427897) [10:18:31] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Migrate remaining container build/report steps from build2001 to build2004 - https://phabricator.wikimedia.org/T417389#12264875 (10MoritzMuehlenhoff) [10:25:08] (03CR) 10JMeybohm: [C:03+1] Remove obsolete Cumin alias for parsoid-testing [puppet] - 10https://gerrit.wikimedia.org/r/1329157 (owner: 10Muehlenhoff) [10:29:25] (03CR) 10Muehlenhoff: [C:03+2] Remove obsolete Cumin alias for parsoid-testing [puppet] - 10https://gerrit.wikimedia.org/r/1329157 (owner: 10Muehlenhoff) [10:29:36] (03CR) 10FNegri: [C:03+1] aptrepo: Stop mirroring Helm upstream repos [puppet] - 10https://gerrit.wikimedia.org/r/1330356 (owner: 10Majavah) [10:32:41] (03CR) 10Atsuko: [C:04-2] "I will update the patch today with the description and members of the group. Waiting @mpopov@wikimedia.org for the list of users. The reas" [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [10:33:36] (03CR) 10FNegri: [C:03+1] "LGTM" [docker-images/toollabs-images] - 10https://gerrit.wikimedia.org/r/1310212 (https://phabricator.wikimedia.org/T432078) (owner: 10Raymond Ndibe) [10:34:08] (03CR) 10Muehlenhoff: [C:03+2] Add VM for cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1331525 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [10:35:42] (03PS1) 10Atsuko: Add airflow-experiment-platform-ops LDAP group [puppet] - 10https://gerrit.wikimedia.org/r/1331529 (https://phabricator.wikimedia.org/T416709) [10:43:58] !log aokoth@cumin1003 START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T436069 [10:44:40] (03CR) 10Muehlenhoff: [C:03+1] "Looks good. I can also create the LDAP group, but I need to know which one is supposed to be the initial group member (our LDAP scheme req" [puppet] - 10https://gerrit.wikimedia.org/r/1331529 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [10:46:38] (03CR) 10Atsuko: "I think @mpopov@wikimedia.org as the requestor of the instance should be initial group member." [puppet] - 10https://gerrit.wikimedia.org/r/1331529 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [11:00:05] Deploy window No deploys all day! See Deployments/Emergencies if things are broken. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260828T0700) [11:00:05] jelto, arnoldokoth, mutante, and arnaudb: gettimeofday() says it's time for GitLab version upgrades. (https://wikitech.wikimedia.org/wiki/Deployments#deploycal-item-20260828T1100) [11:01:20] aokoth@cumin1003 aokoth: The backup on gitlab1004 is complete, ready to proceed with upgrade. [11:07:56] PROBLEM - Gitlab HTTPS SSL Expiry on gitlab.wikimedia.org is CRITICAL: connect to address gitlab.wikimedia.org and port 443: Connection refused https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [11:08:22] PROBLEM - Gitlab HTTPS healthcheck on gitlab.wikimedia.org is CRITICAL: HTTP CRITICAL: HTTP/1.1 502 Bad Gateway - 2353 bytes in 0.013 second response time https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [11:08:56] RECOVERY - Gitlab HTTPS SSL Expiry on gitlab.wikimedia.org is OK: OK - Certificate gitlab.wikimedia.org will expire on Sat 31 Oct 2026 08:10:49 AM GMT +0000. https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [11:09:22] RECOVERY - Gitlab HTTPS healthcheck on gitlab.wikimedia.org is OK: HTTP OK: HTTP/1.1 200 OK - 28821 bytes in 0.202 second response time https://wikitech.wikimedia.org/wiki/GitLab%23Monitoring [11:12:05] !log aokoth@cumin1003 END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - T436069 [11:13:06] PROBLEM - OSPF status on cr2-eqsin is CRITICAL: OSPFv2: 3/4 UP : OSPFv3: 3/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:14:06] RECOVERY - OSPF status on cr2-eqsin is OK: OSPFv2: 4/4 UP : OSPFv3: 4/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [11:23:06] !log jmm@cumin2003 START - Cookbook sre.ganeti.makevm for new host cumin1004.eqiad.wmnet [11:23:09] !log jmm@cumin2003 START - Cookbook sre.dns.netbox [11:26:10] FIRING: BFDdown: BFD session down between cr1-codfw and fe80::669:8f07:ece:f4e7 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:28:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 6.997% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:28:41] jmm@cumin2003 makevm (PID 1379823) is awaiting input [11:29:15] FIRING: MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?panelId=18&fullscreen&orgId=1&var-datasource=eqiad%20prometheus/ops - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [11:31:10] RESOLVED: BFDdown: BFD session down between cr1-codfw and fe80::669:8f07:ece:f4e7 - https://wikitech.wikimedia.org/wiki/Network_monitoring#BFD_status - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-codfw:9804 - https://alerts.wikimedia.org/?q=alertname%3DBFDdown [11:31:15] FIRING: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 0% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:31:50] !ack [11:31:51] 8311 (ACKED) PHPFPMTooBusy sre (mw-api-ext main eqiad) [11:32:02] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-internal-scholarly_443: Servers wdqs1027.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [11:33:02] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [11:33:15] FIRING: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-ext releases routed via main (k8s) 2.364s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [11:34:15] FIRING: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [11:36:57] FIRING: ProbeDown: Service mw-api-ext:4447 has failed probes (http_mw-api-ext_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#mw-api-ext:4447 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:38:07] !log jmm@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM cumin1004.eqiad.wmnet - jmm@cumin2003" [11:38:11] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM cumin1004.eqiad.wmnet - jmm@cumin2003" [11:38:12] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [11:38:12] !log jmm@cumin2003 START - Cookbook sre.dns.wipe-cache cumin1004.eqiad.wmnet on all recursors [11:38:12] FIRING: VarnishUnavailable: varnish-text has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/Varnish#Diagnosing_Varnish_alerts - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=3 - https://alerts.wikimedia.org/?q=alertname%3DVarnishUnavailable [11:38:13] FIRING: HaproxyUnavailable: HAProxy (cache_text) has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyUnavailable [11:38:15] !log jmm@cumin2003 END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cumin1004.eqiad.wmnet on all recursors [11:38:42] !ack [11:38:43] 8312 (ACKED) VarnishUnavailable global sre (varnish-text thanos-rule@main) [11:38:43] 8313 (ACKED) HaproxyUnavailable cache_text global sre (thanos-rule@main) [11:38:51] !log jmm@cumin2003 START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM cumin1004.eqiad.wmnet - jmm@cumin2003" [11:38:54] !log jmm@cumin2003 END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM cumin1004.eqiad.wmnet - jmm@cumin2003" [11:38:58] !log add new LDAP group cn=airflow-experiment-platform-ops,ou=groups,dc=wikimedia,dc=org T416709 [11:40:02] (03CR) 10Muehlenhoff: [C:03+1] "The LDAP group has been added, good to merge" [puppet] - 10https://gerrit.wikimedia.org/r/1331529 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [11:41:55] jmm@cumin2003 makevm (PID 1379823) is awaiting input [11:41:57] RESOLVED: ProbeDown: Service mw-api-ext:4447 has failed probes (http_mw-api-ext_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#mw-api-ext:4447 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/service&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [11:43:12] RESOLVED: VarnishUnavailable: varnish-text has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/Varnish#Diagnosing_Varnish_alerts - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=3 - https://alerts.wikimedia.org/?q=alertname%3DVarnishUnavailable [11:43:13] RESOLVED: HaproxyUnavailable: HAProxy (cache_text) has reduced HTTP availability #page - https://wikitech.wikimedia.org/wiki/HAProxy#HAProxy_for_edge_caching - https://grafana.wikimedia.org/d/000000479/frontend-traffic?viewPanel=13 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyUnavailable [11:43:27] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [11:43:27] T416709: Airflow instance for Experiment Platform - https://phabricator.wikimedia.org/T416709 [11:44:17] lovely I didn't get any text and/or notification [11:46:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 1.736% idle #page - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:48:04] FIRING: MediaWikiElevatedUnknownLogins: Elevated number of login successes (source unknown) via mw-api-ext - TODO - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?from=now-6h&orgId=1&to=now&viewPanel=26 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiElevatedUnknownLogins [11:48:15] RESOLVED: PHPFPMTooBusy: Not enough idle PHP-FPM workers for Mediawiki mw-api-ext releases routed via main at eqiad: 1.562% idle - https://bit.ly/wmf-fpmsat - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=84&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-container_name=All&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DPHPFPMTooBusy [11:48:15] RESOLVED: MediaWikiLatencyExceeded: p75 latency high: eqiad mw-api-ext releases routed via main (k8s) 2.124s - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook#Average_latency_exceeded - https://grafana.wikimedia.org/d/U7JT--knk/mw-on-k8s?orgId=1&viewPanel=55&var-dc=eqiad%20prometheus/k8s&var-service=mediawiki&var-namespace=mw-api-ext&var-release=main - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiLatencyExceeded [11:49:15] RESOLVED: [2x] MediaWikiHighErrorRate: Elevated rate of MediaWiki errors - kube-mw-api-ext - https://wikitech.wikimedia.org/wiki/Application_servers/Runbook - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiHighErrorRate [11:53:04] RESOLVED: MediaWikiElevatedUnknownLogins: Elevated number of login successes (source unknown) via mw-api-ext - TODO - https://grafana.wikimedia.org/d/000000438/mediawiki-exceptions-alerts?from=now-6h&orgId=1&to=now&viewPanel=26 - https://alerts.wikimedia.org/?q=alertname%3DMediaWikiElevatedUnknownLogins [11:56:25] !log jmm@cumin2003 START - Cookbook sre.hosts.reimage for host cumin1004.eqiad.wmnet with OS trixie [11:56:44] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Upgrade Cumin hosts to Trixie - https://phabricator.wikimedia.org/T427897#12265204 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by jmm@cumin2003 for host cumin1004.eqiad.wmnet with OS trixie [12:06:40] FIRING: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:07:57] !log jmm@cumin2003 START - Cookbook sre.hosts.downtime for 2:00:00 on cumin1004.eqiad.wmnet with reason: host reimage [12:10:43] (03CR) 10Atsuko: [C:03+2] Add airflow-experiment-platform-ops LDAP group [puppet] - 10https://gerrit.wikimedia.org/r/1331529 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [12:14:19] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cumin1004.eqiad.wmnet with reason: host reimage [12:27:05] (03CR) 10Ladsgroup: [C:03+1] "Will deploy next week" [software/thumbor-plugins] - 10https://gerrit.wikimedia.org/r/1270150 (https://phabricator.wikimedia.org/T290345) (owner: 10TheDJ) [12:29:09] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12265248 (10MoritzMuehlenhoff) [12:31:00] !log jmm@cumin2003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cumin1004.eqiad.wmnet with OS trixie [12:31:01] !log jmm@cumin2003 END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host cumin1004.eqiad.wmnet [12:31:12] 06SRE, 06Infrastructure-Foundations, 13Patch-For-Review: Upgrade Cumin hosts to Trixie - https://phabricator.wikimedia.org/T427897#12265254 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by jmm@cumin2003 for host cumin1004.eqiad.wmnet with OS trixie completed: - cumin1004 (**PASS**) -... [12:33:51] (03CR) 10JMeybohm: [V:03+2 C:03+2] Update pause image to 3.10.1 (k8s 1.34) [docker-images/production-images] - 10https://gerrit.wikimedia.org/r/1330459 (https://phabricator.wikimedia.org/T427069) (owner: 10JMeybohm) [12:34:17] (03CR) 10JMeybohm: [C:03+2] Update to v1.34.11 [debs/kubernetes] (v1.34) - 10https://gerrit.wikimedia.org/r/1330382 (https://phabricator.wikimedia.org/T427069) (owner: 10JMeybohm) [12:40:52] FIRING: GitLabReplicaDataStale: GitLab - replica gitlab1003:0 serves data older than 30h - https://wikitech.wikimedia.org/wiki/GitLab/Backup_and_Restore - https://grafana.wikimedia.org/d/R_1IvBZnz/gitlab-omnibus-overview - https://alerts.wikimedia.org/?q=alertname%3DGitLabReplicaDataStale [12:41:16] (03CR) 10Sadiya.mohammed13: [C:03+1] wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331503 (https://phabricator.wikimedia.org/T435632) (owner: 10Lucas Werkmeister (WMDE)) [12:41:40] (03CR) 10Elukey: [C:03+1] turnilo: Expose X-analytics thumb_generated in webrequest_sampled_live (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1328212 (https://phabricator.wikimedia.org/T435634) (owner: 10Ladsgroup) [12:41:52] FIRING: GitLabRestoreVersionMismatch: GitLab - restore blocked on gitlab1003:0 by a version mismatch - https://wikitech.wikimedia.org/wiki/GitLab/Backup_and_Restore - https://grafana.wikimedia.org/d/R_1IvBZnz/gitlab-omnibus-overview - https://alerts.wikimedia.org/?q=alertname%3DGitLabRestoreVersionMismatch [12:45:09] (03PS1) 10Slyngshede: site.pp move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1331590 (https://phabricator.wikimedia.org/T436363) [12:57:20] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12265378 (10MoritzMuehlenhoff) [12:58:33] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): krb1002 booting up extremely slow - https://phabricator.wikimedia.org/T435354#12265385 (10Gehel) [12:58:42] 06SRE, 10Infrastructure Security, 06Data-Platform-SRE (2026-08-28 - 2026-09-18), 07SecTeam-Processed, and 2 others: July 2026 Bookworm/Trixie reboots: Data Platform SRE - https://phabricator.wikimedia.org/T431826#12265389 (10Gehel) [13:01:25] RESOLVED: SystemdUnitFailed: send_tile_invalidations.service on maps1011:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [13:01:31] 10SRE-SLO, 10observability, 10Wikidata, 06Wikidata Platform Team, and 3 others: Update WDQS SLOs to reflect graph split changes - https://phabricator.wikimedia.org/T393966#12265449 (10Gehel) [13:01:51] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Attempt to use a VM as Kerberos replica - https://phabricator.wikimedia.org/T435873#12265459 (10Gehel) [13:02:11] 10ops-eqiad, 06SRE, 06DC-Ops, 10Kafka-Infrastructure, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Heterogeneous kafka-jumbo-eqiad rack placement - https://phabricator.wikimedia.org/T435775#12265469 (10Gehel) [13:02:19] 06SRE, 06Data-Engineering, 10Kafka-Infrastructure, 06serviceops-radar, and 3 others: Configuration Management for Kafka settings - https://phabricator.wikimedia.org/T276088#12265465 (10Gehel) [13:03:07] PROBLEM - OSPF status on cr2-eqsin is CRITICAL: OSPFv2: 3/4 UP : OSPFv3: 3/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:03:08] 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Follow up on multiple RAID / drive issues - https://phabricator.wikimedia.org/T426610#12265489 (10Gehel) [13:04:05] RECOVERY - OSPF status on cr2-eqsin is OK: OSPFv2: 4/4 UP : OSPFv3: 4/4 UP https://wikitech.wikimedia.org/wiki/Network_monitoring%23OSPF_status [13:05:04] 06SRE, 06Infrastructure-Foundations, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): install1005 running out of disk due to squid log volume from an-worker* webproxy workload - https://phabricator.wikimedia.org/T435555#12265537 (10Gehel) [13:06:51] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): eqiad row A&B host migration details request for Search Platform - https://phabricator.wikimedia.org/T432651#12265562 (10Gehel) [13:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [13:10:52] FIRING: [2x] GitLabReplicaDataStale: GitLab - replica gitlab1003:0 serves data older than 30h - https://wikitech.wikimedia.org/wiki/GitLab/Backup_and_Restore - https://grafana.wikimedia.org/d/R_1IvBZnz/gitlab-omnibus-overview - https://alerts.wikimedia.org/?q=alertname%3DGitLabReplicaDataStale [13:11:52] FIRING: [2x] GitLabRestoreVersionMismatch: GitLab - restore blocked on gitlab1003:0 by a version mismatch - https://wikitech.wikimedia.org/wiki/GitLab/Backup_and_Restore - https://grafana.wikimedia.org/d/R_1IvBZnz/gitlab-omnibus-overview - https://alerts.wikimedia.org/?q=alertname%3DGitLabRestoreVersionMismatch [13:17:30] (03PS1) 10Muehlenhoff: Apply cluster::management role to cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1331611 (https://phabricator.wikimedia.org/T427897) [13:24:14] 10SRE-swift-storage, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): RdfStreamingUpdaterSpaceUsageTooHigh - https://phabricator.wikimedia.org/T431506#12265731 (10Gehel) [13:24:24] 07sre-alert-triage, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Alert in need of triage: AlertLintProblem (instance localhost:9123) - https://phabricator.wikimedia.org/T430139#12265737 (10Gehel) [13:30:03] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [13:30:21] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1018.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1022.eqiad.wmnet, wdqs1017.eqiad.wmnet, wdqs1020.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [13:31:03] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [13:31:21] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [13:33:57] (03CR) 10Bking: [C:03+2] "PCC looks good, self-merging in the interest of time." [puppet] - 10https://gerrit.wikimedia.org/r/1330644 (https://phabricator.wikimedia.org/T436258) (owner: 10Bking) [13:34:10] (03PS1) 10Jelto: etherpad-next: add discovery ingress records [dns] - 10https://gerrit.wikimedia.org/r/1331613 (https://phabricator.wikimedia.org/T435509) [13:35:03] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1017.eqiad.wmnet, wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [13:35:21] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1011.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1018.eqiad.wmnet, wdqs1016.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1017.eqiad.wmnet, wdqs1019.eqiad.wmnet, wdqs1020.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [13:35:24] (03PS4) 10Atsuko: Grant sudo privileges for the analytics-experiment-users group [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) [13:35:33] (03PS2) 10Jelto: etherpad-next: add discovery ingress records [dns] - 10https://gerrit.wikimedia.org/r/1331613 (https://phabricator.wikimedia.org/T435509) [13:36:02] (03CR) 10Atsuko: "removing -2 as it is ready for review" [puppet] - 10https://gerrit.wikimedia.org/r/1328595 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [13:37:03] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [13:37:09] 10ops-codfw, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Q1:rack/setup/install cirrussearch21[16-20] - https://phabricator.wikimedia.org/T436287#12265814 (10Gehel) [13:37:21] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [13:37:35] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Q1:rack/setup/install cirrussearch11[26-30] - https://phabricator.wikimedia.org/T436285#12265821 (10Gehel) [13:41:12] (03PS1) 10Atsuko: Typo in airflow-experiment-platform-ops LDAP group manager list [puppet] - 10https://gerrit.wikimedia.org/r/1331615 (https://phabricator.wikimedia.org/T416709) [13:45:42] (03PS1) 10Zabe: Set remote virtual domain for globalusage [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1331617 (https://phabricator.wikimedia.org/T422940) [13:46:22] (03CR) 10Bking: [C:03+2] dse-k8s-eqiad: Add Matomo namespaces [deployment-charts] - 10https://gerrit.wikimedia.org/r/1330653 (https://phabricator.wikimedia.org/T436258) (owner: 10Bking) [13:55:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [13:57:49] (03PS1) 10Jelto: service: add etherpad-next to service catalog [puppet] - 10https://gerrit.wikimedia.org/r/1331623 (https://phabricator.wikimedia.org/T435509) [14:01:21] !log bking@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [14:03:11] !log bking@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [14:04:53] (03CR) 10CDanis: [C:03+1] "Looks good, thanks! Just this one thing" [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330487 (owner: 10Hnowlan) [14:09:29] (03CR) 10Thiemo Kreuz (WMDE): [C:03+1] wikidata-query-builder: bump image version [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331503 (https://phabricator.wikimedia.org/T435632) (owner: 10Lucas Werkmeister (WMDE)) [14:09:38] (03CR) 10Atsuko: [C:03+1] Update X-is-browser webrequest_sample turnilo conf [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331524 (https://phabricator.wikimedia.org/T435152) (owner: 10Joal) [14:12:41] (03CR) 10Elukey: [C:03+1] Apply cluster::management role to cumin1004 [puppet] - 10https://gerrit.wikimedia.org/r/1331611 (https://phabricator.wikimedia.org/T427897) (owner: 10Muehlenhoff) [14:13:01] (03CR) 10Atsuko: [C:03+1] "Charts/helmfiles are fine, don't have any contexr for the detail tho" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1327557 (https://phabricator.wikimedia.org/T433973) (owner: 10Dpogorzelski) [14:14:33] 10ops-eqiad, 06SRE, 06DC-Ops, 06Data-Platform-SRE (2026-08-28 - 2026-09-18): Q1:rack/setup/install cirrussearch11[26-30] - https://phabricator.wikimedia.org/T436285#12265961 (10bking) [14:18:19] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499#12265983 (10elukey) I uploaded the images running on all k8s clusters (with the "future" tags up to latest) and I ended up with t... [14:22:26] !log sukhe@cumin1003 START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie [14:22:30] !log sukhe@cumin1003 START - Cookbook sre.hosts.move-vlan for host cp5022 [14:22:30] !log sukhe@cumin1003 END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cp5022 [14:22:44] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12265995 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by sukhe@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [14:26:53] (03CR) 10FNegri: [C:03+1] "I removed the reference to our custom package from https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Kubernetes/Upgrading_Kuberne" [puppet] - 10https://gerrit.wikimedia.org/r/1330356 (owner: 10Majavah) [14:27:12] (03CR) 10CDanis: "Mostly looking good, just a couple more comments. Today I'll finish up cleanups on the patches that are underneath this one and send them" [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330630 (https://phabricator.wikimedia.org/T434729) (owner: 10Andrea Denisse) [14:29:48] (03PS1) 10Bking: cirrussearch: Add new hosts to site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1331647 (https://phabricator.wikimedia.org/T436285) [14:33:55] (03CR) 10CDanis: [C:03+1] "The code is reasonable and the Puppet diffs look good: https://puppet-compiler.wmflabs.org/output/1329558/7662/ms-fe2018.codfw.wmnet/fulld" [puppet] - 10https://gerrit.wikimedia.org/r/1329558 (https://phabricator.wikimedia.org/T431597) (owner: 10Cparle) [14:36:36] (03PS1) 10Jelto: sre.gitlab.upgrade: extend replica downtime to 60h [cookbooks] - 10https://gerrit.wikimedia.org/r/1331653 (https://phabricator.wikimedia.org/T436361) [14:36:57] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499#12266044 (10elukey) To keep archives happy - I am using the following two bash scripts on deploy1003 to get the list of docker im... [14:41:03] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1013.eqiad.wmnet, wdqs1022.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:42:03] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:42:38] (03CR) 10Muehlenhoff: [C:03+1] "LGTM" [puppet] - 10https://gerrit.wikimedia.org/r/1331615 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [14:43:01] 06SRE, 10Maps, 06Traffic: Allow Wikimedia Maps usage on bicirio.somosmovilidad.gov.co - https://phabricator.wikimedia.org/T436378 (10Devstevenmarin) 03NEW [14:43:47] 06SRE, 10Maps, 06Traffic: Allow Wikimedia Maps usage on bicirio.somosmovilidad.gov.co - https://phabricator.wikimedia.org/T436378#12266098 (10Pppery) a:05Devstevenmarin→03None [14:44:43] 06SRE, 10Maps, 06Traffic: Allow Wikimedia Maps usage on bicirio.somosmovilidad.gov.co - https://phabricator.wikimedia.org/T436378#12266108 (10Pppery) 05Open→03Declined > Wikimedia Affiliate supporting project: None You didn't fill in that field. > If you do not fill in all three fields your reques... [14:45:21] PROBLEM - PyBal backends health check on lvs1019 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1021.eqiad.wmnet, wdqs1013.eqiad.wmnet, wdqs1012.eqiad.wmnet, wdqs1017.eqiad.wmnet, wdqs1019.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:46:21] RECOVERY - PyBal backends health check on lvs1019 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:47:33] (03PS1) 10Muehlenhoff: Add cumin1004 as mysql root client / grant [puppet] - 10https://gerrit.wikimedia.org/r/1331666 (https://phabricator.wikimedia.org/T427897) [14:50:03] PROBLEM - PyBal backends health check on lvs1020 is CRITICAL: PYBAL CRITICAL - CRITICAL - wdqs-main_443: Servers wdqs1013.eqiad.wmnet, wdqs1014.eqiad.wmnet, wdqs1016.eqiad.wmnet are marked down but pooled https://wikitech.wikimedia.org/wiki/PyBal [14:50:10] (03PS1) 10Bking: dse-k8s: Fix TLS hostnames for matomo [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331667 (https://phabricator.wikimedia.org/T436258) [14:50:41] 06SRE, 06Infrastructure-Foundations: Integrate Trixie 13.6 point update - https://phabricator.wikimedia.org/T434866#12266127 (10MoritzMuehlenhoff) [14:51:03] RECOVERY - PyBal backends health check on lvs1020 is OK: PYBAL OK - All pools are healthy https://wikitech.wikimedia.org/wiki/PyBal [14:52:17] (03CR) 10CDanis: Add support to store SplunkOnCall users by email (031 comment) [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330630 (https://phabricator.wikimedia.org/T434729) (owner: 10Andrea Denisse) [14:52:54] (03CR) 10AOkoth: [C:03+1] sre.gitlab.upgrade: extend replica downtime to 60h [cookbooks] - 10https://gerrit.wikimedia.org/r/1331653 (https://phabricator.wikimedia.org/T436361) (owner: 10Jelto) [14:53:36] (03CR) 10Jelto: [C:03+2] sre.gitlab.upgrade: extend replica downtime to 60h [cookbooks] - 10https://gerrit.wikimedia.org/r/1331653 (https://phabricator.wikimedia.org/T436361) (owner: 10Jelto) [14:54:13] !log sukhe@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage [14:57:06] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations, 13Patch-For-Review: Move Docker images under the /v2/releng prefix to S3 - https://phabricator.wikimedia.org/T432829#12266152 (10elukey) @jnuche Hi! I was talking with @dancy about the following image prefixes: /v2/repos/test-platform/catalyst/pat... [14:59:05] !log sukhe@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage [14:59:24] (03PS4) 10Filippo Giunchedi: sre/cdn: create recording rules in advance of moving to ratio for CDN [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [14:59:25] (03PS3) 10Filippo Giunchedi: sre/cdn: use traffic ratios in ATSBackendErrorsHigh [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [14:59:39] (03Merged) 10jenkins-bot: sre.gitlab.upgrade: extend replica downtime to 60h [cookbooks] - 10https://gerrit.wikimedia.org/r/1331653 (https://phabricator.wikimedia.org/T436361) (owner: 10Jelto) [15:00:47] (03CR) 10Bking: [C:03+2] "self-merging since this is actively incorrect and involves certificates." [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331667 (https://phabricator.wikimedia.org/T436258) (owner: 10Bking) [15:01:34] (03CR) 10Gehel: [C:03+1] "lgtm" [puppet] - 10https://gerrit.wikimedia.org/r/1331647 (https://phabricator.wikimedia.org/T436285) (owner: 10Bking) [15:03:13] !log bking@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [15:03:41] !log bking@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [15:06:27] (03CR) 10Elukey: [C:03+1] Puppet 8: Replace legacy facts in add_ip6_mapped [puppet] - 10https://gerrit.wikimedia.org/r/1329664 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:08:49] (03CR) 10JHathaway: [C:03+2] Puppet 8: Replace legacy facts in add_ip6_mapped [puppet] - 10https://gerrit.wikimedia.org/r/1329664 (https://phabricator.wikimedia.org/T435225) (owner: 10JHathaway) [15:09:44] (03CR) 10Hnowlan: [C:03+1] "One nit but otherwise lgtm!" [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [15:12:36] (03PS5) 10Filippo Giunchedi: sre/cdn: create recording rules in advance of moving to ratio for CDN [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [15:12:36] (03PS4) 10Filippo Giunchedi: sre/cdn: use traffic ratios in ATSBackendErrorsHigh [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [15:13:11] (03CR) 10Filippo Giunchedi: sre/cdn: create recording rules in advance of moving to ratio for CDN (031 comment) [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [15:17:29] PROBLEM - Host pki1002 is DOWN: PING CRITICAL - Packet loss = 100% [15:18:22] (03CR) 10Filippo Giunchedi: [C:03+2] sre/cdn: create recording rules in advance of moving to ratio for CDN [alerts] - 10https://gerrit.wikimedia.org/r/1325528 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [15:18:42] FIRING: [29x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:20:55] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations, 13Patch-For-Review: Move Docker images under the /v2/releng prefix to S3 - https://phabricator.wikimedia.org/T432829#12266243 (10jnuche) >>! In T432829#12266151, @elukey wrote: > @jnuche Hi! I was talking with @dancy about the following image prefi... [15:22:12] (03CR) 10Hnowlan: [C:03+1] "lgtm! One musing, but can be ignored or changed." [alerts] - 10https://gerrit.wikimedia.org/r/1326811 (https://phabricator.wikimedia.org/T400675) (owner: 10Hnowlan) [15:23:08] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations, 13Patch-For-Review: Move Docker images under the /v2/releng prefix to S3 - https://phabricator.wikimedia.org/T432829#12266247 (10elukey) @jnuche thanks for the quick reply! At the moment the bucket gets /v2/repos/releng/ and /v2/releng/, so we can... [15:23:42] FIRING: [44x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:27:42] FIRING: JobUnavailable: Reduced availability for job cfssl in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:32:45] FIRING: WidespreadPuppetFailure: Puppet has failed in magru - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:33:51] uh [15:33:55] https://puppetboard.wikimedia.org/nodes?status=failed [15:34:13] Could not evaluate: Could not retrieve file metadata for http://pki.discovery.wmnet [15:34:17] seems to be the common them [15:34:20] 11:17:29 <+icinga-wm> PROBLEM - Host pki1002 is DOWN: PING CRITICAL - Packet loss = 100% [15:34:23] ah [15:35:56] oh no [15:35:56] Description: CPU 1 machine check error detected. [15:36:54] I will file a task [15:38:03] 10SRE-swift-storage, 10Ceph, 06Infrastructure-Foundations: Move the majority of the Registry's docker image prefixes to a new s3 bucket - https://phabricator.wikimedia.org/T435499#12266301 (10elukey) Very interestingly, the registry clone script proposes to copy 2019 image tags for things like citoid and cxs... [15:38:38] sukhe: oh noes :( it broke also the other day, it seems a recurrent problem [15:38:46] lemme gdns-depool it [15:38:57] elukey: yep, it's an A/A just checked (didn't know) [15:39:04] I am filing the hardware task, will add you to it [15:39:20] sukhe: I think there is already one, lemme find it [15:39:45] https://phabricator.wikimedia.org/T434268 [15:40:00] ah thanks! [15:40:33] so Valerie wanted to perform a firmware upgrade [15:40:58] and yeah, if this was behind LVS, that would be nice for sure [15:40:59] !log elukey@puppetserver1001 conftool action : set/pooled=false; selector: dnsdisc=pki,name=eqiad [15:41:47] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12266320 (10elukey) I just depooled pki1002 since it failed again :( [15:42:12] 10ops-eqiad, 06SRE, 06DC-Ops, 06Infrastructure-Foundations: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268#12266324 (10ssingh) This happened again today (Fri Aug 28) at 15:37:20, `<+icinga-wm> PROBLEM - Host pki1002 is DOWN: PIN... [15:42:14] sukhe: I'll reboot it to have it ready in case something horrible happens during the weekend, and the only remaining pki host gets down as well [15:42:19] <3 [15:42:45] FIRING: [2x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:43:37] !log powercycle pki1002 after CPU-related failures (host completely unresponsive) - T434268 [15:43:39] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [15:43:39] T434268: pki1002 became unresponsive causing several hosts to alert on failed puppet runs. - https://phabricator.wikimedia.org/T434268 [15:46:17] RECOVERY - Host pki1002 is UP: PING OK - Packet loss = 0%, RTA = 0.34 ms [15:47:42] RESOLVED: JobUnavailable: Reduced availability for job cfssl in ops@eqiad - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [15:47:45] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:48:42] RESOLVED: [44x] ProbeDown: Service pki1002:443 has failed probes (http_PKI_aux_front_proxy_ip4) - https://wikitech.wikimedia.org/wiki/Runbook#pki1002:443 - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://alerts.wikimedia.org/?q=alertname%3DProbeDown [15:51:25] FIRING: [24x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki1002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [15:52:45] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [15:54:54] (03CR) 10Atsuko: [C:03+2] Typo in airflow-experiment-platform-ops LDAP group manager list [puppet] - 10https://gerrit.wikimedia.org/r/1331615 (https://phabricator.wikimedia.org/T416709) (owner: 10Atsuko) [15:57:45] FIRING: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [16:00:50] (03CR) 10Ssingh: [C:03+1] site.pp move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1331590 (https://phabricator.wikimedia.org/T436363) (owner: 10Slyngshede) [16:02:45] RESOLVED: [3x] WidespreadPuppetFailure: Puppet has failed in drmrs - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet?orgId=1&viewPanel=6 - https://alerts.wikimedia.org/?q=alertname%3DWidespreadPuppetFailure [16:03:48] !log sudo ipmitool -I lanplus -H "cp5022.mgmt.eqsin.wmnet" -U root -E chassis power cycle: trying to debug why it doesn't come up after reimage [16:03:49] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [16:11:42] FIRING: JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:16:42] FIRING: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:21:42] RESOLVED: [2x] JobUnavailable: Reduced availability for job sidekiq in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [16:26:09] (03CR) 10CDanis: [C:03+2] "I'll do it in a followup, no need for a roundtrip on it" [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330487 (owner: 10Hnowlan) [16:27:14] (03Merged) 10jenkins-bot: Add button for mgmt escalation [software/klaxon] - 10https://gerrit.wikimedia.org/r/1330487 (owner: 10Hnowlan) [16:27:17] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12266439 (10JAATPH) [16:28:12] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12266440 (10JArguello-WMF) [16:28:37] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12266443 (10JArguello-WMF) [16:28:53] 06SRE, 10SRE-Access-Requests: Requesting access to Stat host stat1010 for Jose Aleman - https://phabricator.wikimedia.org/T436298#12266444 (10JArguello-WMF) [16:29:33] (03PS1) 10Cwhite: opensearch: set provide_chain on pki certificates [puppet] - 10https://gerrit.wikimedia.org/r/1331697 (https://phabricator.wikimedia.org/T350516) [16:31:18] 10ops-eqsin, 06Infrastructure-Foundations: attempts to reimage cp5022 are failing - https://phabricator.wikimedia.org/T436386 (10CDobbins) 03NEW [16:32:09] (03CR) 10Cwhite: [C:03+2] opensearch: set provide_chain on pki certificates [puppet] - 10https://gerrit.wikimedia.org/r/1331697 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [16:34:03] (03PS1) 10CDanis: fix: typo [software/klaxon] - 10https://gerrit.wikimedia.org/r/1331700 [16:34:20] (03CR) 10CDanis: [C:03+2] fix: typo [software/klaxon] - 10https://gerrit.wikimedia.org/r/1331700 (owner: 10CDanis) [16:35:29] (03Merged) 10jenkins-bot: fix: typo [software/klaxon] - 10https://gerrit.wikimedia.org/r/1331700 (owner: 10CDanis) [16:38:32] !log sukhe@cumin1003 END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie [16:38:49] 10ops-eqsin, 06SRE, 06DC-Ops, 06Traffic, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12266478 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by sukhe@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (**F... [16:41:07] (03PS1) 10Cwhite: Revert "opensearch: set provide_chain on pki certificates" [puppet] - 10https://gerrit.wikimedia.org/r/1331703 [16:45:30] (03CR) 10Jasmine: [C:03+1] "Thanks Simon!" [cookbooks] - 10https://gerrit.wikimedia.org/r/1331446 (https://phabricator.wikimedia.org/T433363) (owner: 10Slyngshede) [16:51:25] RESOLVED: [24x] SystemdUnitFailed: cfssl-ocsprefresh-Wikimedia_Internal_Root_CA.service on pki1002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [17:03:12] (03CR) 10Cwhite: [C:03+2] Revert "opensearch: set provide_chain on pki certificates" [puppet] - 10https://gerrit.wikimedia.org/r/1331703 (owner: 10Cwhite) [17:06:09] (03CR) 10Bking: [C:03+2] cirrussearch: Add new hosts to site.pp [puppet] - 10https://gerrit.wikimedia.org/r/1331647 (https://phabricator.wikimedia.org/T436285) (owner: 10Bking) [17:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [17:24:33] 06SRE, 10Internet-Archive, 06Traffic, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12266567 (10AlexisJazz) >>! In T435743#12258203, @ssingh wrote: > I have tried to reproduce now with some obscure paths as well but "Save Page" is working fine. @AlexisJazz: ca... [17:32:36] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm [17:50:13] !log eevans@cumin1003 END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cassandra-dev2003.codfw.wmnet with OS bookworm [17:50:35] !log eevans@cumin1003 START - Cookbook sre.hosts.provision for host cassandra-dev2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [17:52:55] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL [17:53:25] !log eevans@cumin1003 START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm [17:55:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [17:59:19] 06SRE, 10Pywikibot, 06Traffic, 10Wikidata, and 3 others: Pywikibot reports maxlag retry error on Wikidata - https://phabricator.wikimedia.org/T421642#12266660 (10Xqt) @Epidosis: The impact is not solely for Wikidata. [18:03:14] (03PS2) 10Slyngshede: site.pp move cp3074 from upload to text [puppet] - 10https://gerrit.wikimedia.org/r/1331590 (https://phabricator.wikimedia.org/T436363) [18:09:56] !log eevans@cumin1003 START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage [18:13:43] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage [18:33:38] !log eevans@cumin1003 END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm [18:41:16] (03PS1) 10Andrew Bogott: Add cloudvps-tenant-usage.py [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) [18:41:52] (03CR) 10CI reject: [V:04-1] Add cloudvps-tenant-usage.py [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) (owner: 10Andrew Bogott) [18:46:05] (03PS2) 10Andrew Bogott: Add cloudvps-tenant-usage.py [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) [18:50:29] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) (owner: 10Andrew Bogott) [18:53:14] (03PS3) 10Andrew Bogott: Add cloudvps-tenant-usage.py [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) [18:53:20] (03CR) 10Andrew Bogott: "check experimental" [puppet] - 10https://gerrit.wikimedia.org/r/1331749 (https://phabricator.wikimedia.org/T436275) (owner: 10Andrew Bogott) [18:56:09] (03CR) 10BCornwall: [C:03+1] etherpad-next: add discovery ingress records [dns] - 10https://gerrit.wikimedia.org/r/1331613 (https://phabricator.wikimedia.org/T435509) (owner: 10Jelto) [19:04:06] (03PS1) 10Lerickson: Update default Eventgate URL value in the WDQS chart. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331764 (https://phabricator.wikimedia.org/T433375) [19:08:59] (03CR) 10Lerickson: "Hey, feel free to merge this in my absence next week if you'd like to. Also OK to wait till I'm back. This is just cleanup that doesn't af" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331764 (https://phabricator.wikimedia.org/T433375) (owner: 10Lerickson) [19:09:02] (03PS1) 10Cwhite: cfssl: support transforming the certificate key to PKCS#8 format [puppet] - 10https://gerrit.wikimedia.org/r/1331768 (https://phabricator.wikimedia.org/T436393) [19:09:22] (03PS2) 10Cwhite: cfssl: support transforming the certificate key to PKCS#8 format [puppet] - 10https://gerrit.wikimedia.org/r/1331768 (https://phabricator.wikimedia.org/T436393) [19:30:33] (03PS1) 10Cwhite: cfssl: update cert chain to optionally include the root ca [puppet] - 10https://gerrit.wikimedia.org/r/1331783 (https://phabricator.wikimedia.org/T436394) [19:31:30] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [19:33:06] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [19:33:13] (03PS2) 10Cwhite: cfssl: update cert chain to optionally include the root ca [puppet] - 10https://gerrit.wikimedia.org/r/1331783 (https://phabricator.wikimedia.org/T436394) [19:34:58] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [19:35:44] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [19:36:19] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [19:36:26] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [19:42:00] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [19:43:21] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [19:44:08] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [19:45:21] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [19:45:47] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [19:47:01] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [19:50:35] (03PS1) 10Lerickson: Allow up to 3 redirects for federated endpoints. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331787 (https://phabricator.wikimedia.org/T433133) [19:53:13] (03PS3) 10Cwhite: cfssl: support transforming the certificate key to PKCS#8 format [puppet] - 10https://gerrit.wikimedia.org/r/1331768 (https://phabricator.wikimedia.org/T436393) [20:03:33] (03CR) 10Lerickson: "I deployed this to staging (homedir) to check that it worked (it did!), and then put staging back. See the linked phab for a test query to" [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331787 (https://phabricator.wikimedia.org/T433133) (owner: 10Lerickson) [20:04:24] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [20:05:38] !log lerickson@deploy1003 helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [20:12:57] (03PS1) 10Ryan Kemper: cirrussearch: Prepare codfw expansion [puppet] - 10https://gerrit.wikimedia.org/r/1331792 (https://phabricator.wikimedia.org/T436287) [20:13:08] (03PS1) 10Ryan Kemper: cirrussearch: Prepare eqiad expansion [puppet] - 10https://gerrit.wikimedia.org/r/1331793 (https://phabricator.wikimedia.org/T436285) [20:13:19] (03CR) 10CI reject: [V:04-1] cirrussearch: Prepare codfw expansion [puppet] - 10https://gerrit.wikimedia.org/r/1331792 (https://phabricator.wikimedia.org/T436287) (owner: 10Ryan Kemper) [20:14:41] (03PS2) 10Lerickson: Allow up to 3 redirects for federated endpoints. [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331787 (https://phabricator.wikimedia.org/T433133) [20:18:37] (03Abandoned) 10Ryan Kemper: cirrussearch: Prepare eqiad expansion [puppet] - 10https://gerrit.wikimedia.org/r/1331793 (https://phabricator.wikimedia.org/T436285) (owner: 10Ryan Kemper) [20:18:47] (03Abandoned) 10Ryan Kemper: cirrussearch: Prepare codfw expansion [puppet] - 10https://gerrit.wikimedia.org/r/1331792 (https://phabricator.wikimedia.org/T436287) (owner: 10Ryan Kemper) [20:20:29] (03PS1) 10Ryan Kemper: cirrussearch: Fix codfw task reference [puppet] - 10https://gerrit.wikimedia.org/r/1331801 (https://phabricator.wikimedia.org/T436287) [20:20:49] (03PS2) 10Ryan Kemper: cirrussearch: Fix codfw task reference [puppet] - 10https://gerrit.wikimedia.org/r/1331801 (https://phabricator.wikimedia.org/T436287) [20:21:17] (03PS3) 10Ryan Kemper: cirrussearch: Fix codfw task reference [puppet] - 10https://gerrit.wikimedia.org/r/1331801 (https://phabricator.wikimedia.org/T436287) [20:21:58] (03CR) 10Bking: [C:03+1] cirrussearch: Fix codfw task reference [puppet] - 10https://gerrit.wikimedia.org/r/1331801 (https://phabricator.wikimedia.org/T436287) (owner: 10Ryan Kemper) [20:24:14] (03CR) 10Ryan Kemper: [C:03+2] cirrussearch: Fix codfw task reference [puppet] - 10https://gerrit.wikimedia.org/r/1331801 (https://phabricator.wikimedia.org/T436287) (owner: 10Ryan Kemper) [20:35:04] (03PS4) 10Cwhite: cfssl: support transforming the certificate key to PKCS#8 format [puppet] - 10https://gerrit.wikimedia.org/r/1331768 (https://phabricator.wikimedia.org/T436393) [20:36:35] (03CR) 10Cwhite: cfssl: support transforming the certificate key to PKCS#8 format (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1331768 (https://phabricator.wikimedia.org/T436393) (owner: 10Cwhite) [20:53:49] (03CR) 10Ryan Kemper: [C:03+2] wdqs: drop dangling query-legacy-full cert SAN [deployment-charts] - 10https://gerrit.wikimedia.org/r/1278562 (https://phabricator.wikimedia.org/T415073) (owner: 10Ryan Kemper) [21:03:20] (03Merged) 10jenkins-bot: wdqs: drop dangling query-legacy-full cert SAN [deployment-charts] - 10https://gerrit.wikimedia.org/r/1278562 (https://phabricator.wikimedia.org/T415073) (owner: 10Ryan Kemper) [21:05:45] !log T415073 beginning admin_ng namespace-certificates rollout (gerrit 1278562, drop query-legacy-full.wikidata.org SAN from wikidata-query-gui cert), starting with staging [21:05:47] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:05:48] T415073: Cleanup after decommission of the WDQS full graph endpoint - https://phabricator.wikimedia.org/T415073 [21:06:22] !log ryankemper@deploy1003 helmfile [staging-codfw] START helmfile.d/admin 'apply'. [21:06:43] !log ryankemper@deploy1003 helmfile [staging-codfw] DONE helmfile.d/admin 'apply'. [21:06:51] FIRING: NetworkDeviceAlarmActive: Alarm active on cr1-eqiad - https://wikitech.wikimedia.org/wiki/Network_monitoring#Juniper_alarm - https://grafana.wikimedia.org/d/fb403d62-5f03-434a-9dff-bd02b9fff504/network-device-overview?var-instance=cr1-eqiad:9804 - https://alerts.wikimedia.org/?q=alertname%3DNetworkDeviceAlarmActive [21:07:05] !log ryankemper@deploy1003 helmfile [staging-eqiad] START helmfile.d/admin 'apply'. [21:07:18] !log ryankemper@deploy1003 helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'. [21:07:28] !log ryankemper@deploy1003 helmfile [codfw] START helmfile.d/admin 'apply'. [21:07:44] !log ryankemper@deploy1003 helmfile [codfw] DONE helmfile.d/admin 'apply'. [21:07:56] !log ryankemper@deploy1003 helmfile [eqiad] START helmfile.d/admin 'apply'. [21:08:05] !log ryankemper@deploy1003 helmfile [eqiad] DONE helmfile.d/admin 'apply'. [21:11:19] (03PS1) 10C. Scott Ananian: Keep Balinese Palm Leaf variants enabled on wikisource [mediawiki-config] - 10https://gerrit.wikimedia.org/r/1331830 (https://phabricator.wikimedia.org/T436398) [21:16:21] !log T415073 admin_ng namespace-certificates applied on `staging-codfw`, `staging-eqiad`, `codfw`, `eqiad`; `wikidata-query-gui` cert reissued without `query-legacy-full.wikidata.org` SAN, verified on the wire in both DCs [21:16:24] Logged the message at https://wikitech.wikimedia.org/wiki/Server_Admin_Log [21:16:24] T415073: Cleanup after decommission of the WDQS full graph endpoint - https://phabricator.wikimedia.org/T415073 [21:39:50] (03CR) 10Ryan Kemper: "This change is guaranteed to be safe per the following, so I'm just going to merge this to get it out of my queue:" [puppet] - 10https://gerrit.wikimedia.org/r/1328211 (https://phabricator.wikimedia.org/T419831) (owner: 10Ryan Kemper) [21:39:56] (03CR) 10Ryan Kemper: [C:03+2] cumin: drop duplicate mediabackup-storage alias [puppet] - 10https://gerrit.wikimedia.org/r/1328211 (https://phabricator.wikimedia.org/T419831) (owner: 10Ryan Kemper) [21:55:38] FIRING: CertAlmostExpired: gNMI TLS certificate for lsw1-d6-eqiad.mgmt.eqiad.wmnet is going to expire in 0s - https://wikitech.wikimedia.org/wiki/Network_monitoring#CertAlmostExpired - https://grafana.wikimedia.org/d/eab73c60-a402-4f9b-a4a7-ea489b374458/gnmic?var-site=eqiad - https://alerts.wikimedia.org/?q=alertname%3DCertAlmostExpired [22:14:45] (03PS1) 10Scott French: Rakefile: Fix 'upstream' mock in services_proxy data [deployment-charts] - 10https://gerrit.wikimedia.org/r/1331857 (https://phabricator.wikimedia.org/T427666) [22:55:45] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:55:45] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:55:45] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:56:39] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:56:39] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:56:39] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 3.083 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:59:45] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [22:59:45] PROBLEM - SSH on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/SSH/monitoring [22:59:45] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [23:00:17] (03CR) 10Cwhite: cfssl: support transforming the certificate key to PKCS#8 format (031 comment) [puppet] - 10https://gerrit.wikimedia.org/r/1331768 (https://phabricator.wikimedia.org/T436393) (owner: 10Cwhite) [23:02:47] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 9.511 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [23:03:42] FIRING: JobUnavailable: Reduced availability for job atlas_exporter in ops@codfw - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/NEJu05xZz/prometheus-targets - https://alerts.wikimedia.org/?q=alertname%3DJobUnavailable [23:04:11] (03PS12) 10Cwhite: cfssl: update cert chain to optionally include the root ca [puppet] - 10https://gerrit.wikimedia.org/r/1331783 (https://phabricator.wikimedia.org/T436394) [23:05:35] (03CR) 10Cwhite: [C:03+2] beta-logs: install CA root [puppet] - 10https://gerrit.wikimedia.org/r/1330687 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:05:47] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [23:09:37] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [23:09:37] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 1.005 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [23:12:47] PROBLEM - librenms.wikimedia.org tls expiry on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [23:12:47] PROBLEM - librenms.wikimedia.org requires authentication on netmon2002 is CRITICAL: CRITICAL - Socket timeout after 10 seconds https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [23:12:54] (03PS5) 10Cwhite: opensearch: expose a chained certificate [puppet] - 10https://gerrit.wikimedia.org/r/1329379 (https://phabricator.wikimedia.org/T350516) [23:18:17] (03CR) 10Cwhite: [C:03+2] opensearch: expose a chained certificate [puppet] - 10https://gerrit.wikimedia.org/r/1329379 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:18:39] (03PS4) 10Cwhite: beta-logs: add logstash_real role [puppet] - 10https://gerrit.wikimedia.org/r/1330688 (https://phabricator.wikimedia.org/T350516) [23:21:18] (03PS5) 10Cwhite: beta-logs: add logstash_real role [puppet] - 10https://gerrit.wikimedia.org/r/1330688 (https://phabricator.wikimedia.org/T350516) [23:34:42] (03PS1) 10Cwhite: opensearch: try again to construct a certificate chain [puppet] - 10https://gerrit.wikimedia.org/r/1331897 (https://phabricator.wikimedia.org/T350516) [23:34:52] (03CR) 10Cwhite: [C:03+2] beta-logs: add logstash_real role [puppet] - 10https://gerrit.wikimedia.org/r/1330688 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:37:10] (03CR) 10Cwhite: [C:03+2] opensearch: try again to construct a certificate chain [puppet] - 10https://gerrit.wikimedia.org/r/1331897 (https://phabricator.wikimedia.org/T350516) (owner: 10Cwhite) [23:55:37] RECOVERY - librenms.wikimedia.org tls expiry on netmon2002 is OK: OK - Certificate librenms.wikimedia.org will expire on Mon 09 Nov 2026 02:18:41 AM GMT +0000. https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration [23:55:37] RECOVERY - SSH on netmon2002 is OK: SSH OK - OpenSSH_9.2p1 Debian-2+deb12u10 (protocol 2.0) https://wikitech.wikimedia.org/wiki/SSH/monitoring [23:55:37] RECOVERY - librenms.wikimedia.org requires authentication on netmon2002 is OK: HTTP OK: Status line output matched HTTP/1.1 302 - 701 bytes in 0.131 second response time https://wikitech.wikimedia.org/wiki/CAS-SSO/Administration