[00:04:16] 10Cloud-VPS, 06tools-infrastructure-team, 06tools-platform-team, 10Maps: wmf-auto-restart[2162435]: INFO: Could not query the PID(s) of exim4: -11 - https://phabricator.wikimedia.org/T435262#12229230 (10saper) Reported to Debian as https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1144806 Do we really ne... [00:10:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [00:15:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [00:20:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [01:52:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [01:57:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:02:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:09:56] !log andrew@cloudcumin1001 admin START - Cookbook wmcs.ceph.set_cluster_in_maintenance [02:10:00] !log andrew@cloudcumin1001 admin END (PASS) - Cookbook wmcs.ceph.set_cluster_in_maintenance (exit_code=0) [02:12:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:15:21] PROBLEM - Host cloudcephosd1042 is DOWN: PING CRITICAL - Packet loss = 100% [02:17:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [02:17:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:18:47] FIRING: NodeDown: Node cloudcephosd1042 is down. - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/NodeDown - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&var-server=cloudcephosd1042 - https://alerts.wikimedia.org/?q=alertname%3DNodeDown [02:20:23] RECOVERY - Host cloudcephosd1042 is UP: PING OK - Packet loss = 0%, RTA = 0.24 ms [02:22:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:23:47] RESOLVED: NodeDown: Node cloudcephosd1042 is down. - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/NodeDown - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&var-server=cloudcephosd1042 - https://alerts.wikimedia.org/?q=alertname%3DNodeDown [02:26:35] PROBLEM - Host cloudcephosd1041 is DOWN: PING CRITICAL - Packet loss = 100% [02:29:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:31:17] FIRING: [2x] NodeDown: Node cloudcephosd1041 is down. - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/NodeDown - https://alerts.wikimedia.org/?q=alertname%3DNodeDown [02:39:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:39:50] RECOVERY - Host cloudcephosd1041 is UP: PING OK - Packet loss = 0%, RTA = 2.13 ms [02:41:17] RESOLVED: NodeDown: Node cloudcephosd1041 is down. - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/NodeDown - https://grafana.wikimedia.org/d/000000377/host-overview?orgId=1&var-server=cloudcephosd1041 - https://alerts.wikimedia.org/?q=alertname%3DNodeDown [02:44:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:49:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [02:51:07] !log andrew@cloudcumin1001 admin START - Cookbook wmcs.ceph.unset_cluster_maintenance [02:51:07] !log andrew@cloudcumin1001 admin END (PASS) - Cookbook wmcs.ceph.unset_cluster_maintenance (exit_code=0) [02:57:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [03:54:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [03:59:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [04:04:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [04:09:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:08:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:13:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:23:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:28:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:33:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:50:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:55:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [06:05:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [07:03:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [07:08:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [07:32:11] (03merge) 10samwilson: Update GitLab CI from Debian bookworm to Debian trixie [toolforge-repos/ocr] - 10https://gitlab.wikimedia.org/toolforge-repos/ocr/-/merge_requests/7 (owner: 10sweil) [08:03:49] (03open) 10samwilson: Update CONTRIBUTING.md [toolforge-repos/ocr] - 10https://gitlab.wikimedia.org/toolforge-repos/ocr/-/merge_requests/13 [08:05:26] (03update) 10dcaro: tests: add behavior tests for previously untested code branches [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/362 (https://phabricator.wikimedia.org/T434166) (owner: 10raymond-ndibe) [08:09:23] (03merge) 10samwilson: Update CONTRIBUTING.md [toolforge-repos/ocr] - 10https://gitlab.wikimedia.org/toolforge-repos/ocr/-/merge_requests/13 [08:23:49] (03PS10) 10Krinkle: [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) [08:24:21] (03CR) 10CI reject: [V:04-1] [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) (owner: 10Krinkle) [08:50:57] (03update) 10raymond-ndibe: models: support webservice job-type [repos/cloud/toolforge/jobs-api] (models_refactor) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/264 (https://phabricator.wikimedia.org/T428898) [09:40:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [09:43:34] 10Data-Services, 06tools-platform-team, 06Data-Persistence: Make sure wmf-pt-kill service waits for mariadb to be ready - https://phabricator.wikimedia.org/T435178#12230094 (10Marostegui) >>! In T435178#12226003, @fgiunchedi wrote: > What do you think of having wmf-pt-kill start alongside other "sidecars" li... [09:45:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [09:54:18] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12230109 (10fgiunchedi) Ok I re-ran the provision cookbook and after a few retries it worked for cloudvirt1055 / cloudvi... [10:10:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [10:16:56] 10PAWS, 06tools-platform-team: PAWS notebook error - https://phabricator.wikimedia.org/T435304 (10DreamRimmer) 03NEW [10:21:11] (03approved) 10dcaro: [T398425] Treat invalid YAML as a user error [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/91 (https://phabricator.wikimedia.org/T398425) (owner: 10mahveotm) [10:27:18] (03update) 10dcaro: [T398425] Treat invalid YAML as a user error [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/91 (https://phabricator.wikimedia.org/T398425) (owner: 10mahveotm) [10:28:39] (03merge) 10dcaro: [T398425] Treat invalid YAML as a user error [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/91 (https://phabricator.wikimedia.org/T398425) (owner: 10mahveotm) [10:29:51] (03open) 10dcaro: d/changelog: bump to 0.0.19 [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/93 (https://phabricator.wikimedia.org/T398425) [10:30:44] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component components-cli [10:31:04] 10Cloud-VPS, 06tools-infrastructure-team: Fix pacct rotation properly everywhere - https://phabricator.wikimedia.org/T410410#12230328 (10taavi) a:03taavi [10:33:56] !log dcaro@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-cli [10:35:44] 10Toolforge, 06tools-platform-team, 07OKR-Work: Define alert conditions in Prometheus (if possible) - https://phabricator.wikimedia.org/T432863#12230357 (10fnegri) I agree on the three principles, and I think having a dedicated alertmanager instance that can be deployed in lima-kilo is the desirable end stat... [10:37:36] !log dcaro@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component components-cli [10:39:52] 10Toolforge, 06tools-platform-team, 07OKR-Work: Set up a Toolforge “alerting service” that emails tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865#12230378 (10dcaro) I would lean towards the service polling alertmanager, to avoid having to "teach" alermanager t... [10:40:36] 10Toolforge, 06tools-platform-team, 07OKR-Work: Set up a Toolforge “alerting service” that emails tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865#12230380 (10dcaro) >>! In T432865#12230378, @dcaro wrote: > I would lean towards the service polling alertmanager,... [10:40:53] !log dcaro@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-cli [10:42:59] (03approved) 10dcaro: d/changelog: bump to 0.0.19 [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/93 (https://phabricator.wikimedia.org/T398425) [10:43:06] (03merge) 10dcaro: d/changelog: bump to 0.0.19 [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/93 (https://phabricator.wikimedia.org/T398425) [10:46:40] 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 06Infrastructure-Foundations, 10Puppet-Core: Normalise hiera default values - https://phabricator.wikimedia.org/T289665#12230417 (10LSobanski) p:05Medium→03Triage [10:47:09] 06cloud-services-team, 10Toolforge, 06tools-platform-team, 13Patch-For-Review: [components-cli] Invalid YAML file error should not encourage reporting the issue to admins - https://phabricator.wikimedia.org/T398425#12230423 (10dcaro) 05In progress→03Resolved Deployed \o/ [10:48:00] 10Cloud Services Proposals, 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, and 3 others: Easing pain points caused by divergence between cloudservices and production puppet usecases - https://phabricator.wikimedia.org/T285539#12230443 (10LSobanski) p:05Medium→03Triage [10:50:01] (03update) 10raymond-ndibe: test continuous job webservice [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1131 (https://phabricator.wikimedia.org/T348755) [10:51:25] (03update) 10sweil: Update package-lock.json [toolforge-repos/wiki-ocr] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-ocr/-/merge_requests/1 [10:58:33] RESOLVED: PuppetAgentStaleLastRun: Last Puppet run was over 24 hours ago on instance proxy-6 in project project-proxy - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetAgentStaleLastRun [11:04:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:09:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:12:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:17:24] 10Toolforge, 06tools-platform-team, 07OKR-Work: Define alert conditions in Prometheus (if possible) - https://phabricator.wikimedia.org/T432863#12230604 (10taavi) The metricsinfra AM deployment is already highly available and designed to be used by multiple projects with different alerting needs and targets.... [11:17:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:20:30] (03update) 10raymond-ndibe: test continuous job webservice [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1131 (https://phabricator.wikimedia.org/T348755) [11:21:12] 10Toolforge, 06tools-platform-team, 07OKR-Work: Define alert conditions in Prometheus (if possible) - https://phabricator.wikimedia.org/T432863#12230631 (10fnegri) > I guess you could run Alertmanager in Kubernetes, although that still leaves the Prometheus server which is at the moment in production best de... [11:22:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:23:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:25:35] 10Data-Services, 06tools-platform-team, 06Data-Persistence: Make sure wmf-pt-kill service waits for mariadb to be ready - https://phabricator.wikimedia.org/T435178#12230649 (10fgiunchedi) >>! In T435178#12230094, @Marostegui wrote: >>>! In T435178#12226003, @fgiunchedi wrote: >> What do you think of having w... [11:28:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:30:13] (03update) 10raymond-ndibe: test continuous job webservice [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1131 (https://phabricator.wikimedia.org/T348755) [11:31:21] 10Cloud-VPS, 06tools-infrastructure-team: Fix pacct rotation properly everywhere - https://phabricator.wikimedia.org/T410410#12230655 (10taavi) 05Open→03Resolved [11:49:21] 10Tool-inteGraality: Reference column drill-down queries time out on QLever - https://phabricator.wikimedia.org/T435101#12230729 (10JeanFred) [11:50:03] FIRING: PuppetSyncFailure: Failed to update Puppet repository /srv/git/operations/puppet on instance metricsinfra-puppetserver-1 in project metricsinfra - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetSyncFailure [11:52:58] (03update) 10raymond-ndibe: support --webservice option [repos/cloud/toolforge/jobs-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/merge_requests/143 (https://phabricator.wikimedia.org/T428898) [11:55:03] RESOLVED: PuppetSyncFailure: Failed to update Puppet repository /srv/git/operations/puppet on instance metricsinfra-puppetserver-1 in project metricsinfra - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetSyncFailure [11:57:33] 10Toolforge, 06tools-platform-team, 07OKR-Work: [logs-api] Implement an endpoint to retrieve system-level event logs - https://phabricator.wikimedia.org/T432566#12230755 (10TLepage-WMF) 05Open→03In progress a:03TLepage-WMF [12:07:44] FIRING: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [12:10:24] (03approved) 10dcaro: tests: add behavior tests for previously untested code branches [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/362 (https://phabricator.wikimedia.org/T434166) (owner: 10raymond-ndibe) [12:10:26] (03update) 10dcaro: tests: add behavior tests for previously untested code branches [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/362 (https://phabricator.wikimedia.org/T434166) (owner: 10raymond-ndibe) [12:20:25] 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 10Observability-Metrics, and 2 others: Modernise memcached systemd unit / sync, and make it presentable - https://phabricator.wikimedia.org/T273950#12230867 (10taavi) Tagging what I believe are the right projects for the two remaining profiles. [12:23:15] (03update) 10raymond-ndibe: test continuous job webservice [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1131 (https://phabricator.wikimedia.org/T348755) [12:32:44] RESOLVED: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [12:33:35] (03open) 10dcaro: dont allow extras by default [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/369 (https://phabricator.wikimedia.org/T434166) [12:33:55] 10Tool-wikiloves: Add Wiki Loves Villages into the Wikiloves statistics tool - https://phabricator.wikimedia.org/T435316#12230950 (10Nemoralis) [12:34:09] (03update) 10dcaro: core.models: declare the models not accepting extras [repos/cloud/toolforge/jobs-api] (rename_common_job_to_common_options) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/369 (https://phabricator.wikimedia.org/T434166) [12:34:46] (03update) 10dcaro: core.models: declare the models not accepting extras [repos/cloud/toolforge/jobs-api] (rename_common_job_to_common_options) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/369 (https://phabricator.wikimedia.org/T434166) [12:38:54] (03update) 10dcaro: refactor: move observability fields out of CommonOptions into ObservabilityOptions [repos/cloud/toolforge/jobs-api] (rename_common_job_to_common_options) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/363 (https://phabricator.wikimedia.org/T434166) (owner: 10raymond-ndibe) [12:53:25] (03open) 10mahveotm: maintain-harbor: ignore volatile retention schedule time [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1382 (https://phabricator.wikimedia.org/T393878) [12:57:05] 06cloud-services-team, 10Toolforge, 06tools-platform-team, 13Patch-For-Review: [functional-tests] maintain-harbor tests are a bit flaky - https://phabricator.wikimedia.org/T393878#12231028 (10Mahveotm) 05Open→03In progress [13:05:59] 06cloud-services-team, 10wikitech.wikimedia.org, 10Observability-Alerting, 13Patch-For-Review: Move wikitech-static monitoring off Icinga - https://phabricator.wikimedia.org/T362397#12231156 (10fgiunchedi) 05Open→03Resolved a:03fgiunchedi I'm calling this one done -- we're monitoring the wikitech... [13:08:01] 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: openstack: alert for cloudvirts without aggregate or with unexpected set of them - https://phabricator.wikimedia.org/T284747#12231175 (10fgiunchedi) 05Open→03Resolved a:03fgiunchedi This is done, we're now alerting on cloudvirt disabled... [13:15:36] (03update) 10dcaro: refactor: move observability fields out of CommonOptions into ObservabilityOptions [repos/cloud/toolforge/jobs-api] (rename_common_job_to_common_options) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/363 (https://phabricator.wikimedia.org/T434166) (owner: 10raymond-ndibe) [13:16:50] 06cloud-services-team, 10Toolforge, 06tools-platform-team: [builds-builder] Upgrade to "heroku-26" stack - https://phabricator.wikimedia.org/T431127#12231186 (10dcaro) [13:18:57] 10PAWS, 06tools-platform-team: PAWS notebook error - https://phabricator.wikimedia.org/T435304#12231191 (10dcaro) [13:19:14] 10PAWS, 06tools-platform-team: PAWS notebook error - https://phabricator.wikimedia.org/T435304#12231193 (10dcaro) 05Open→03Resolved p:05Triage→03High a:03dcaro Thanks @Max! Yes, this was a false positive by the autoflagging system, the account has been restored. [13:19:44] 10PAWS, 06tools-platform-team: [paws] false positive account flagged as misbehaving - https://phabricator.wikimedia.org/T435304#12231200 (10dcaro) [13:20:52] 10Cloud-VPS, 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team, and 3 others: Move WMCS off of Icinga and introduce alertmanager - https://phabricator.wikimedia.org/T328502#12231204 (10fgiunchedi) Ok we're very close to the finish line! What's left are the DNS checks from production, which hi... [13:21:50] 10Toolforge, 06tools-platform-team, 07OKR-Work: Email tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865#12231208 (10fnegri) [13:31:18] (03update) 10raymond-ndibe: test continuous job webservice [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1131 (https://phabricator.wikimedia.org/T348755) [13:43:37] (03update) 10dcaro: maintain-harbor: ignore volatile retention schedule time [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1382 (https://phabricator.wikimedia.org/T393878) (owner: 10mahveotm) [13:44:31] 10Toolforge, 06tools-platform-team, 13Patch-For-Review: [components-api] split cancel endpoint to avoid duplicated operation id - https://phabricator.wikimedia.org/T434499#12231291 (10dcaro) p:05Triage→03Medium [13:44:38] 10Toolforge (Push-to-Deploy), 06tools-platform-team, 13Patch-For-Review: [jobs-api] refactor models in preparation for webservice job-type - https://phabricator.wikimedia.org/T434166#12231293 (10dcaro) p:05Triage→03High [13:45:10] 06cloud-services-team, 10Toolforge, 06tools-platform-team, 13Patch-For-Review: [functional-tests] maintain-harbor tests are a bit flaky - https://phabricator.wikimedia.org/T393878#12231297 (10dcaro) a:03Mahveotm [13:45:32] 10Cloud-VPS, 06tools-infrastructure-team, 10Maps: wmf-auto-restart[2162435]: INFO: Could not query the PID(s) of exim4: -11 - https://phabricator.wikimedia.org/T435262#12231299 (10dcaro) [13:52:03] 10VPS-project-Codesearch, 07Epic: make a new frontend for codesearch - https://phabricator.wikimedia.org/T435106#12231365 (10ekrem) p:05Triage→03Low [13:55:05] (03open) 10dcaro: gen: update toolforge models [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/183 [13:57:06] (03update) 10raymond-ndibe: test continuous job webservice [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1131 (https://phabricator.wikimedia.org/T348755) [14:07:02] 10Toolforge, 06tools-platform-team: [toolsdb] 2026-08-18 Transaction History Length growing too much - https://phabricator.wikimedia.org/T435220#12231445 (10fnegri) p:05Triage→03High [14:18:43] 06cloud-services-team, 10Toolforge, 06tools-platform-team: [builds-builder] Upgrade to "heroku-26" stack - https://phabricator.wikimedia.org/T431127#12231517 (10fnegri) p:05Triage→03Low [14:20:56] 10Data-Services, 06tools-infrastructure-team, 06tools-platform-team, 06Data-Persistence: Make sure wmf-pt-kill service waits for mariadb to be ready - https://phabricator.wikimedia.org/T435178#12231526 (10fnegri) [14:21:48] 10Quarry, 06tools-platform-team: quarry: Upgrade Python libraries - https://phabricator.wikimedia.org/T397331#12231529 (10fnegri) p:05Triage→03Medium [14:24:45] (03open) 10filippo: Fix task template division by zero [toolforge-repos/cloudvps-quota] - 10https://gitlab.wikimedia.org/toolforge-repos/cloudvps-quota/-/merge_requests/2 (https://phabricator.wikimedia.org/T433836) [14:24:57] (03open) 10filippo: Fix task template division by zero [toolforge-repos/cloudvps-quota] - 10https://gitlab.wikimedia.org/toolforge-repos/cloudvps-quota/-/merge_requests/3 (https://phabricator.wikimedia.org/T433836 https://phabricator.wikimedia.org/T433837) [14:25:19] (03update) 10filippo: Fix task template division by zero [toolforge-repos/cloudvps-quota] - 10https://gitlab.wikimedia.org/toolforge-repos/cloudvps-quota/-/merge_requests/2 (https://phabricator.wikimedia.org/T433836) [14:25:21] (03close) 10filippo: Fix task template division by zero [toolforge-repos/cloudvps-quota] - 10https://gitlab.wikimedia.org/toolforge-repos/cloudvps-quota/-/merge_requests/2 (https://phabricator.wikimedia.org/T433836) [14:26:11] (03update) 10filippo: Fix task template division by zero and API failure display [toolforge-repos/cloudvps-quota] - 10https://gitlab.wikimedia.org/toolforge-repos/cloudvps-quota/-/merge_requests/3 (https://phabricator.wikimedia.org/T433836 https://phabricator.wikimedia.org/T433837) [14:30:11] 10Toolforge, 06tools-platform-team: Deploying built static apps to toolforge - https://phabricator.wikimedia.org/T435088#12231591 (10fnegri) p:05Triage→03Low [14:31:19] (03update) 10raymond-ndibe: test continuous job webservice [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1131 (https://phabricator.wikimedia.org/T348755) [14:36:32] 10Toolforge, 06tools-platform-team: Deploying built static apps to toolforge - https://phabricator.wikimedia.org/T435088#12231636 (10fnegri) As a workaround, you should be able to add uwsgi following the guide at https://wikitech-static.wikimedia.org/wiki/Help_Toolforge/My_first_static_tool.html#Step_2:_Instal... [14:37:32] 10PAWS, 06tools-platform-team: Notebook error - https://phabricator.wikimedia.org/T435068#12231642 (10fnegri) p:05Triage→03Low [14:46:34] 06cloud-services-team, 10Toolforge, 06tools-platform-team: [toolsdb] Add db-level and user-level monitoring - https://phabricator.wikimedia.org/T428087#12231676 (10dcaro) [14:50:03] 06cloud-services-team, 10PAWS, 06tools-platform-team: PAWS: restore nbgitpuller — removing it was never necessary - https://phabricator.wikimedia.org/T434973#12231694 (10fnegri) p:05Triage→03Low PAWS is not currently a priority for the #tools-platform-team, marking as Low. This can be revisited in the fu... [15:08:41] FIRING: CloudVPSDesignateLeaks: Detected 2 stray dns records - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/Designate_record_leaks - https://grafana.wikimedia.org/d/ebJoA6VWz/wmcs-openstack-eqiad-nova-fullstack - https://alerts.wikimedia.org/?q=alertname%3DCloudVPSDesignateLeaks [15:15:55] 10Toolforge, 06tools-platform-team, 07OKR-Work: Email tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865#12232019 (10fnegri) I'm honestly still undecided between the 3 possible implementations (webhook, polling, no service). I tried summarizing the pros/cons belo... [15:17:27] 10Toolforge, 06tools-platform-team, 07OKR-Work: Email tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865#12232024 (10fnegri) [15:21:11] 10Toolforge, 06tools-platform-team, 07OKR-Work: Email tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865#12232171 (10taavi) My vote goes to option A. Option B to me seems to re-implement features that Alertmanager already does well with little benefits, and optio... [15:21:29] 10Toolforge, 06tools-platform-team, 07OKR-Work: Email tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865#12232186 (10dcaro) Another advantage of option B: * Easy decoupling from alertmanager, allowing us to easily use other systems instead or besides alertmanager... [15:26:18] 10Toolforge, 06tools-platform-team: [components-api] add the option to the tool configuration bulid section to allow passing built time envvars - https://phabricator.wikimedia.org/T434496#12232211 (10dcaro) 05Open→03In progress [15:26:20] 10Toolforge, 06tools-platform-team: [components-api] add the option to the tool configuration bulid section to allow passing built time envvars - https://phabricator.wikimedia.org/T434496#12232212 (10dcaro) a:03dcaro [15:35:37] (03open) 10dcaro: api: add envvars to source build [repos/cloud/toolforge/components-api] (update_toolforge_models) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/184 (https://phabricator.wikimedia.org/T434496) [15:37:10] (03update) 10dcaro: api: add envvars to source build [repos/cloud/toolforge/components-api] (update_toolforge_models) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/184 (https://phabricator.wikimedia.org/T434496) [15:44:58] (03open) 10mahveotm: cli: preserve option separators for subcommands [repos/cloud/toolforge/toolforge-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-cli/-/merge_requests/74 (https://phabricator.wikimedia.org/T370184) [15:45:15] (03update) 10dcaro: gen: update toolforge models [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/183 [15:53:36] 10Toolforge, 06tools-platform-team, 13Patch-For-Review: [cli] the generic cli swallows the `--` from other commands - https://phabricator.wikimedia.org/T370184#12232448 (10Mahveotm) 05Open→03In progress [15:53:47] 10Toolforge, 06tools-platform-team, 13Patch-For-Review: [cli] the generic cli swallows the `--` from other commands - https://phabricator.wikimedia.org/T370184#12232449 (10Mahveotm) a:03Mahveotm [16:18:45] 10Toolforge, 06tools-platform-team, 07OKR-Work: Email tool maintainers when an alert condition is detected - https://phabricator.wikimedia.org/T432865#12232604 (10fnegri) a:03fnegri [16:24:07] FIRING: ToolsDBHistoryLengthGrowing: ToolsDB History Length is above the desired threshold - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolsDBHistoryLengthGrowing - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolsDBHistoryLengthGrowing [16:29:44] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects paulina.toolforge.org performance - https://phabricator.wikimedia.org/T432878#12232660 (10bd808) 05Open→03In progress a:03Pepe_piton [16:36:29] 10Toolforge, 06tools-platform-team: [toolsdb] 2026-08-18 Transaction History Length growing too much - https://phabricator.wikimedia.org/T435220#12232739 (10fnegri) I think this is caused by `mix-n-match` like in {T428139}. In that task, I manually scaled the deployment of `rustbot` to zero, but I see that the... [16:45:39] 10Toolforge, 06tools-platform-team: [toolsdb] 2026-08-18 Transaction History Length growing too much - https://phabricator.wikimedia.org/T435220#12232852 (10fnegri) 05Open→03In progress [17:21:23] 10Cloud-VPS (Debian Bullseye Deprecation), 10Continuous-Integration-Infrastructure: Re-build integration-cumin.integration.eqiad1.wikimedia.cloud on something newer than bullseye - https://phabricator.wikimedia.org/T433592#12233165 (10Jdforrester-WMF) I created the instance, but I can't shell into the new box... [17:27:49] 10Cloud-VPS (Debian Bullseye Deprecation), 10Continuous-Integration-Infrastructure: Re-build integration-cumin.integration.eqiad1.wikimedia.cloud on something newer than bullseye - https://phabricator.wikimedia.org/T433592#12233216 (10taavi) The initial Puppet run on the host is failing to compile, which means... [17:29:39] 10Cloud-VPS (Debian Bullseye Deprecation), 10Continuous-Integration-Infrastructure: Re-build integration-cumin.integration.eqiad1.wikimedia.cloud on something newer than bullseye - https://phabricator.wikimedia.org/T433592#12233220 (10Jdforrester-WMF) I applied them but it didn't let me in, so I removed them a... [17:32:44] FIRING: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [17:36:36] 10Cloud-VPS (Debian Bullseye Deprecation), 10Continuous-Integration-Infrastructure: Re-build integration-cumin.integration.eqiad1.wikimedia.cloud on something newer than bullseye - https://phabricator.wikimedia.org/T433592#12233263 (10taavi) Yeah, if the initial Puppet run fails then the instance usually gets... [17:37:30] 10Cloud-VPS (Debian Bullseye Deprecation), 10Continuous-Integration-Infrastructure: Re-build integration-cumin.integration.eqiad1.wikimedia.cloud on something newer than bullseye - https://phabricator.wikimedia.org/T433592#12233266 (10Jdforrester-WMF) Aha! OK, that's fine. [17:46:50] !log taavi@cloudcumin1001 onfire START - Cookbook wmcs.vps.delete_project for project onfire in eqiad1 [17:47:20] (03open) 10group_199_bot_f98be072172e323ae6d1441939d3e461: projects: delete project onfire [repos/cloud/cloud-vps/tofu-infra] - 10https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/344 [17:49:00] (03merge) 10taavi: projects: delete project onfire [repos/cloud/cloud-vps/tofu-infra] - 10https://gitlab.wikimedia.org/repos/cloud/cloud-vps/tofu-infra/-/merge_requests/344 (owner: 10group_199_bot_f98be072172e323ae6d1441939d3e461) [17:49:29] !log taavi@cloudcumin1001 onfire END (PASS) - Cookbook wmcs.vps.delete_project (exit_code=0) for project onfire in eqiad1 [17:53:40] 10Cloud-VPS (Debian Bullseye Deprecation), 10Continuous-Integration-Infrastructure: Re-build integration-cumin.integration.eqiad1.wikimedia.cloud on something newer than bullseye - https://phabricator.wikimedia.org/T433592#12233336 (10Jdforrester-WMF) OK, newly-re-created integration-cumin-01 is now set up, wi... [17:55:47] 06cloud-services-team, 10Toolforge, 06tools-platform-team, 10Elasticsearch, 07Epic: Deploy multi-tenant OpenSearch cluster as replacement for Elasticsearch - https://phabricator.wikimedia.org/T348943#12233364 (10Izno) [18:01:23] (03update) 10sweil: Add new OCR parameter to normalize the result text [toolforge-repos/ocr] - 10https://gitlab.wikimedia.org/toolforge-repos/ocr/-/merge_requests/3 [18:08:14] 06cloud-services-team, 10Toolforge, 06tools-platform-team, 10Elasticsearch, 07Epic: Deploy multi-tenant OpenSearch cluster as replacement for Elasticsearch - https://phabricator.wikimedia.org/T348943#12233405 (10BLiviero-WMF) A new OpenSearch cluster was made available (and announced to cloud-announce) o... [18:22:44] RESOLVED: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [18:24:36] (03update) 10lucaswerkmeister: toolforge-cd: generate deployment description from Git [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/97 (https://phabricator.wikimedia.org/T393169 https://phabricator.wikimedia.org/T401993) [18:24:49] (03update) 10lucaswerkmeister: toolforge-cd: generate deployment description from Git [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/97 (https://phabricator.wikimedia.org/T393169 https://phabricator.wikimedia.org/T401993) [18:30:04] 06tools-infrastructure-team: Decommission "onfire" Cloud VPS project - https://phabricator.wikimedia.org/T435387 (10BLiviero-WMF) 03NEW [18:42:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [18:52:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [18:53:32] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12233630 (10VRiley-WMF) [19:02:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [19:07:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [19:08:56] FIRING: CloudVPSDesignateLeaks: Detected 8 stray dns records - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/Designate_record_leaks - https://grafana.wikimedia.org/d/ebJoA6VWz/wmcs-openstack-eqiad-nova-fullstack - https://alerts.wikimedia.org/?q=alertname%3DCloudVPSDesignateLeaks [19:12:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [19:22:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [19:27:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [19:30:21] 10Cloud-VPS (Debian Bullseye Deprecation), 10Beta-Cluster-Infrastructure, 06MediaWiki-Engineering: Beta Cluster: Switch to PHP 8.5, includes writing Puppet changes to support PHP 8.5 on baremetal - https://phabricator.wikimedia.org/T435393 (10bd808) 03NEW [19:30:56] 10Cloud-VPS (Debian Bullseye Deprecation), 10Beta-Cluster-Infrastructure, 06MediaWiki-Engineering: Beta Cluster: Switch to PHP 8.5, includes writing Puppet changes to support PHP 8.5 on baremetal - https://phabricator.wikimedia.org/T435393#12233843 (10bd808) [19:30:59] 10Cloud-VPS (Debian Bullseye Deprecation), 10Beta-Cluster-Infrastructure, 07Epic, 06Release-Engineering-Team (Priority Backlog 📥): Migrate deployment-prep away from Debian Bullseye to Bookworm/Trixie - https://phabricator.wikimedia.org/T401839#12233844 (10bd808) [19:37:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [20:40:37] RESOLVED: ToolsDBHistoryLengthGrowing: ToolsDB History Length is above the desired threshold - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolsDBHistoryLengthGrowing - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolsDBHistoryLengthGrowing [20:50:03] FIRING: PuppetAgentStaleLastRun: Last Puppet run was over 24 hours ago on instance proxy-5 in project project-proxy - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetAgentStaleLastRun [21:49:23] 10Cloud-VPS, 06tools-infrastructure-team: Request for object_storage role - https://phabricator.wikimedia.org/T435414 (10Gabinaluz) 03NEW [21:52:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [21:57:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [22:01:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [22:06:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [22:42:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [22:52:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [22:54:44] FIRING: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [22:57:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [23:02:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [23:07:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [23:09:01] FIRING: CloudVPSDesignateLeaks: Detected 8 stray dns records - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/Designate_record_leaks - https://grafana.wikimedia.org/d/ebJoA6VWz/wmcs-openstack-eqiad-nova-fullstack - https://alerts.wikimedia.org/?q=alertname%3DCloudVPSDesignateLeaks [23:17:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [23:19:44] RESOLVED: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [23:48:26] 10Cloud-Services: Unable to subscribe to mailing lists anonymously - https://phabricator.wikimedia.org/T435424 (10PacmanD) 03NEW The #Cloud-Services project tag is not intended to have any tasks. Please check the list on https://phabricator.wikimedia.org/project/profile/832/ and replace it with a more specific...