[00:14:28] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/namehistory] - 10https://gitlab.wikimedia.org/toolforge-repos/namehistory/-/merge_requests/12 [00:14:31] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/namehistory] - 10https://gitlab.wikimedia.org/toolforge-repos/namehistory/-/merge_requests/12 [00:14:42] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/allblocksever] - 10https://gitlab.wikimedia.org/toolforge-repos/allblocksever/-/merge_requests/24 [00:14:46] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/allblocksever] - 10https://gitlab.wikimedia.org/toolforge-repos/allblocksever/-/merge_requests/24 [00:14:57] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/checkusertools] - 10https://gitlab.wikimedia.org/toolforge-repos/checkusertools/-/merge_requests/19 [00:15:00] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/checkusertools] - 10https://gitlab.wikimedia.org/toolforge-repos/checkusertools/-/merge_requests/19 [00:17:09] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/rangetree] - 10https://gitlab.wikimedia.org/toolforge-repos/rangetree/-/merge_requests/51 [00:17:13] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/rangetree] - 10https://gitlab.wikimedia.org/toolforge-repos/rangetree/-/merge_requests/51 [00:17:33] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/multiuserinfo] - 10https://gitlab.wikimedia.org/toolforge-repos/multiuserinfo/-/merge_requests/77 [00:17:34] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/multiuserinfo] - 10https://gitlab.wikimedia.org/toolforge-repos/multiuserinfo/-/merge_requests/77 [00:17:39] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/multiuserinfo] - 10https://gitlab.wikimedia.org/toolforge-repos/multiuserinfo/-/merge_requests/77 [00:18:24] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/wiki-mail-verify] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-mail-verify/-/merge_requests/35 [00:18:27] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/wiki-mail-verify] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-mail-verify/-/merge_requests/35 [00:19:11] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/dewiki-voterightsnotifier] - 10https://gitlab.wikimedia.org/toolforge-repos/dewiki-voterightsnotifier/-/merge_requests/40 [00:19:11] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/dewiki-voterightsnotifier] - 10https://gitlab.wikimedia.org/toolforge-repos/dewiki-voterightsnotifier/-/merge_requests/40 [00:19:20] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/dewiki-voterightsnotifier] - 10https://gitlab.wikimedia.org/toolforge-repos/dewiki-voterightsnotifier/-/merge_requests/40 [00:19:41] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/dewiki-rangeblock] - 10https://gitlab.wikimedia.org/toolforge-repos/dewiki-rangeblock/-/merge_requests/25 [00:19:46] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/dewiki-rangeblock] - 10https://gitlab.wikimedia.org/toolforge-repos/dewiki-rangeblock/-/merge_requests/25 [00:19:53] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/flaggedrevspromotioncheck] - 10https://gitlab.wikimedia.org/toolforge-repos/flaggedrevspromotioncheck/-/merge_requests/28 [00:20:02] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/flaggedrevspromotioncheck] - 10https://gitlab.wikimedia.org/toolforge-repos/flaggedrevspromotioncheck/-/merge_requests/28 [00:20:04] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/dewikisignbot] - 10https://gitlab.wikimedia.org/toolforge-repos/dewikisignbot/-/merge_requests/31 [00:20:05] (03update) 10renovatebot: Lock file maintenance [toolforge-repos/dewikisignbot] - 10https://gitlab.wikimedia.org/toolforge-repos/dewikisignbot/-/merge_requests/31 [00:20:16] (03open) 10renovatebot: Lock file maintenance [toolforge-repos/dewikisignbot] - 10https://gitlab.wikimedia.org/toolforge-repos/dewikisignbot/-/merge_requests/31 [00:38:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [00:43:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [01:17:02] (03PS5) 10Krinkle: [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) [01:17:35] (03CR) 10CI reject: [V:04-1] [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) (owner: 10Krinkle) [02:04:28] 10Cloud-VPS, 06tools-infrastructure-team: Advice for Hybrid Architecture for AI Tool - https://phabricator.wikimedia.org/T416590#12219567 (10Andrew) That sounds right, and very doable. One thing to keep in mind is that toolforge won't run an externally-built container, you'll need to use the toolforge build se... [02:14:32] (03PS6) 10Krinkle: [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) [02:14:46] (03CR) 10CI reject: [V:04-1] [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) (owner: 10Krinkle) [02:20:04] 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: Network unavailable on a few VMs - https://phabricator.wikimedia.org/T432426#12219581 (10Andrew) today a repeat for demo-patrollers.wmf-research-tools.eqiad1.wikimedia.cloud [02:28:17] 10Toolforge, 06tools-infrastructure-team: [disable-tools] Disabled Toolforge tool deletions not occurring - https://phabricator.wikimedia.org/T431241#12219594 (10Andrew) There's at least one example of this working again: https://toolsadmin.wikimedia.org/tools/id/wikiblame [02:37:16] (03PS7) 10Krinkle: [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) [02:37:49] (03CR) 10CI reject: [V:04-1] [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) (owner: 10Krinkle) [03:29:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [03:36:44] FIRING: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [03:39:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [03:41:44] RESOLVED: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [03:59:04] FIRING: PuppetAgentStaleLastRun: Last Puppet run was over 24 hours ago on instance tools-k8s-haproxy-8 in project tools - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetAgentStaleLastRun [04:09:03] 10Tool-mwcli: mwcli containers fail on arm systems - https://phabricator.wikimedia.org/T435037 (10PixDeVl) 03NEW [04:10:39] 10Tool-mwcli: mwcli containers fail on arm systems - https://phabricator.wikimedia.org/T435037#12219633 (10PixDeVl) [04:12:48] 10Tool-mwcli, 10a Wikimedia CLI: mwcli containers fail on arm systems - https://phabricator.wikimedia.org/T435037#12219634 (10PixDeVl) [04:54:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [05:04:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:09:29] (03PS8) 10Krinkle: [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) [05:09:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:10:01] (03CR) 10CI reject: [V:04-1] [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) (owner: 10Krinkle) [05:49:29] FIRING: ToolforgeToolviewsStale: Toolviews data is stale - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolforgeToolviewsStale - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolforgeToolviewsStale [06:40:39] (03PS9) 10Krinkle: [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) [06:41:13] (03CR) 10CI reject: [V:04-1] [WIP] Rewrite CVNBot in python [labs/countervandalism/CVNBot] - 10https://gerrit.wikimedia.org/r/1324908 (https://phabricator.wikimedia.org/T327136) (owner: 10Krinkle) [06:42:48] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12219727 (10fgiunchedi) [06:44:47] 10Tool-wiki-gender-stats: Add alternative texts to images on the website - https://phabricator.wikimedia.org/T434311#12219730 (10ShivangiU2) Hi @Danya, yes, I’ll go through the contribution guide before proceeding. Thanks for the heads-up about the 0.4.2 release and the updated requirement for localized alt texts! [06:55:07] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12219735 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by filippo@cumin1003 for hosts: `cloudvirt1049... [07:02:28] FIRING: [2x] KeyholderUnarmed: 2 unarmed Keyholder key(s) on cloudcumin1001:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [07:05:20] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12219761 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by filippo@cumin1003 for hosts: `cloudvirt1051... [07:07:28] RESOLVED: [2x] KeyholderUnarmed: 2 unarmed Keyholder key(s) on cloudcumin1001:9100 - https://wikitech.wikimedia.org/wiki/Keyholder - TODO - https://alerts.wikimedia.org/?q=alertname%3DKeyholderUnarmed [07:07:50] FIRING: PawsJupyterHubDown: PAWS JupyterHub is down https://wikitech.wikimedia.org/wiki/PAWS/Admin - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPawsJupyterHubDown [07:08:03] FIRING: TargetDown: Job jupyterhub is unreachable in project paws instance hub-paws.wmcloud.org:443 - https://prometheus-alerts.wmcloud.org/?q=alertname%3DTargetDown [07:09:22] 10Data-Services, 06tools-infrastructure-team, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Remove hdfs fuse mounts from clouddumps - https://phabricator.wikimedia.org/T434794#12219764 (10Gehel) [07:15:58] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12219772 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by filippo@cumin1003 for hosts: `cloudvirt1054... [07:24:21] !log dcaro@acme paws START - Cookbook wmcs.openstack.cloudvirt.vm_console [07:24:24] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Paws/SAL [07:25:24] !log dcaro@acme paws END (ERROR) - Cookbook wmcs.openstack.cloudvirt.vm_console (exit_code=255) [07:25:25] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Paws/SAL [07:28:36] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12219799 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by filippo@cumin1003 for hosts: `cloudvirt1055... [07:39:19] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12219841 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by filippo@cumin1003 for hosts: `cloudvirt1056... [07:47:50] RESOLVED: PawsJupyterHubDown: PAWS JupyterHub is down https://wikitech.wikimedia.org/wiki/PAWS/Admin - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPawsJupyterHubDown [07:48:03] RESOLVED: TargetDown: Job jupyterhub is unreachable in project paws instance hub-paws.wmcloud.org:443 - https://prometheus-alerts.wmcloud.org/?q=alertname%3DTargetDown [07:54:34] 06cloud-services-team, 10Tools, 07Performance Issue: Stalktoy is often slow or not responding - https://phabricator.wikimedia.org/T427941#12219898 (10Framawiki) An alternative to Oauth login of the users using their wiki account, could be to block crawlers could be a simple html page sent when a cookie is mi... [08:00:44] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12219926 (10ops-monitoring-bot) cookbooks.sre.hosts.decommission executed by filippo@cumin1003 for hosts: `cloudvirt1057... [08:09:03] RESOLVED: PuppetAgentStaleLastRun: Last Puppet run was over 24 hours ago on instance tools-k8s-haproxy-8 in project tools - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetAgentStaleLastRun [08:19:00] 06tools-infrastructure-team, 06Wikidata Platform Team, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Migrate to dumps-nfs.w.o in production - https://phabricator.wikimedia.org/T432212#12219970 (10fgiunchedi) >>! In T432212#12210764, @fgiunchedi wrote: >>>! In T432212#12210704, @Gehel w... [08:21:23] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12219988 (10fgiunchedi) >>! In T431682#12210883, @fgiunchedi wrote: >>>! In T431682#12209331, @VRiley-WMF wrote: >> Hey... [08:24:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:29:15] 10Tool-lexeme-forms, 06translatewiki.net, 10LPL Projects (Ongoing maintenance), 10LPL sprints (S2 2026 August): l10n-bot cannot create GitLab merge requests due to gitlab-ssh migration - https://phabricator.wikimedia.org/T432838#12220058 (10Nikerabbit) I've renamed the variable back. [08:29:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:31:22] (03open) 10dcaro: start-devenv: Show that yes is the default [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/339 [08:31:44] FIRING: NovaComputeUnavailable: #page nova-compute agent is unavailable - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/NovaComputeUnavailable - https://grafana.wikimedia.org/d/000000579/wmcs-openstack-eqiad-summary - https://alerts.wikimedia.org/?q=alertname%3DNovaComputeUnavailable [08:32:33] 10Tool-wiki-gender-stats: Add alternative texts to images on the website - https://phabricator.wikimedia.org/T434311#12220073 (10Danya) If you struggle with French and/or Esperanto translations, don’t hesitate to ask me for help. Please do not use automated translations tools, especially for Esperanto, they are... [08:32:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:33:09] (03update) 10dcaro: start-devenv: Show that yes is the default [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/339 [08:36:46] 10Tool-wiki-gender-stats: [TF] Integrate May 2026 data from Danÿa’s local copy - https://phabricator.wikimedia.org/T431933#12220090 (10Danya) 05Open→03Stalled Blocked by bug T434400. [08:37:17] 10Tool-wiki-gender-stats: [TF] Ingest data for August 2026 - https://phabricator.wikimedia.org/T434328#12220095 (10Danya) 05In progress→03Stalled [08:38:30] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [08:39:28] 10Tool-wiki-gender-stats, 13Patch-For-Review: Dumps ingester crashes because of too many prepared statements - https://phabricator.wikimedia.org/T434400#12220102 (10Danya) p:05High→03Unbreak! [08:39:29] RESOLVED: ToolforgeToolviewsStale: Toolviews data is stale - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolforgeToolviewsStale - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolforgeToolviewsStale [08:39:57] 10Tool-wiki-gender-stats, 13Patch-For-Review: Dumps ingester crashes because of too many prepared statements - https://phabricator.wikimedia.org/T434400#12220117 (10Danya) p:05Unbreak!→03High [08:40:38] 10Tool-wiki-gender-stats, 13Patch-For-Review: Dumps ingester crashes because of too many prepared statements - https://phabricator.wikimedia.org/T434400#12220123 (10Danya) a:03Danya [08:42:56] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [08:47:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:52:30] (03open) 10dcaro: global: move all stages to shared yaml [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/94 [08:52:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:54:24] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [08:54:45] (03update) 10dcaro: global: move all stages to shared yaml [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/94 [08:59:03] FIRING: TargetDown: Job main-haproxy is unreachable in project project-proxy instance proxy-5 - https://prometheus-alerts.wmcloud.org/?q=alertname%3DTargetDown [09:04:03] RESOLVED: TargetDown: Job main-haproxy is unreachable in project project-proxy instance proxy-5 - https://prometheus-alerts.wmcloud.org/?q=alertname%3DTargetDown [09:04:38] (03open) 10dcaro: ci: move to trixie images for tests [toolforge-repos/toolstats] - 10https://gitlab.wikimedia.org/toolforge-repos/toolstats/-/merge_requests/5 [09:04:45] (03close) 10dcaro: gitlab: move to bookworm images [toolforge-repos/toolstats] - 10https://gitlab.wikimedia.org/toolforge-repos/toolstats/-/merge_requests/4 [09:05:21] (03approved) 10dcaro: ci: move to trixie images for tests [toolforge-repos/toolstats] - 10https://gitlab.wikimedia.org/toolforge-repos/toolstats/-/merge_requests/5 [09:05:23] (03merge) 10dcaro: ci: move to trixie images for tests [toolforge-repos/toolstats] - 10https://gitlab.wikimedia.org/toolforge-repos/toolstats/-/merge_requests/5 [09:10:38] 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team, 07OKR-Work: [metricsinfra] allow custom routing for Toolforge alerts - https://phabricator.wikimedia.org/T434801#12220354 (10fnegri) Thanks @taavi, I will attempt creating a patch for option 2, and send it to you and @fgiunchedi for review. [09:10:45] (03PS1) 10Tiziano Fogli: grafana/image-renderer: add renderer token [labs/private] - 10https://gerrit.wikimedia.org/r/1326222 (https://phabricator.wikimedia.org/T432970) [09:11:01] 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team, 07OKR-Work: [metricsinfra] allow custom routing for Toolforge alerts - https://phabricator.wikimedia.org/T434801#12220355 (10fnegri) 05Open→03In progress a:03fnegri [09:11:11] (03CR) 10Tiziano Fogli: [C:03+2] grafana/image-renderer: add renderer token [labs/private] - 10https://gerrit.wikimedia.org/r/1326222 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [09:11:14] (03CR) 10Tiziano Fogli: [V:03+2 C:03+2] grafana/image-renderer: add renderer token [labs/private] - 10https://gerrit.wikimedia.org/r/1326222 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [09:18:57] (03open) 10dcaro: add logging timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1373 [09:19:44] (03update) 10dcaro: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] (helmfileorder) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1373 [09:20:43] 10Tool-lexeme-forms, 06translatewiki.net, 10LPL Projects (Ongoing maintenance), 10LPL sprints (S2 2026 August): l10n-bot cannot create GitLab merge requests due to gitlab-ssh migration - https://phabricator.wikimedia.org/T432838#12220392 (10abi_) Lets see if the MR gets created today. [09:21:07] (03open) 10dcaro: ci: use trixie image [repos/cloud/toolforge/toolforge-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-cli/-/merge_requests/73 [09:21:52] (03update) 10dcaro: ci: use trixie image [repos/cloud/toolforge/toolforge-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-cli/-/merge_requests/73 [09:23:34] (03merge) 10dcaro: ci: use trixie image [repos/cloud/toolforge/toolforge-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-cli/-/merge_requests/73 [09:25:12] (03open) 10dcaro: ci: use trixie image [repos/cloud/toolforge/webservice-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/webservice-cli/-/merge_requests/120 [09:28:07] (03approved) 10dcaro: ci: use trixie image [repos/cloud/toolforge/webservice-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/webservice-cli/-/merge_requests/120 [09:28:12] (03merge) 10dcaro: ci: use trixie image [repos/cloud/toolforge/webservice-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/webservice-cli/-/merge_requests/120 [09:29:30] (03open) 10dcaro: ci: use trixie images [toolforge-repos/fourohfour] - 10https://gitlab.wikimedia.org/toolforge-repos/fourohfour/-/merge_requests/19 [09:30:14] (03merge) 10dcaro: ci: use trixie images [toolforge-repos/fourohfour] - 10https://gitlab.wikimedia.org/toolforge-repos/fourohfour/-/merge_requests/19 [09:32:14] (03open) 10dcaro: ci: use trixie [toolforge-repos/k8s-status] - 10https://gitlab.wikimedia.org/toolforge-repos/k8s-status/-/merge_requests/8 [09:34:19] (03update) 10dcaro: ci: use trixie [toolforge-repos/k8s-status] - 10https://gitlab.wikimedia.org/toolforge-repos/k8s-status/-/merge_requests/8 [09:35:16] (03update) 10dcaro: ci: use trixie [toolforge-repos/k8s-status] - 10https://gitlab.wikimedia.org/toolforge-repos/k8s-status/-/merge_requests/8 [09:44:14] 10Tool-itwiki: BotCancellazioni: rollback "cllimit=500" workaround after T433922 has been resolved - https://phabricator.wikimedia.org/T434598#12220501 (10valerio.bozzolan) Do you propose a wontfix? That's fine for me [09:47:44] (03PS1) 10Tiziano Fogli: grafana/image-renderer: adjust dummy secret [labs/private] - 10https://gerrit.wikimedia.org/r/1326229 (https://phabricator.wikimedia.org/T432970) [09:47:46] (03CR) 10Tiziano Fogli: [V:03+2 C:03+2] grafana/image-renderer: adjust dummy secret [labs/private] - 10https://gerrit.wikimedia.org/r/1326229 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [09:54:05] (03PS1) 10Tiziano Fogli: grafana/image-renderer: adjust dummy secret [labs/private] - 10https://gerrit.wikimedia.org/r/1326233 (https://phabricator.wikimedia.org/T432970) [09:54:08] (03CR) 10Tiziano Fogli: [V:03+2 C:03+2] grafana/image-renderer: adjust dummy secret [labs/private] - 10https://gerrit.wikimedia.org/r/1326233 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [09:54:38] 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: openstack: alert for cloudvirts without aggregate or with unexpected set of them - https://phabricator.wikimedia.org/T284747#12220520 (10fgiunchedi) Since we ditched the `maintenance` aggregate in {T424802} we can use the openstack-exporter metric... [09:59:50] (03PS1) 10Clément Goubert: Dummy pass for redis::lock::instance [labs/private] - 10https://gerrit.wikimedia.org/r/1326236 (https://phabricator.wikimedia.org/T427999) [10:01:27] (03PS1) 10Tiziano Fogli: grafana/image-renderer: use only hex chars for renderer token [labs/private] - 10https://gerrit.wikimedia.org/r/1326237 (https://phabricator.wikimedia.org/T432970) [10:01:32] (03CR) 10Tiziano Fogli: [C:03+2] grafana/image-renderer: use only hex chars for renderer token [labs/private] - 10https://gerrit.wikimedia.org/r/1326237 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [10:01:34] (03CR) 10Tiziano Fogli: [V:03+2 C:03+2] grafana/image-renderer: use only hex chars for renderer token [labs/private] - 10https://gerrit.wikimedia.org/r/1326237 (https://phabricator.wikimedia.org/T432970) (owner: 10Tiziano Fogli) [10:01:48] (03CR) 10Majavah: [C:03+2] vps: create_instance: Add codfw1dev support [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1320917 (owner: 10Majavah) [10:04:54] (03Merged) 10jenkins-bot: vps: create_instance: Add codfw1dev support [cloud/wmcs-cookbooks] - 10https://gerrit.wikimedia.org/r/1320917 (owner: 10Majavah) [10:06:00] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [10:08:34] (03CR) 10Clément Goubert: [C:03+2] Dummy pass for redis::lock::instance [labs/private] - 10https://gerrit.wikimedia.org/r/1326236 (https://phabricator.wikimedia.org/T427999) (owner: 10Clément Goubert) [10:08:38] (03CR) 10Clément Goubert: [V:03+2 C:03+2] Dummy pass for redis::lock::instance [labs/private] - 10https://gerrit.wikimedia.org/r/1326236 (https://phabricator.wikimedia.org/T427999) (owner: 10Clément Goubert) [10:08:45] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [10:11:05] 10PAWS, 06tools-platform-team: Notebook error - https://phabricator.wikimedia.org/T435068 (10Gtda132) 03NEW [10:25:32] 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: Convert dynamicproxy from nginx to haproxy - https://phabricator.wikimedia.org/T429930#12220692 (10taavi) [10:27:38] (03update) 10dcaro: Alloy relabeling [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1371 (https://phabricator.wikimedia.org/T432969) (owner: 10tlepage) [10:31:22] 10PAWS, 06tools-platform-team: Notebook error - https://phabricator.wikimedia.org/T435068#12220719 (10dcaro) Hi @Gtda132, your account was automatically locked due to misuse of resources. Can you elaborate on what are you using it for, why and how does it fit with paws terms of usage? (see https://wikitech.wi... [10:31:37] 10PAWS, 06tools-platform-team: Notebook error - https://phabricator.wikimedia.org/T435068#12220722 (10dcaro) a:03Gtda132 [10:43:49] FIRING: NeutronAgentDownForLong: Neutron neutron-openvswitch-agent on cloudvirt1049 has been down for more than 2h - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDownForLong [10:44:49] FIRING: NeutronAgentDown: Neutron neutron-openvswitch-agent on cloudvirt1049 is down - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDown [10:58:49] FIRING: [2x] NeutronAgentDownForLong: Neutron neutron-openvswitch-agent on cloudvirt1049 has been down for more than 2h - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDownForLong [10:59:49] FIRING: [2x] NeutronAgentDown: Neutron neutron-openvswitch-agent on cloudvirt1049 is down - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDown [11:02:21] 10Cloud-VPS, 06tools-infrastructure-team: Migrate cloudidp-dev to Redis ticket storage - https://phabricator.wikimedia.org/T433963#12220810 (10taavi) 05Open→03Resolved [11:08:49] FIRING: [3x] NeutronAgentDownForLong: Neutron neutron-openvswitch-agent on cloudvirt1049 has been down for more than 2h - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDownForLong [11:09:50] FIRING: [3x] NeutronAgentDown: Neutron neutron-openvswitch-agent on cloudvirt1049 is down - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDown [11:09:56] 10Cloud-VPS (Debian Bullseye Deprecation), 06tools-infrastructure-team, 13Patch-For-Review: Upgrade IDP server in cloudinfra project off of Bullseye - https://phabricator.wikimedia.org/T402006#12220821 (10taavi) 05Open→03Resolved [11:10:57] 06cloud-services-team, 10Cloud-VPS (Debian Bullseye Deprecation): Migrate cloudinfra project off of Debian Bullseye - https://phabricator.wikimedia.org/T401811#12220831 (10taavi) 05Open→03Resolved a:03taavi [11:11:19] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12220833 (10VRiley-WMF) Perfect. I will be starting to take these down and moving them now. [11:11:20] 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 06SRE, 13Patch-For-Review: Modernise memcached systemd unit / sync, and make it presentable - https://phabricator.wikimedia.org/T273950#12220834 (10taavi) [11:15:03] FIRING: PuppetAgentNoResources: No Puppet resources found on instance metricsinfra-grafana-2 on project metricsinfra - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetAgentNoResources [11:18:50] FIRING: [4x] NeutronAgentDownForLong: Neutron neutron-openvswitch-agent on cloudvirt1049 has been down for more than 2h - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDownForLong [11:19:49] FIRING: [4x] NeutronAgentDown: Neutron neutron-openvswitch-agent on cloudvirt1049 is down - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDown [11:29:50] FIRING: [5x] NeutronAgentDown: Neutron neutron-openvswitch-agent on cloudvirt1049 is down - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDown [11:33:43] (03approved) 10dcaro: k8s v1.33.13, helm v4.1.4, helmfile v1.7.3, kindest/node v1.32.11 [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/338 (https://phabricator.wikimedia.org/T433132) (owner: 10andrew) [11:33:50] FIRING: [5x] NeutronAgentDownForLong: Neutron neutron-openvswitch-agent on cloudvirt1049 has been down for more than 2h - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDownForLong [11:38:10] !log tlepage@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component logging [11:38:15] !log tlepage@cloudcumin1001 toolsbeta END (FAIL) - Cookbook wmcs.toolforge.component.deploy (exit_code=99) for component logging [11:39:47] !log tlepage@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component logging [11:42:09] (03update) 10dcaro: auth: Allow specifying allowed urls for superusers [repos/cloud/toolforge/api-gateway] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/api-gateway/-/merge_requests/103 [11:52:12] !log tlepage@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component logging [11:53:28] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12221186 (10VRiley-WMF) Servers have been physically relocated. Will start powering on the units and updating netbox sho... [11:53:49] FIRING: [6x] NeutronAgentDownForLong: Neutron neutron-openvswitch-agent on cloudvirt1049 has been down for more than 2h - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDownForLong [11:54:50] FIRING: [6x] NeutronAgentDown: Neutron neutron-openvswitch-agent on cloudvirt1049 is down - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Troubleshooting#Networking_failures - https://grafana.wikimedia.org/d/wKnDJf97z/wmcs-neutron-eqiad1 - https://alerts.wikimedia.org/?q=alertname%3DNeutronAgentDown [12:00:47] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [12:01:52] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [12:03:37] (03approved) 10dcaro: [T401993] Add deployment descriptions to the CLI [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/90 (https://phabricator.wikimedia.org/T401993) (owner: 10mahveotm) [12:06:32] (03approved) 10dcaro: [T401993] Add a "description" field to the deployment [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/181 (https://phabricator.wikimedia.org/T401993) (owner: 10mahveotm) [12:06:37] (03update) 10raymond-ndibe: models: support publish option [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/180 (https://phabricator.wikimedia.org/T433355) [12:06:41] (03approved) 10raymond-ndibe: models: support publish option [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/180 (https://phabricator.wikimedia.org/T433355) [12:07:00] (03merge) 10raymond-ndibe: models: support publish option [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/180 (https://phabricator.wikimedia.org/T433355) [12:07:10] (03update) 10raymond-ndibe: components-api: test continuous job publish [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1367 (https://phabricator.wikimedia.org/T433355) [12:09:54] (03open) 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce: components-api: bump to 0.0.206-20260817120711-a9c79b26 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1374 (https://phabricator.wikimedia.org/T433355) [12:12:19] (03update) 10dcaro: auth: Allow specifying allowed urls for superusers [repos/cloud/toolforge/api-gateway] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/api-gateway/-/merge_requests/103 [12:16:52] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [12:17:36] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [12:19:43] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [12:21:45] !log raymond-ndibe@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component components-api [12:25:11] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [12:26:05] !log raymond-ndibe@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-api [12:26:35] !log raymond-ndibe@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component components-api [12:29:52] (03open) 10l10n-bot: Localisation updates from https://translatewiki.net. [toolforge-repos/wd-image-positions] - 10https://gitlab.wikimedia.org/toolforge-repos/wd-image-positions/-/merge_requests/75 [12:29:53] (03open) 10l10n-bot: Localisation updates from https://translatewiki.net. [toolforge-repos/lexeme-forms] - 10https://gitlab.wikimedia.org/toolforge-repos/lexeme-forms/-/merge_requests/45 [12:31:47] !log raymond-ndibe@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-api [12:33:00] (03update) 10raymond-ndibe: components-api: bump to 0.0.206-20260817120711-a9c79b26 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1374 (https://phabricator.wikimedia.org/T433355) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [12:33:01] (03approved) 10raymond-ndibe: components-api: bump to 0.0.206-20260817120711-a9c79b26 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1374 (https://phabricator.wikimedia.org/T433355) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [12:33:08] (03merge) 10raymond-ndibe: components-api: bump to 0.0.206-20260817120711-a9c79b26 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1374 (https://phabricator.wikimedia.org/T433355) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [12:33:34] (03update) 10raymond-ndibe: components-api: test continuous job publish [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1367 (https://phabricator.wikimedia.org/T433355) [12:34:13] !log raymond-ndibe@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component components-api [12:37:18] !log raymond-ndibe@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-api [12:38:39] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12221367 (10VRiley-WMF) 1049 CableID 5295 Row 14 Port 9 1051 CableID 5297 Row 15 Port 7 1054 CableID 5352 Row 16 Port... [12:39:35] (03update) 10mahveotm: [T401993] Add a "description" field to the deployment [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/181 (https://phabricator.wikimedia.org/T401993) [12:53:05] 10Tool-lexeme-forms, 06translatewiki.net, 07Essential-Work, 10LPL Projects (Ongoing maintenance), 10LPL sprints (S2 2026 August): l10n-bot cannot create GitLab merge requests due to gitlab-ssh migration - https://phabricator.wikimedia.org/T432838#12221435 (10Nikerabbit) 05Open→03Resolved This is... [12:54:24] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [12:57:44] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12221469 (10VRiley-WMF) Updated Netbox device page with new rack location Deleted the cables connected to the device's i... [12:59:29] !log raymond-ndibe@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component components-api [13:03:18] !log raymond-ndibe@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-api [13:06:21] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [13:07:17] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [13:08:08] 10Toolforge, 06tools-platform-team: Deploying static apps to toolforge - https://phabricator.wikimedia.org/T435088 (10Gouvernathor) 03NEW [13:09:14] 10Toolforge, 06tools-platform-team: Deploying static apps to toolforge - https://phabricator.wikimedia.org/T435088#12221576 (10Gouvernathor) [13:10:27] (03merge) 10raymond-ndibe: components-api: test continuous job publish [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1367 (https://phabricator.wikimedia.org/T433355) [13:13:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [13:18:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [13:20:24] 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: "osd.224 observed stalled read indications in DB device" - https://phabricator.wikimedia.org/T434402#12221626 (10Andrew) Today, osd.353 observed stalled read indications in DB device [13:23:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [13:26:28] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12221666 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1... [13:26:54] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 3 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12221673 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1049.... [13:28:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [13:35:00] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12221763 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1... [13:39:22] (03update) 10dcaro: auth: Allow specifying allowed urls for superusers [repos/cloud/toolforge/api-gateway] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/api-gateway/-/merge_requests/103 [13:40:10] (03update) 10dcaro: auth: Allow specifying allowed urls for superusers [repos/cloud/toolforge/api-gateway] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/api-gateway/-/merge_requests/103 [13:42:51] !log tlepage@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component logging [13:48:05] !log tlepage@cloudcumin1001 tools END (FAIL) - Cookbook wmcs.toolforge.component.deploy (exit_code=99) for component logging [13:49:30] (03update) 10dcaro: auth: Allow specifying allowed urls for superusers [repos/cloud/toolforge/api-gateway] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/api-gateway/-/merge_requests/103 [13:56:39] 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: "osd.224 observed stalled read indications in DB device" - https://phabricator.wikimedia.org/T434402#12221872 (10Andrew) There are some decent docs about this alert here: https://oneuptime.com/blog/post/2026-03-31-rook-fix-db-device-stalled-read-alert-ceph/view [14:08:26] (03update) 10dcaro: global: add log write endpoint [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/32 (https://phabricator.wikimedia.org/T432565) [14:10:11] (03update) 10dcaro: global: add log write endpoint [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/32 [14:12:22] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12221981 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1049.... [14:15:30] (03update) 10raymond-ndibe: tests: add behavior tests for previously untested code branches [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/362 (https://phabricator.wikimedia.org/T434166) [14:24:38] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12222083 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1... [14:29:33] 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: "osd.224 observed stalled read indications in DB device" - https://phabricator.wikimedia.org/T434402#12222104 (10Andrew) There's a reassuring screed about other admins getting tangled up by these new alerts, here: https://www.spinics.net/lists/ceph-users/msg... [14:29:47] !log tlepage@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component logging [14:32:23] (03merge) 10countcount: Lock file maintenance [toolforge-repos/dewikisignbot] - 10https://gitlab.wikimedia.org/toolforge-repos/dewikisignbot/-/merge_requests/31 (owner: 10renovatebot) [14:32:33] (03merge) 10countcount: Lock file maintenance [toolforge-repos/flaggedrevspromotioncheck] - 10https://gitlab.wikimedia.org/toolforge-repos/flaggedrevspromotioncheck/-/merge_requests/28 (owner: 10renovatebot) [14:32:36] (03merge) 10countcount: Lock file maintenance [toolforge-repos/dewiki-rangeblock] - 10https://gitlab.wikimedia.org/toolforge-repos/dewiki-rangeblock/-/merge_requests/25 (owner: 10renovatebot) [14:32:50] (03merge) 10countcount: Lock file maintenance [toolforge-repos/dewiki-voterightsnotifier] - 10https://gitlab.wikimedia.org/toolforge-repos/dewiki-voterightsnotifier/-/merge_requests/40 (owner: 10renovatebot) [14:32:52] (03merge) 10countcount: Lock file maintenance [toolforge-repos/wiki-mail-verify] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-mail-verify/-/merge_requests/35 (owner: 10renovatebot) [14:33:01] (03merge) 10countcount: Lock file maintenance [toolforge-repos/multiuserinfo] - 10https://gitlab.wikimedia.org/toolforge-repos/multiuserinfo/-/merge_requests/77 (owner: 10renovatebot) [14:33:09] (03merge) 10countcount: Lock file maintenance [toolforge-repos/rangetree] - 10https://gitlab.wikimedia.org/toolforge-repos/rangetree/-/merge_requests/51 (owner: 10renovatebot) [14:33:14] (03merge) 10countcount: Lock file maintenance [toolforge-repos/checkusertools] - 10https://gitlab.wikimedia.org/toolforge-repos/checkusertools/-/merge_requests/19 (owner: 10renovatebot) [14:33:17] (03merge) 10countcount: Lock file maintenance [toolforge-repos/allblocksever] - 10https://gitlab.wikimedia.org/toolforge-repos/allblocksever/-/merge_requests/24 (owner: 10renovatebot) [14:33:20] (03merge) 10countcount: Lock file maintenance [toolforge-repos/namehistory] - 10https://gitlab.wikimedia.org/toolforge-repos/namehistory/-/merge_requests/12 (owner: 10renovatebot) [14:33:51] 10Tool-inteGraality: GoodReferenceCheck (S!) drill-down queries time out on QLever - https://phabricator.wikimedia.org/T435101 (10JeanFred) 03NEW [14:35:13] (03update) 10raymond-ndibe: tests: add behavior tests for previously untested code branches [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/362 (https://phabricator.wikimedia.org/T434166) [14:35:19] (03update) 10raymond-ndibe: tests: add behavior tests for previously untested code branches [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/362 (https://phabricator.wikimedia.org/T434166) [14:38:22] !log tlepage@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component logging [14:46:12] 10Cloud-VPS, 06tools-infrastructure-team: cloudcephosd1043 drives are very very busy - https://phabricator.wikimedia.org/T434334#12222266 (10Andrew) RAM usage on OSDs is running around 50% for older servers and much, much less than that for newer servers. It would be reasonable to increase bluestore_cache_... [14:46:44] 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: "osd.224 observed stalled read indications in DB device" - https://phabricator.wikimedia.org/T434402#12222267 (10Andrew) RAM usage on OSDs is running around 50% for older servers and much, much less than that for newer servers. It would be reasonable to incre... [14:48:30] 10Cloud-VPS, 06tools-infrastructure-team, 10Cumin, 06Infrastructure-Foundations: Enable cumin hostfile backend on cloudcumin hosts - https://phabricator.wikimedia.org/T433916#12222277 (10LSobanski) p:05Triage→03Low [14:51:18] (03update) 10raymond-ndibe: models: move cmd field out of CommonJob into job-type models [repos/cloud/toolforge/jobs-api] (increase_test_coverage) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/364 (https://phabricator.wikimedia.org/T434166) [14:51:32] (03update) 10raymond-ndibe: models: move cmd field out of CommonJob into job-type models [repos/cloud/toolforge/jobs-api] (increase_test_coverage) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/364 (https://phabricator.wikimedia.org/T434166) [14:51:35] (03update) 10raymond-ndibe: models: move cmd field out of CommonJob into job-type models [repos/cloud/toolforge/jobs-api] (increase_test_coverage) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/364 (https://phabricator.wikimedia.org/T434166) [14:53:53] 10Tool-wikinewsie: Dropdown for continents/regions on Wikinewsie - https://phabricator.wikimedia.org/T434608#12222342 (10waldyrious) This sounds very useful but I'm not sure a fixed list would be the best approach. For example, that split of European regions isn't particularly interesting to me — I'd much rather... [14:54:35] 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: "osd.224 observed stalled read indications in DB device" - https://phabricator.wikimedia.org/T434402#12222345 (10Andrew) ` root@cloudcephosd1045:/var/log# free -h total used free shared buff/cache available Mem:... [14:58:13] 10VPS-project-Codesearch: make a new frontend for codesearch - https://phabricator.wikimedia.org/T435106 (10ekrem) 03NEW [15:03:48] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12222395 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1... [15:06:29] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12222406 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1051.... [15:06:52] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12222417 (10VRiley-WMF) [15:11:43] 10VPS-project-Codesearch, 07Epic: make a new frontend for codesearch - https://phabricator.wikimedia.org/T435106#12222459 (10ekrem) [15:14:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [15:14:20] (03update) 10renovatebot: Update pnpm to v11.22.0 [toolforge-repos/namehistory] - 10https://gitlab.wikimedia.org/toolforge-repos/namehistory/-/merge_requests/11 [15:15:09] (03update) 10renovatebot: Update pnpm to v11.22.0 [toolforge-repos/allblocksever] - 10https://gitlab.wikimedia.org/toolforge-repos/allblocksever/-/merge_requests/23 [15:15:26] (03update) 10renovatebot: Update pnpm to v11.22.0 [toolforge-repos/checkusertools] - 10https://gitlab.wikimedia.org/toolforge-repos/checkusertools/-/merge_requests/18 [15:16:44] FIRING: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [15:17:09] (03update) 10renovatebot: Update pnpm to v11.22.0 [toolforge-repos/rangetree] - 10https://gitlab.wikimedia.org/toolforge-repos/rangetree/-/merge_requests/50 [15:18:17] (03update) 10renovatebot: Update pnpm to v11.22.0 [toolforge-repos/wiki-mail-verify] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-mail-verify/-/merge_requests/34 [15:19:54] (03update) 10renovatebot: Update python Docker tag to v3.14.7 [toolforge-repos/flaggedrevspromotioncheck] - 10https://gitlab.wikimedia.org/toolforge-repos/flaggedrevspromotioncheck/-/merge_requests/18 [15:21:12] (03update) 10dcaro: Alloy relabeling [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1371 (https://phabricator.wikimedia.org/T432969) (owner: 10tlepage) [15:21:44] RESOLVED: MaintainDBUsersManyErrors: Maintain-dbusers is having sustained errors - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/MaintainDBUsersManyErrors - https://grafana.wikimedia.org/d/ae240a06-c13e-49f3-b12c-58432c551e85/wmcs-maintain-dbusers - https://alerts.wikimedia.org/?q=alertname%3DMaintainDBUsersManyErrors [15:24:20] (03update) 10dcaro: [T401993] Add a "description" field to the deployment [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/181 (https://phabricator.wikimedia.org/T401993) (owner: 10mahveotm) [15:27:00] (03merge) 10dcaro: [T401993] Add a "description" field to the deployment [repos/cloud/toolforge/components-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/merge_requests/181 (https://phabricator.wikimedia.org/T401993) (owner: 10mahveotm) [15:27:37] 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: "osd.224 observed stalled read indications in DB device" - https://phabricator.wikimedia.org/T434402#12222574 (10Andrew) I think cache resizing is a dead end. We already try to use autosizing with a cap of 6G per osd. [15:30:17] (03open) 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce: components-api: bump to 0.0.207-20260817152713-c090522b [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1375 (https://phabricator.wikimedia.org/T401993) [15:30:21] (03PS6) 10Arendpieter: Use IDP for authentication [labs/striker] - 10https://gerrit.wikimedia.org/r/1250537 (https://phabricator.wikimedia.org/T359554) [15:30:40] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component components-api [15:35:35] !log dcaro@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-api [15:36:05] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12222609 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1... [15:36:43] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12222612 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1... [15:39:03] !log dcaro@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component components-api [15:43:47] !log dcaro@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-api [15:44:34] (03approved) 10dcaro: components-api: bump to 0.0.207-20260817152713-c090522b [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1375 (https://phabricator.wikimedia.org/T401993) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [15:44:44] (03merge) 10dcaro: components-api: bump to 0.0.207-20260817152713-c090522b [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1375 (https://phabricator.wikimedia.org/T401993) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [15:53:56] (03merge) 10dcaro: [T401993] Add deployment descriptions to the CLI [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/90 (https://phabricator.wikimedia.org/T401993) (owner: 10mahveotm) [15:55:18] (03update) 10tlepage: Alloy relabeling [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1371 (https://phabricator.wikimedia.org/T432969) [15:56:24] (03merge) 10tlepage: Alloy relabeling [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1371 (https://phabricator.wikimedia.org/T432969) [15:58:38] (03open) 10dcaro: d/changelog: bump to 0.0.18 [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/92 (https://phabricator.wikimedia.org/T401993) [16:01:44] 10Toolforge, 06tools-platform-team: Deploying static apps to toolforge - https://phabricator.wikimedia.org/T435088#12222800 (10bd808) https://wikitech.wikimedia.org/wiki/Help:Toolforge/Web#Serving_static_files is the current solution for serving static files without a webserver. This is pretty different from t... [16:06:22] 10Toolforge (Push-to-Deploy), 06tools-platform-team: [components-api] add job publish support - https://phabricator.wikimedia.org/T433355#12222818 (10Raymond_Ndibe) 05Open→03Resolved [16:10:41] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [16:17:10] 10Toolforge (Push-to-Deploy), 06tools-platform-team: [components-api] add job publish support - https://phabricator.wikimedia.org/T433355#12222910 (10dcaro) 05Resolved→03In progress Forgot to update the docs: https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Changelog and https://wikitech.wikimedia.org/... [16:20:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:23:59] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12222959 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1055.... [16:25:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:30:15] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12223002 (10VRiley-WMF) Hey @fgiunchedi it seems as though that cloudvirt1054, cloudvirt1055, cloudvirt1056, cloudvirt10... [16:31:39] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12223011 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by vriley@cumin1003 for host cloudvirt1... [16:36:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:41:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:47:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:52:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:55:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:56:18] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12223134 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1056.... [16:57:00] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12223142 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1054.... [17:00:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:05:02] (03approved) 10lucaswerkmeister: Localisation updates from https://translatewiki.net. [toolforge-repos/lexeme-forms] - 10https://gitlab.wikimedia.org/toolforge-repos/lexeme-forms/-/merge_requests/45 (owner: 10l10n-bot) [17:05:05] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12223184 (10Pepe_piton) @bd808 I migrated the Paulina tool to the new build service container, and started it with --mount=none The container st... [17:05:06] (03merge) 10lucaswerkmeister: Localisation updates from https://translatewiki.net. [toolforge-repos/lexeme-forms] - 10https://gitlab.wikimedia.org/toolforge-repos/lexeme-forms/-/merge_requests/45 (owner: 10l10n-bot) [17:07:51] 10Tool-lexeme-forms, 06translatewiki.net, 07Essential-Work, 10LPL Projects (Ongoing maintenance), 10LPL sprints (S2 2026 August): l10n-bot cannot create GitLab merge requests due to gitlab-ssh migration - https://phabricator.wikimedia.org/T432838#12223205 (10LucasWerkmeister) Thanks! \o/ [17:09:13] 10Tool-lexeme-forms, 10MediaWiki-extensions-OAuth, 06MediaWiki-Core-Platform-Team (Kanban): Error “The refresh token is invalid.” after migrating Wikidata Lexeme Forms to OAuth 2 - https://phabricator.wikimedia.org/T431146#12223220 (10LucasWerkmeister) (Still no trace of the error in the logs… I’m starting t... [17:11:00] 06cloud-services-team, 10Quarry, 06tools-platform-team, 06RoadToWiki, and 3 others: "1 rows" should be "1 row" - https://phabricator.wikimedia.org/T419564#12223223 (10Framawiki) This change looks good to me. Looks like I lost merge rights since Gerrit>Github transition, so if some power user could merge th... [17:13:44] (03update) 10renovatebot: Update pnpm to v11.22.0 [toolforge-repos/namehistory] - 10https://gitlab.wikimedia.org/toolforge-repos/namehistory/-/merge_requests/11 [17:16:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:21:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:24:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:26:27] 10Toolforge, 06tools-platform-team: Deploying static apps to toolforge - https://phabricator.wikimedia.org/T435088#12223340 (10Gouvernathor) Interesting, though it doesn't fit the bill for my use case, which is to replace an existing tool with the same URL. And, it still requires me to build the files, which n... [17:26:38] 10Toolforge, 06tools-platform-team: Deploying built static apps to toolforge - https://phabricator.wikimedia.org/T435088#12223341 (10Gouvernathor) [17:28:50] 10Tool-wikinewsie: Dropdown for continents/regions on Wikinewsie - https://phabricator.wikimedia.org/T434608#12223355 (10Pharos) >>! In T434608#12222342, @waldyrious wrote: > This sounds very useful but I'm not sure a fixed list would be the best approach. For example, that split of European regions isn't partic... [17:29:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:32:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:37:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:38:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:43:13] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12223427 (10bd808) >>! In T432878#12223184, @Pepe_piton wrote: > However, the tool runs slower than with the legacy system, and after some time,... [17:43:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [17:51:21] 06tools-infrastructure-team, 06SRE, 10SRE-Access-Requests: Requesting access to wmcs-roots for bliviero - https://phabricator.wikimedia.org/T435123 (10Andrew) 03NEW [17:52:00] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12223455 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by vriley@cumin1003 for host cloudvirt1057.... [17:52:34] 06tools-infrastructure-team, 06SRE, 10SRE-Access-Requests: Requesting access to wmcs-roots for bliviero - https://phabricator.wikimedia.org/T435123#12223459 (10Andrew) See https://phabricator.wikimedia.org/T428815 for a lot of this having already been done for a different access group. [18:08:14] 06tools-infrastructure-team, 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to wmcs-roots for bliviero - https://phabricator.wikimedia.org/T435123#12223530 (10BLiviero-WMF) [18:14:22] (03update) 10renovatebot: Update pnpm to v11.22.0 [toolforge-repos/namehistory] - 10https://gitlab.wikimedia.org/toolforge-repos/namehistory/-/merge_requests/11 [18:14:32] 06tools-infrastructure-team, 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to wmcs-roots for bliviero - https://phabricator.wikimedia.org/T435123#12223552 (10Andrew) @mark in theory Belinda is the approver of record for this group but we would like your vote before adding her. Thank you! [19:03:28] 10Toolforge (Push-to-Deploy): deploy-to-toolforge.yaml cannot be used with GitLab CI that defines `default: image:` - https://phabricator.wikimedia.org/T435130 (10LucasWerkmeister) 03NEW [19:08:54] (03open) 10lucaswerkmeister: toolforge-cd: move image into jobs [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/95 (https://phabricator.wikimedia.org/T435130) [19:39:20] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12223845 (10Pepe_piton) Thank you @bd808. The big step up on August 12 probably reflects the uwsgi.ini settings I added that day, while still usi... [20:08:29] 10Tool-wikinewsie: Wikinewsie doesn't do proper image attribution - https://phabricator.wikimedia.org/T435136 (10Pharos) 03NEW [20:11:33] 10Cloud-VPS (Debian Bullseye Deprecation): Shutdown timeline extension requests for maintained projects - https://phabricator.wikimedia.org/T434103#12223945 (10bd808) deployment-prep will need more time. {T433998} and {T401839} are places where WMCS folks should be able to see some amount of progress on planing... [20:12:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [20:13:04] 10Cloud-VPS (Debian Bullseye Deprecation): Shutdown timeline extension requests for maintained projects - https://phabricator.wikimedia.org/T434103#12223952 (10bd808) [20:13:05] 10Cloud-VPS (Debian Bullseye Deprecation), 10Beta-Cluster-Infrastructure, 06Release-Engineering-Team (Doing 😎): Planning for a plan to migrate Beta Cluster again - https://phabricator.wikimedia.org/T433998#12223951 (10bd808) [20:22:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [20:31:21] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12224036 (10Pepe_piton) While tinkering with Grafana, I see that the application's memory usage is actually low (around 500 MB of 6GB allocated),... [20:37:22] 10Cloud-VPS (Debian Bullseye Deprecation), 06tools-infrastructure-team: Cloud VPS Debian Bullseye deprecation - https://phabricator.wikimedia.org/T401804#12224062 (10bd808) [20:43:57] 10Cloud-VPS (Debian Bullseye Deprecation), 10Continuous-Integration-Infrastructure: Re-build integration-cumin.integration.eqiad1.wikimedia.cloud on something newer than bullseye - https://phabricator.wikimedia.org/T433592#12224109 (10bd808) This seems to be the only Bullseye box we still need to get out of In... [21:07:07] !log andrew@cloudcumin1001 admin START - Cookbook wmcs.ceph.roll_reboot_mons (T434750) [21:13:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [21:19:50] !log andrew@cloudcumin1001 admin END (PASS) - Cookbook wmcs.ceph.roll_reboot_mons (exit_code=0) (T434750) [21:20:25] !log andrew@cloudcumin1001 admin START - Cookbook wmcs.ceph.roll_reboot_osds (T434750) [21:23:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [21:28:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [22:54:19] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12224408 (10bd808) >>! In T432878#12224036, @Pepe_piton wrote: > While tinkering with Grafana, I see that the application's memory usage is actua... [23:28:42] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12224456 (10Pepe_piton) Thank you @bd808! I just tried your suggestion: ` toolforge webservice --cpu=1 --mem=1Gi --replicas=6 --mount=none build... [23:46:04] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12224504 (10bd808) >>! In T432878#12224456, @Pepe_piton wrote: > Thank you @bd808! I just tried your suggestion: > And it performs really much be...