[00:09:23] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12224526 (10Pepe_piton) > I don't expect this will change the bot access rates much, but I would recommend changing your robots.txt to: > > ROBO... [01:01:03] 10Cloud-VPS (Debian Bullseye Deprecation): Shutdown timeline extension requests for maintained projects - https://phabricator.wikimedia.org/T434103#12224583 (10komla) [01:13:54] !log andrew@cloudcumin1001 admin END (PASS) - Cookbook wmcs.ceph.roll_reboot_osds (exit_code=0) (T434750) [01:24:54] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [03:31:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [03:36:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [03:41:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [03:46:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [03:56:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [05:59:35] (03merge) 10countcount: Update pnpm to v11.22.0 [toolforge-repos/namehistory] - 10https://gitlab.wikimedia.org/toolforge-repos/namehistory/-/merge_requests/11 (owner: 10renovatebot) [05:59:40] (03merge) 10countcount: Update pnpm to v11.22.0 [toolforge-repos/wiki-mail-verify] - 10https://gitlab.wikimedia.org/toolforge-repos/wiki-mail-verify/-/merge_requests/34 (owner: 10renovatebot) [06:00:02] (03merge) 10countcount: Update pnpm to v11.22.0 [toolforge-repos/allblocksever] - 10https://gitlab.wikimedia.org/toolforge-repos/allblocksever/-/merge_requests/23 (owner: 10renovatebot) [06:00:09] (03merge) 10countcount: Update pnpm to v11.22.0 [toolforge-repos/checkusertools] - 10https://gitlab.wikimedia.org/toolforge-repos/checkusertools/-/merge_requests/18 (owner: 10renovatebot) [06:01:09] (03merge) 10countcount: Update pnpm to v11.22.0 [toolforge-repos/rangetree] - 10https://gitlab.wikimedia.org/toolforge-repos/rangetree/-/merge_requests/50 (owner: 10renovatebot) [06:17:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [06:27:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [06:37:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [06:41:43] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224869 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt... [06:42:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [06:46:39] 10Tool-inteGraality, 10Toolhub: Deduplicate integraality Toolhub records - https://phabricator.wikimedia.org/T435159 (10JeanFred) 03NEW [06:46:44] 10Tool-inteGraality, 10Toolhub: Deduplicate integraality Toolhub records - https://phabricator.wikimedia.org/T435159#12224893 (10JeanFred) [06:52:27] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224896 (10fgiunchedi) >>! In T431682#12223002, @VRiley-WMF wrote: > Hey @fgiunchedi it seems as though that cloudvirt1... [06:52:31] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224897 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1054... [06:53:26] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224898 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt... [06:59:22] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224919 (10fgiunchedi) cloudvirt1055 can't find a network device to boot from: ` Booting from HTTP Device 1: Embedded... [06:59:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [06:59:44] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224921 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1055... [07:00:14] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224922 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt... [07:00:30] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224923 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1056... [07:01:23] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224926 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by filippo@cumin1003 for host cloudvirt... [07:01:37] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224927 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by filippo@cumin1003 for host cloudvirt1056... [07:04:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [07:07:47] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224932 (10fgiunchedi) cloudvirt1056 initially failed the reimage when resetting a system that's off: ` Resetting chas... [07:13:29] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224944 (10fgiunchedi) And cloudvirt1057 also doesn't boot from the network: ` Booting from HTTP Device 1: Embedded NI... [07:16:24] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12224945 (10fgiunchedi) @VRiley-WMF to recap my findings: * cloudvirt1055 / cloudvirt1056 / cloudvirt1057 all don't fin... [08:00:00] 10Quarry, 06tools-platform-team: quarry: Upgrade Python libraries - https://phabricator.wikimedia.org/T397331#12224998 (10dcaro) it stopped installing already due to celery 5.1.2 from pypi not being supported anymore: ` Using cached celery-5.1.2-py3-none-any.whl.metadata (20 kB) WARNING: Ignoring version 5.... [08:02:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:12:07] 06tools-infrastructure-team, 06Wikidata Platform Team, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Migrate to dumps-nfs.w.o in production - https://phabricator.wikimedia.org/T432212#12225020 (10Gehel) Those mount points are not needed immediately on w[cd]qs nodes, let's remove them in this case. [08:12:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:13:15] (03update) 10raymond-ndibe: refactor: move observability fields out of CommonOptions into ObservabilityOptions [repos/cloud/toolforge/jobs-api] (rename_common_job_to_common_options) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/363 (https://phabricator.wikimedia.org/T434166) [08:23:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:24:26] (03update) 10raymond-ndibe: refactor: rename CommonJob to CommonOptions [repos/cloud/toolforge/jobs-api] (move_cmd_out_of_common_job) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/367 (https://phabricator.wikimedia.org/T434166) [08:28:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [08:48:03] FIRING: PuppetAgentStaleLastRun: Last Puppet run was over 24 hours ago on instance proxy-6 in project project-proxy - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetAgentStaleLastRun [09:05:25] FIRING: [2x] SystemdUnitFailed: wmf-pt-kill@s1.service on clouddb1013:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:15:25] RESOLVED: [4x] SystemdUnitFailed: wmf-pt-kill@s1.service on clouddb1013:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:15:55] FIRING: [5x] SystemdUnitFailed: wmf-pt-kill@s1.service on clouddb1013:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:20:40] FIRING: [7x] SystemdUnitFailed: wmf-pt-kill@s1.service on clouddb1013:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:25:40] FIRING: [9x] SystemdUnitFailed: wmf-pt-kill@s1.service on clouddb1013:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:28:29] FIRING: SystemdUnitCrashLoop: wmf-pt-kill@s7.service crashloop on clouddb1018:9100 - TODO - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitCrashLoop [09:30:40] FIRING: [11x] SystemdUnitFailed: wmf-pt-kill@s1.service on clouddb1013:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:33:29] RESOLVED: SystemdUnitCrashLoop: wmf-pt-kill@s7.service crashloop on clouddb1018:9100 - TODO - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitCrashLoop [09:35:40] FIRING: [10x] SystemdUnitFailed: wmf-pt-kill@s5.service on clouddb1016:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:39:37] (03approved) 10dcaro: start-devenv: Show that yes is the default [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/339 [09:39:42] (03merge) 10dcaro: start-devenv: Show that yes is the default [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/339 [09:40:40] FIRING: [10x] SystemdUnitFailed: wmf-pt-kill@s5.service on clouddb1016:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:40:55] FIRING: [12x] SystemdUnitFailed: wmf-pt-kill@s5.service on clouddb1016:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:41:10] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [09:41:23] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [09:41:37] (03update) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [09:42:54] (03approved) 10tlepage: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 (owner: 10dcaro) [09:45:40] FIRING: [12x] SystemdUnitFailed: wmf-pt-kill@s5.service on clouddb1016:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:46:10] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component components-cli [09:46:22] FIRING: [2x] HAProxyWikiReplicaSectionUnavailable: Wiki replica section x4 has no available servers on cloudlb1001:9900 - https://wikitech.wikimedia.org/wiki/HAProxy - TODO - https://alerts.wikimedia.org/?q=alertname%3DHAProxyWikiReplicaSectionUnavailable [09:49:00] !log dcaro@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-cli [09:50:40] FIRING: [14x] SystemdUnitFailed: wmf-pt-kill@s5.service on clouddb1016:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:51:22] RESOLVED: [2x] HAProxyWikiReplicaSectionUnavailable: Wiki replica section x4 has no available servers on cloudlb1001:9900 - https://wikitech.wikimedia.org/wiki/HAProxy - TODO - https://alerts.wikimedia.org/?q=alertname%3DHAProxyWikiReplicaSectionUnavailable [09:54:54] 10Tool-wiki-gender-stats: Update wiki_langs.json to integrate new bol.wikipedia.org - https://phabricator.wikimedia.org/T435176 (10Danya) 03NEW [09:55:03] 10Tool-wiki-gender-stats: Update wiki_langs.json to integrate new bol.wikipedia.org - https://phabricator.wikimedia.org/T435176#12225512 (10Danya) p:05Triage→03Medium [09:55:14] 10Tool-wiki-gender-stats: Update wiki_langs.json to integrate new bol.wikipedia.org - https://phabricator.wikimedia.org/T435176#12225513 (10Danya) [09:55:40] RESOLVED: [12x] SystemdUnitFailed: wmf-pt-kill@s5.service on clouddb1016:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [09:56:26] !log dcaro@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component components-cli [09:56:29] FIRING: [2x] SystemdUnitCrashLoop: wmf-pt-kill@s4.service crashloop on clouddb1025:9100 - TODO - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitCrashLoop [09:56:36] 10Tool-wiki-gender-stats: Update wiki_langs.json to integrate new bol.wikipedia.org - https://phabricator.wikimedia.org/T435176#12225517 (10Danya) [09:57:23] 10Tool-wiki-gender-stats: Update wiki_langs.json to integrate new bol.wikipedia.org - https://phabricator.wikimedia.org/T435176#12225522 (10Danya) [09:58:02] 10Tool-wiki-gender-stats: [VERSION] 0.5.0 - https://phabricator.wikimedia.org/T432922#12225525 (10Danya) [09:58:03] 10Tool-wiki-gender-stats: Update wiki_langs.json to integrate new bol.wikipedia.org - https://phabricator.wikimedia.org/T435176#12225524 (10Danya) [09:58:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [09:59:43] !log dcaro@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component components-cli [10:00:34] (03approved) 10dcaro: d/changelog: bump to 0.0.18 [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/92 (https://phabricator.wikimedia.org/T401993) [10:00:40] FIRING: [11x] SystemdUnitFailed: wmf-pt-kill@s5.service on clouddb1016:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:00:42] (03merge) 10dcaro: d/changelog: bump to 0.0.18 [repos/cloud/toolforge/components-cli] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/components-cli/-/merge_requests/92 (https://phabricator.wikimedia.org/T401993) [10:01:29] RESOLVED: [2x] SystemdUnitCrashLoop: wmf-pt-kill@s4.service crashloop on clouddb1025:9100 - TODO - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitCrashLoop [10:03:20] (03merge) 10dcaro: toolforge-cd: adds a new release workflow for packages [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/93 [10:03:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip6) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [10:03:52] 10Data-Services, 06tools-platform-team, 06Data-Persistence: Make sure multiinstance wmf-pt-kill service starts together with mariadb - https://phabricator.wikimedia.org/T435178 (10fgiunchedi) 03NEW [10:05:02] 10Data-Services, 06tools-platform-team, 06Data-Persistence: Make sure wmf-pt-kill service starts together with mariadb - https://phabricator.wikimedia.org/T435178#12225567 (10fgiunchedi) [10:05:46] 10Data-Services, 06tools-platform-team, 06Data-Persistence: Make sure wmf-pt-kill service waits for mariadb to be ready - https://phabricator.wikimedia.org/T435178#12225570 (10fgiunchedi) [10:12:42] (03PS1) 10Majavah: Add fake Grafana rendered token for metricsinfra [labs/private] - 10https://gerrit.wikimedia.org/r/1326791 [10:13:25] (03CR) 10Majavah: [V:03+2 C:03+2] Add fake Grafana rendered token for metricsinfra [labs/private] - 10https://gerrit.wikimedia.org/r/1326791 (owner: 10Majavah) [10:15:09] 10Data-Services, 06tools-platform-team, 06Data-Persistence: Make sure wmf-pt-kill service waits for mariadb to be ready - https://phabricator.wikimedia.org/T435178#12225597 (10Marostegui) Yes, it happens also on production hosts. You'd need to first start mariadb and then puppet can start pt-kill (if mariadb... [10:17:37] (03PS1) 10FNegri: api: fix remaining joinedload() for SQLAlchemy 2.0 [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326794 [10:17:37] (03PS1) 10FNegri: Fix migration chain [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326795 [10:17:37] (03PS1) 10FNegri: extra_labels: add column default to model [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326796 [10:17:38] (03PS1) 10FNegri: Add alert_routes table [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326797 (https://phabricator.wikimedia.org/T434801) [10:20:01] (03update) 10dcaro: toolforge-cd: move image into jobs [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/95 (https://phabricator.wikimedia.org/T435130) (owner: 10lucaswerkmeister) [10:20:40] RESOLVED: [2x] SystemdUnitFailed: wmf-pt-kill@s4.service on clouddb1025:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:22:11] (03approved) 10dcaro: toolforge-cd: move image into jobs [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/95 (https://phabricator.wikimedia.org/T435130) (owner: 10lucaswerkmeister) [10:22:24] (03merge) 10dcaro: toolforge-cd: move image into jobs [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/95 (https://phabricator.wikimedia.org/T435130) (owner: 10lucaswerkmeister) [10:23:09] (03PS1) 10FNegri: Replace pkg_resources with importlib.metadata [cloud/metricsinfra/prometheus-configurator] - 10https://gerrit.wikimedia.org/r/1326801 [10:23:09] (03PS1) 10FNegri: alertmanager.py: Render alert_routes as nested routes [cloud/metricsinfra/prometheus-configurator] - 10https://gerrit.wikimedia.org/r/1326802 (https://phabricator.wikimedia.org/T434801) [10:24:04] (03CR) 10CI reject: [V:04-1] alertmanager.py: Render alert_routes as nested routes [cloud/metricsinfra/prometheus-configurator] - 10https://gerrit.wikimedia.org/r/1326802 (https://phabricator.wikimedia.org/T434801) (owner: 10FNegri) [10:27:11] (03PS2) 10FNegri: alertmanager.py: Render alert_routes as nested routes [cloud/metricsinfra/prometheus-configurator] - 10https://gerrit.wikimedia.org/r/1326802 (https://phabricator.wikimedia.org/T434801) [10:27:35] (03CR) 10Majavah: [C:03+2] api: fix remaining joinedload() for SQLAlchemy 2.0 [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326794 (owner: 10FNegri) [10:28:26] (03CR) 10Majavah: [C:03+2] Fix migration chain [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326795 (owner: 10FNegri) [10:28:33] (03Merged) 10jenkins-bot: api: fix remaining joinedload() for SQLAlchemy 2.0 [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326794 (owner: 10FNegri) [10:28:56] 06tools-infrastructure-team, 06Wikidata Platform Team, 06Data-Platform-SRE (2026-08-07 - 2026-08-28), 13Patch-For-Review: Migrate to dumps-nfs.w.o in production - https://phabricator.wikimedia.org/T432212#12225644 (10fgiunchedi) >>! In T432212#12225020, @Gehel wrote: > Those mount points are not needed imm... [10:28:58] (03CR) 10Majavah: [C:03+2] extra_labels: add column default to model [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326796 (owner: 10FNegri) [10:29:37] (03Merged) 10jenkins-bot: Fix migration chain [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326795 (owner: 10FNegri) [10:29:41] (03Merged) 10jenkins-bot: extra_labels: add column default to model [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326796 (owner: 10FNegri) [10:29:57] RESOLVED: PuppetAgentNoResources: No Puppet resources found on instance metricsinfra-grafana-2 on project metricsinfra - https://prometheus-alerts.wmcloud.org/?q=alertname%3DPuppetAgentNoResources [10:32:48] 10Toolforge, 06tools-platform-team: [ci] move 'image' out of the global scope - https://phabricator.wikimedia.org/T435182 (10dcaro) 03NEW [10:33:12] (03open) 10dcaro: global: move `image` inside tests [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/96 [10:33:58] (03update) 10dcaro: global: move `image` inside tests [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/96 [10:34:43] 06cloud-services-team, 10Toolforge, 06tools-platform-team: [components-api] Add a "description" field to the deployment - https://phabricator.wikimedia.org/T401993#12225692 (10dcaro) 05In progress→03Resolved Deployed \o/ [10:39:28] 10Tool-leximap: Add Malayalam language to Leximap - https://phabricator.wikimedia.org/T433956#12225704 (10SGill) a:03Mohammed_Sadat_WMDE [10:42:52] (03CR) 10Majavah: [C:03+2] Add alert_routes table [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326797 (https://phabricator.wikimedia.org/T434801) (owner: 10FNegri) [10:43:07] (03CR) 10Majavah: [C:03+2] Replace pkg_resources with importlib.metadata [cloud/metricsinfra/prometheus-configurator] - 10https://gerrit.wikimedia.org/r/1326801 (owner: 10FNegri) [10:44:27] 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team, 07OKR-Work, 13Patch-For-Review: [metricsinfra] allow custom routing for Toolforge alerts - https://phabricator.wikimedia.org/T434801#12225722 (10fnegri) The patches above implement option 2, I tested this locally and it seems to work as int... [10:44:42] (03Merged) 10jenkins-bot: Add alert_routes table [cloud/metricsinfra/prometheus-manager] - 10https://gerrit.wikimedia.org/r/1326797 (https://phabricator.wikimedia.org/T434801) (owner: 10FNegri) [10:44:55] (03Merged) 10jenkins-bot: Replace pkg_resources with importlib.metadata [cloud/metricsinfra/prometheus-configurator] - 10https://gerrit.wikimedia.org/r/1326801 (owner: 10FNegri) [10:45:13] (03CR) 10Majavah: [C:03+2] alertmanager.py: Render alert_routes as nested routes [cloud/metricsinfra/prometheus-configurator] - 10https://gerrit.wikimedia.org/r/1326802 (https://phabricator.wikimedia.org/T434801) (owner: 10FNegri) [10:45:54] (03Merged) 10jenkins-bot: alertmanager.py: Render alert_routes as nested routes [cloud/metricsinfra/prometheus-configurator] - 10https://gerrit.wikimedia.org/r/1326802 (https://phabricator.wikimedia.org/T434801) (owner: 10FNegri) [10:51:07] 06tools-infrastructure-team, 06SRE, 10SRE-Access-Requests, 13Patch-For-Review: Requesting access to wmcs-roots for bliviero - https://phabricator.wikimedia.org/T435123#12225753 (10JMeybohm) 05Open→03Stalled [11:21:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:26:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:31:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:36:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [11:38:10] 06tools-infrastructure-team, 06Wikidata Platform Team, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Migrate to dumps-nfs.w.o in production - https://phabricator.wikimedia.org/T432212#12225942 (10fgiunchedi) [11:38:32] 06tools-infrastructure-team, 06Wikidata Platform Team, 06Data-Platform-SRE (2026-08-07 - 2026-08-28): Migrate to dumps-nfs.w.o in production - https://phabricator.wikimedia.org/T432212#12225943 (10fgiunchedi) 05Open→03Resolved a:03fgiunchedi This is done ! [11:52:30] 06cloud-services-team, 10Cloud-VPS, 06tools-infrastructure-team, 06SRE, 13Patch-For-Review: Modernise memcached systemd unit / sync, and make it presentable - https://phabricator.wikimedia.org/T273950#12225970 (10taavi) As far as I can tell it's just `profile::swift::proxy` and `profile::thanos::swift::f... [12:04:27] 10Data-Services, 06tools-platform-team, 06Data-Persistence: Make sure wmf-pt-kill service waits for mariadb to be ready - https://phabricator.wikimedia.org/T435178#12226003 (10fgiunchedi) What do you think of having wmf-pt-kill start alongside other "sidecars" like mysqld-exporter instead of unconditionally... [12:09:31] 10Cloud-VPS, 06tools-infrastructure-team, 10Cumin, 06Infrastructure-Foundations: Enable cumin hostfile backend on cloudcumin hosts - https://phabricator.wikimedia.org/T433916#12226024 (10LSobanski) a:03elukey [12:09:38] 10Cloud-VPS, 06tools-infrastructure-team, 10Cumin, 06Infrastructure-Foundations: Enable cumin hostfile backend on cloudcumin hosts - https://phabricator.wikimedia.org/T433916#12226026 (10LSobanski) 05Open→03In progress [12:14:42] 10Toolforge (Push-to-Deploy): deploy-to-toolforge.yaml cannot be used with GitLab CI that defines `default: image:` - https://phabricator.wikimedia.org/T435130#12226043 (10LucasWerkmeister) 05Open→03Resolved a:03LucasWerkmeister Can confirm it’s [working](https://gitlab.wikimedia.org/toolforge-repos/qu... [12:18:04] 06cloud-services-team, 10Toolforge (Push-to-Deploy), 06tools-platform-team: [components-api] optionally log deployments to SAL automatically - https://phabricator.wikimedia.org/T393169#12226052 (10LucasWerkmeister) Given {T401993}, this should probably be built on top of that – the CI pipeline submits an aut... [12:41:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [12:46:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [12:49:12] 10Data-Services, 06tools-platform-team, 06DBA: Decommission clouddb1013-clouddb1020 - https://phabricator.wikimedia.org/T434048#12226182 (10fnegri) 6 of the old hosts were repooled by mistake while rebooting them for {T434750}: clouddb1013, clouddb1014, clouddb1016, clouddb1017, clouddb1018, clouddb1020. I... [12:49:43] 10Data-Services, 06tools-platform-team, 06DBA: Decommission clouddb1013-clouddb1020 - https://phabricator.wikimedia.org/T434048#12226184 (10Marostegui) Thank you [12:53:54] 06cloud-services-team, 10Data-Services, 06tools-platform-team, 06DBA: Productionize new clouddb* hosts (clouddb1022-1033) - https://phabricator.wikimedia.org/T409557#12226198 (10fnegri) clouddb1032 was pooled on 2026-07-27: https://sal.toolforge.org/log/cFFDop8B8tZ8Ohr0dhaj The HAProxy logs show the h... [13:10:26] (03approved) 10dcaro: deployment: added certificate with client rights [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/365 [13:10:36] (03merge) 10dcaro: deployment: added certificate with client rights [repos/cloud/toolforge/jobs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-api/-/merge_requests/365 [13:11:29] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component jobs-api [13:11:34] !log dcaro@cloudcumin1001 toolsbeta END (FAIL) - Cookbook wmcs.toolforge.component.deploy (exit_code=99) for component jobs-api [13:14:14] (03update) 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce: jobs-api: bump to 0.0.559-20260818131053-7a0d4cd9 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1376 (https://phabricator.wikimedia.org/T432565) [13:14:23] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component jobs-api [13:14:23] (03open) 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce: jobs-api: bump to 0.0.559-20260818131053-7a0d4cd9 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1376 (https://phabricator.wikimedia.org/T432565) [13:19:16] (03merge) 10dcaro: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] (helmfileorder) - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1373 [13:19:17] (03update) 10dcaro: api-gateway: allow jobs-api to send logs to logs-api [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1365 [13:19:20] (03update) 10dcaro: Bring component helmfiles into the modern age. [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1370 (owner: 10andrew) [13:19:54] 06cloud-services-team, 10Data-Services, 06tools-platform-team, 06DBA, 13Patch-For-Review: Productionize new clouddb* hosts (clouddb1022-1033) - https://phabricator.wikimedia.org/T409557#12226286 (10fnegri) clouddb1024 and clouddb1025 should have the same config, they diverged when we needed a replace... [13:20:24] (03update) 10raymond-ndibe: Bring component helmfiles into the modern age. [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1370 (owner: 10andrew) [13:22:39] (03update) 10raymond-ndibe: k8s v1.33.13, helm v4.1.4, helmfile v1.7.3, kindest/node v1.32.11 [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/338 (https://phabricator.wikimedia.org/T433132) (owner: 10andrew) [13:22:53] (03update) 10dcaro: Bring component helmfiles into the modern age. [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1370 (owner: 10andrew) [13:28:13] !log dcaro@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component jobs-api [13:28:32] !log dcaro@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component jobs-api [13:42:18] !log dcaro@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component jobs-api [13:44:11] (03approved) 10dcaro: jobs-api: bump to 0.0.559-20260818131053-7a0d4cd9 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1376 (https://phabricator.wikimedia.org/T432565) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [13:44:17] (03merge) 10dcaro: jobs-api: bump to 0.0.559-20260818131053-7a0d4cd9 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1376 (https://phabricator.wikimedia.org/T432565) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [13:56:10] 10Toolforge, 06tools-platform-team, 05Cloud-Services-Origin-Team, 07Cloud-Services-Worktype-Project, 07OKR-Work: [jobs-api] add certificate to authenticate against logs-api - https://phabricator.wikimedia.org/T434120#12226386 (10dcaro) 05Open→03Resolved [13:56:47] (03update) 10dcaro: global: move `image` inside tests [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/96 [13:57:09] (03update) 10dcaro: global: move `image` inside tests [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/96 [13:57:39] (03update) 10dcaro: global: move `image` inside tests [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/96 (https://phabricator.wikimedia.org/T435182) [13:58:14] 10Toolforge, 06tools-platform-team: [ci] move 'image' out of the global scope - https://phabricator.wikimedia.org/T435182#12226409 (10dcaro) p:05Triage→03Low [14:01:40] 06cloud-services-team, 10Striker, 06tools-platform-team, 10CAS-SSO, 13Patch-For-Review: Use IDP for authentication in Striker - https://phabricator.wikimedia.org/T359554#12226452 (10Arendpieter) [14:01:55] 06cloud-services-team, 10Striker, 06tools-platform-team, 10CAS-SSO, 13Patch-For-Review: Use IDP for authentication in Striker - https://phabricator.wikimedia.org/T359554#12226456 (10Arendpieter) 05Open→03In progress [14:06:24] (03approved) 10dcaro: auth: Allow specifying allowed urls for superusers [repos/cloud/toolforge/api-gateway] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/api-gateway/-/merge_requests/103 [14:06:30] (03merge) 10dcaro: auth: Allow specifying allowed urls for superusers [repos/cloud/toolforge/api-gateway] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/api-gateway/-/merge_requests/103 [14:06:39] 10Toolforge, 06tools-infrastructure-team, 06tools-platform-team, 07OKR-Work: [metricsinfra] allow custom routing for Toolforge alerts - https://phabricator.wikimedia.org/T434801#12226470 (10fnegri) 05In progress→03Resolved [14:09:01] (03update) 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce: api-gateway: bump to 0.0.104-20260818140643-0b8520a0 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1377 (https://phabricator.wikimedia.org/T433930) [14:09:05] (03open) 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce: api-gateway: bump to 0.0.104-20260818140643-0b8520a0 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1377 (https://phabricator.wikimedia.org/T433930) [14:09:33] (03open) 10dcaro: logs-api: configure write/read endpoints [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1378 (https://phabricator.wikimedia.org/T432565) [14:11:39] (03update) 10dcaro: api-gateway: bump to 0.0.104-20260818140643-0b8520a0 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1377 (https://phabricator.wikimedia.org/T433930) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [14:12:03] (03close) 10dcaro: api-gateway: allow jobs-api to send logs to logs-api [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1365 [14:19:55] (03open) 10fnegri: maintainer-alerts: re-enable alerts and add label [repos/cloud/toolforge/alerts] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/70 (https://phabricator.wikimedia.org/T432863) [14:22:18] (03update) 10fnegri: maintainer-alerts: re-enable alerts and add label [repos/cloud/toolforge/alerts] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/70 (https://phabricator.wikimedia.org/T432863) [14:22:21] (03update) 10fnegri: maintainer-alerts: re-enable alerts and add label [repos/cloud/toolforge/alerts] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/70 (https://phabricator.wikimedia.org/T432863) [14:30:59] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component api-gateway [14:32:33] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12226752 (10VRiley-WMF) hey @fgiunchedi so the reason I think it may have to do with some of the scripts is because 10... [14:34:37] !log dcaro@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component api-gateway [14:35:50] !log dcaro@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component api-gateway [14:40:02] !log dcaro@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component api-gateway [14:40:43] (03approved) 10dcaro: api-gateway: bump to 0.0.104-20260818140643-0b8520a0 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1377 (https://phabricator.wikimedia.org/T433930) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [14:40:49] (03merge) 10dcaro: api-gateway: bump to 0.0.104-20260818140643-0b8520a0 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1377 (https://phabricator.wikimedia.org/T433930) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [14:41:22] 10Toolforge, 06tools-platform-team, 07OKR-Work, 13Patch-For-Review: [api-gateway] Specify allowed urls for superusers - https://phabricator.wikimedia.org/T433930#12226830 (10dcaro) 05In progress→03Resolved [14:49:25] (03approved) 10taavi: maintainer-alerts: re-enable alerts and add label [repos/cloud/toolforge/alerts] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/70 (https://phabricator.wikimedia.org/T432863) (owner: 10fnegri) [14:50:18] 10Toolforge, 06tools-platform-team: [k8s, kube-proxy] "udpIdleTimeout" KubeProxyConfiguration deprecation - https://phabricator.wikimedia.org/T373537#12226889 (10dcaro) 05Open→03Invalid This does not seem to be happening anymore, I'll close. [14:51:34] (03update) 10dcaro: logs-api: configure write/read endpoints [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1378 (https://phabricator.wikimedia.org/T432565) [14:57:28] (03update) 10andrew: Bring component helmfiles into the modern age. [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1370 [15:04:54] (03update) 10fnegri: maintainer-alerts: re-enable alerts and add label [repos/cloud/toolforge/alerts] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/70 (https://phabricator.wikimedia.org/T432863) [15:05:12] (03merge) 10fnegri: maintainer-alerts: re-enable alerts and add label [repos/cloud/toolforge/alerts] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/alerts/-/merge_requests/70 (https://phabricator.wikimedia.org/T432863) [15:08:17] 10Tool-wikinewsie, 07Spike: Explore Possible Wikinewsie Deploy Pipeline - https://phabricator.wikimedia.org/T434616#12226950 (10ZhaoFJx) A consensus has been reached to auto-deploy every commit on the main branch for now. Should the project expand in the future, a separate QA site would be considered, and a ma... [15:08:27] 10Tool-wikinewsie, 07Spike: Explore Possible Wikinewsie Deploy Pipeline - https://phabricator.wikimedia.org/T434616#12226954 (10ZhaoFJx) 05In progress→03Resolved [15:29:07] FIRING: ToolsDBHistoryLengthGrowing: ToolsDB History Length is above the desired threshold - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolsDBHistoryLengthGrowing - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolsDBHistoryLengthGrowing [15:34:07] RESOLVED: ToolsDBHistoryLengthGrowing: ToolsDB History Length is above the desired threshold - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolsDBHistoryLengthGrowing - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolsDBHistoryLengthGrowing [15:41:17] 10Toolforge, 06tools-platform-team: [toolsdb] 2026-08-18 Transaction History Length growing too much - https://phabricator.wikimedia.org/T435220 (10fnegri) 03NEW [15:41:33] 06cloud-services-team, 10Toolforge, 06tools-platform-team: [toolsdb] Transaction History Length growing too much - https://phabricator.wikimedia.org/T428139#12227139 (10fnegri) The history length is growing again, I opened a new task: {T435220}. [15:43:20] 10Toolforge, 06tools-platform-team: [toolsdb] 2026-08-18 Transaction History Length growing too much - https://phabricator.wikimedia.org/T435220#12227145 (10fnegri) [15:43:47] (03approved) 10dcaro: logs-api: configure write/read endpoints [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1378 (https://phabricator.wikimedia.org/T432565) [15:43:51] (03unapproved) 10dcaro: logs-api: configure write/read endpoints [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1378 (https://phabricator.wikimedia.org/T432565) [15:43:58] (03approved) 10dcaro: global: add log write endpoint [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/32 [15:44:06] (03update) 10dcaro: global: add log write endpoint [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/32 [15:44:09] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12227148 (10bd808) These graphs tell an interesting story about scaling Paulina: {F98990021,size=full} Throwing a lot more RAM and CPU at one c... [15:44:42] (03merge) 10dcaro: global: add log write endpoint [repos/cloud/toolforge/logs-api] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/logs-api/-/merge_requests/32 [15:44:44] (03update) 10andrew: k8s v1.33.13, helm v4.1.4, helmfile v1.7.3, kindest/node v1.32.11 [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/338 (https://phabricator.wikimedia.org/T433132) [15:46:48] (03update) 10andrew: k8s v1.33.13, helm v3.2.0, kindest/node v1.32.11 [repos/cloud/toolforge/lima-kilo] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/lima-kilo/-/merge_requests/338 (https://phabricator.wikimedia.org/T433132) [15:47:19] (03open) 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce: logs-api: bump to 0.0.42-20260818154457-1aca4397 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1379 (https://phabricator.wikimedia.org/T432565) [15:51:53] 10Toolforge, 06tools-platform-team, 07Kubernetes: [infra,k8s] replace admission controllers with an existing policy admin project - https://phabricator.wikimedia.org/T335131#12227205 (10dcaro) 05Open→03Resolved a:03dcaro We went with kyverno [15:52:07] FIRING: ToolsDBHistoryLengthGrowing: ToolsDB History Length is above the desired threshold - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolsDBHistoryLengthGrowing - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolsDBHistoryLengthGrowing [15:52:47] 06cloud-services-team (Hardware), 10Cloud-VPS, 06tools-infrastructure-team, 06DC-Ops, and 2 others: Rebalance cloudvirts out of E4 and into C8 - https://phabricator.wikimedia.org/T431682#12227216 (10Andrew) Any reason we can't just flip these over to uefi? [15:56:30] 10Tool-bash, 10Tool-sal, 10Stashbot: Migrate storage to new OpenSearch v2 Toolforge cluster - https://phabricator.wikimedia.org/T435223 (10bd808) 03NEW [15:56:34] (03update) 10dcaro: logs-api: bump to 0.0.42-20260818154457-1aca4397 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1379 (https://phabricator.wikimedia.org/T432565) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [15:57:40] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component logs-api [15:59:49] 10Toolforge, 06tools-platform-team, 07OKR-Work: Define alert conditions in Prometheus (if possible) - https://phabricator.wikimedia.org/T432863#12227291 (10fnegri) With the metricsinfra patches in {T434801}, alerts can now be routed differently based on labels. In https://gitlab.wikimedia.org/repos/cloud/too... [16:07:13] !log dcaro@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component logs-api [16:08:37] RESOLVED: ToolsDBHistoryLengthGrowing: ToolsDB History Length is above the desired threshold - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolsDBHistoryLengthGrowing - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolsDBHistoryLengthGrowing [16:23:07] FIRING: ToolsDBHistoryLengthGrowing: ToolsDB History Length is above the desired threshold - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/ToolsDBHistoryLengthGrowing - https://prometheus-alerts.wmcloud.org/?q=alertname%3DToolsDBHistoryLengthGrowing [16:29:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:32:14] (03open) 10dcaro: loki-tools: allow logs-api writting [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1380 (https://phabricator.wikimedia.org/T432565) [16:32:24] (03update) 10dcaro: logs-api: bump to 0.0.42-20260818154457-1aca4397 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1379 (https://phabricator.wikimedia.org/T432565) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [16:32:35] (03update) 10dcaro: logs-api: bump to 0.0.42-20260818154457-1aca4397 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1379 (https://phabricator.wikimedia.org/T432565) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [16:33:03] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component logging [16:33:34] !log dcaro@cloudcumin1001 toolsbeta END (FAIL) - Cookbook wmcs.toolforge.component.deploy (exit_code=99) for component logging [16:34:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:35:45] !log dcaro@cloudcumin1001 toolsbeta START - Cookbook wmcs.toolforge.component.deploy for component logging [16:37:39] (03open) 10dcaro: add logging timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 [16:38:01] (03update) 10dcaro: add logging timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 [16:39:27] 10Toolforge, 06tools-platform-team: [logging] Last deployment on tools timed out - https://phabricator.wikimedia.org/T435228 (10dcaro) 03NEW [16:40:40] 10Toolforge, 06tools-platform-team: [logging] Last deployment on tools timed out - https://phabricator.wikimedia.org/T435228#12227491 (10dcaro) p:05Triage→03Medium [16:40:57] (03update) 10dcaro: add logging timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 (https://phabricator.wikimedia.org/T435228) [16:41:35] (03update) 10dcaro: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 (https://phabricator.wikimedia.org/T435228) [16:43:03] (03update) 10dcaro: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 (https://phabricator.wikimedia.org/T435228) [16:44:13] !log dcaro@cloudcumin1001 toolsbeta END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component logging [16:44:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:46:19] (03update) 10dcaro: logs-api: configure write/read endpoints [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1378 (https://phabricator.wikimedia.org/T432565) [16:46:33] !log dcaro@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component logging [16:47:24] 10Tool-containers: Update containers-redis build using newer Ubuntu 2024.04 stack - https://phabricator.wikimedia.org/T435231 (10bd808) 03NEW [16:49:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:54:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [16:55:25] !log dcaro@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component logging [16:56:17] (03approved) 10tlepage: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 (https://phabricator.wikimedia.org/T435228) (owner: 10dcaro) [16:56:20] (03update) 10tlepage: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 (https://phabricator.wikimedia.org/T435228) (owner: 10dcaro) [16:57:16] (03approved) 10tlepage: logs-api: configure write/read endpoints [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1378 (https://phabricator.wikimedia.org/T432565) (owner: 10dcaro) [16:57:29] 10Tool-inteGraality, 10Toolhub: Deduplicate integraality Toolhub records - https://phabricator.wikimedia.org/T435159#12227589 (10bd808) [16:57:54] (03close) 10dcaro: logs-api: configure write/read endpoints [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1378 (https://phabricator.wikimedia.org/T432565) [16:59:53] (03approved) 10dcaro: loki-tools: allow logs-api writting [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1380 (https://phabricator.wikimedia.org/T432565) [16:59:59] (03merge) 10dcaro: loki-tools: allow logs-api writting [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1380 (https://phabricator.wikimedia.org/T432565) [17:00:06] !log dcaro@cloudcumin1001 tools START - Cookbook wmcs.toolforge.component.deploy for component logs-api [17:10:31] !log dcaro@cloudcumin1001 tools END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component logs-api [17:11:05] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects tool performance - https://phabricator.wikimedia.org/T432878#12227649 (10Pepe_piton) {F99004701} Yesterday around 12:00 AM UTC, CPU topped and the tool failed, so I set up 6 * (2 core + 1G). If things go f... [17:11:25] (03approved) 10dcaro: logs-api: bump to 0.0.42-20260818154457-1aca4397 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1379 (https://phabricator.wikimedia.org/T432565) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [17:11:32] (03update) 10dcaro: logs-api: bump to 0.0.42-20260818154457-1aca4397 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1379 (https://phabricator.wikimedia.org/T432565) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [17:12:09] 10Toolforge, 06tools-platform-team, 07OKR-Work, 13Patch-For-Review: [logs-api,loki] Implement a system-level logging endpoint (write) - https://phabricator.wikimedia.org/T432565#12227651 (10dcaro) 05In progress→03Resolved [17:12:15] (03merge) 10dcaro: logs-api: bump to 0.0.42-20260818154457-1aca4397 [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1379 (https://phabricator.wikimedia.org/T432565) (owner: 10group_203_bot_3c0afd0d9fd9529f3b7bc7e69a4a3bce) [17:12:22] (03approved) 10dcaro: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 (https://phabricator.wikimedia.org/T435228) [17:12:34] (03update) 10dcaro: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 (https://phabricator.wikimedia.org/T435228) [17:13:19] (03merge) 10dcaro: logging: add long timeout [repos/cloud/toolforge/toolforge-deploy] - 10https://gitlab.wikimedia.org/repos/cloud/toolforge/toolforge-deploy/-/merge_requests/1381 (https://phabricator.wikimedia.org/T435228) [17:13:26] 10Toolforge, 06tools-platform-team: [logging] Last deployment on tools timed out - https://phabricator.wikimedia.org/T435228#12227662 (10dcaro) 05Open→03Resolved [17:14:37] 10Tool-inteGraality, 10Toolhub: Deduplicate integraality Toolhub records - https://phabricator.wikimedia.org/T435159#12227686 (10bd808) > The integraality entry was deleted, including its community annotations and entry in lists This is actually expected behavior, but also probably very poorly explained. When... [17:15:58] 10Tool-inteGraality, 10Toolhub: Attempting to deduplicate integraality Toolhub records resulted in lost annotations, list removal, and the expected crawler result being ignored - https://phabricator.wikimedia.org/T435159#12227693 (10bd808) [17:20:24] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects paulina.toolforge.org performance - https://phabricator.wikimedia.org/T432878#12227703 (10bd808) [17:20:54] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team, 07Kubernetes: Toolforge: paulina.toolforge.org Python/Flask tool experiencing connection errors - https://phabricator.wikimedia.org/T425784#12227704 (10bd808) [17:21:08] 10Toolforge, 06tools-platform-team, 07OKR-Work: Define alert conditions in Prometheus (if possible) - https://phabricator.wikimedia.org/T432863#12227706 (10dcaro) Yep, there's a couple things we should keep in mind: * redeployability/reproducibility: it's a great advantage to be able to deploy toolforge by i... [17:21:53] 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: cloudceph HEALTH_WARN, multiple OSD(s) experiencing slow operations in BlueStore - https://phabricator.wikimedia.org/T429387#12227709 (10Andrew) ceph docs suggest [[ https://docs.ceph.com/en/latest/start/hardware-recommendations/#write-caches | di... [17:25:41] 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: cloudceph HEALTH_WARN, multiple OSD(s) experiencing slow operations in BlueStore - https://phabricator.wikimedia.org/T429387#12227722 (10Andrew) I'm going to try flipping the cache setting on cloudcephosd1046, which is osd.330-osd.337 [17:33:58] 10Tool-wikinewsie: More user-generated metadata for headlines - https://phabricator.wikimedia.org/T435237 (10Pharos) 03NEW [17:39:45] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects paulina.toolforge.org performance - https://phabricator.wikimedia.org/T432878#12227819 (10bd808) >>! In T432878#12224526, @Pepe_piton wrote: >> I don't expect this will change the bot access rates much, but... [17:40:01] 10Cloud-VPS, 06tools-infrastructure-team, 10Ceph: "osd.224 observed stalled read indications in DB device" - https://phabricator.wikimedia.org/T434402#12227823 (10Andrew) 224 and 168 are back in; I am working on T429387 which I suspect is the same issue. [17:49:53] 10Tool-itwiki: BotCancellazioni: rollback "cllimit=500" workaround after T433922 has been resolved - https://phabricator.wikimedia.org/T434598#12227903 (10Umherirrender) Setting limits is up to the client calling the api. [17:55:49] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects paulina.toolforge.org performance - https://phabricator.wikimedia.org/T432878#12227963 (10bd808) >>! In T432878#12227649, @Pepe_piton wrote: > Is there a maximum number of replicas? `lang=shell-session too... [17:55:51] 06cloud-services-team, 10Tool-paulina, 10Toolforge, 06tools-platform-team: Automated traffic affects paulina.toolforge.org performance - https://phabricator.wikimedia.org/T432878#12227966 (10bd808) >>! In T432878#12227649, @Pepe_piton wrote: > At this point the tool is receiving ~15 requests per second. I... [18:21:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [18:26:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [18:29:45] (03open) 10lucaswerkmeister: toolforge-cd: generate deployment description from Git [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/97 (https://phabricator.wikimedia.org/T393169 https://phabricator.wikimedia.org/T401993) [18:44:05] 10Tool-wikimedia-attribution, 06MediaWiki-API-Platform-Team, 06MediaWiki-Core-Platform-Team, 10MediaWiki-REST-API, and 2 others: Download CTA should be present only on Wikipedia projects - https://phabricator.wikimedia.org/T434214#12228183 (10Zaidusyy) a:03Zaidusyy [18:50:18] PROBLEM - Host cloudcephosd1046 is DOWN: PING CRITICAL - Packet loss = 100% [18:52:46] RECOVERY - Host cloudcephosd1046 is UP: PING OK - Packet loss = 0%, RTA = 0.82 ms [18:55:09] FIRING: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [19:00:09] RESOLVED: CephClusterInWarning: Ceph cluster in eqiad is in warning status - https://wikitech.wikimedia.org/wiki/Portal:Cloud_VPS/Admin/Runbooks/CephClusterInWarning - https://grafana.wikimedia.org/d/P1tFnn3Mk/wmcs-ceph-eqiad-health?orgId=1&search=open&tag=ceph&tag=health&tag=WMCS - https://alerts.wikimedia.org/?q=alertname%3DCephClusterInWarning [19:33:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [19:48:38] RESOLVED: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [19:54:38] FIRING: ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [20:04:38] FIRING: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [20:09:38] RESOLVED: [2x] ProbeDown: Service toolsbeta-test-k8s-haproxy-7:443 has failed probes (http_admin_beta_toolforge_org_ip4) - https://wikitech.wikimedia.org/wiki/Portal:Toolforge/Admin/Runbooks/k8s-haproxy - https://grafana.wikimedia.org/d/O0nHhdhnz/network-probes-overview?var-job=probes/custom&var-module=All - https://prometheus-alerts.wmcloud.org/?q=alertname%3DProbeDown [20:11:56] 10Tools: Broken image in the footer - https://phabricator.wikimedia.org/T435252 (10Nux) 03NEW [21:14:23] (03update) 10lucaswerkmeister: toolforge-cd: generate deployment description from Git [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/97 (https://phabricator.wikimedia.org/T393169 https://phabricator.wikimedia.org/T401993) [21:17:16] (03update) 10lucaswerkmeister: toolforge-cd: generate deployment description from Git [repos/cloud/cicd/gitlab-ci] - 10https://gitlab.wikimedia.org/repos/cloud/cicd/gitlab-ci/-/merge_requests/97 (https://phabricator.wikimedia.org/T393169 https://phabricator.wikimedia.org/T401993) [22:52:04] 10Cloud-VPS, 06tools-infrastructure-team, 06tools-platform-team, 10Maps: wmf-auto-restart[2162435]: INFO: Could not query the PID(s) of exim4: -11 - https://phabricator.wikimedia.org/T435262#12229042 (10saper) The root cause of this as confirmed by strace(1) is that lsof +c 15 -nXd DEL -e /mnt/nfs/se... [23:31:41] 10Cloud-VPS, 06tools-infrastructure-team, 13Patch-For-Review: cloudceph HEALTH_WARN, multiple OSD(s) experiencing slow operations in BlueStore - https://phabricator.wikimedia.org/T429387#12229138 (10Andrew) The lower-number osd servers (1016-1022) are about to be replaced. The higher-number were all purchase... [23:39:23] 10Cloud-VPS, 06tools-infrastructure-team, 06tools-platform-team, 10Maps: wmf-auto-restart[2162435]: INFO: Could not query the PID(s) of exim4: -11 - https://phabricator.wikimedia.org/T435262#12229178 (10saper) ` root@maps-warper5:~# gdb --args lsof +c 15 -nXd DEL -e /mnt/nfs/secondary-maps GNU gdb (Debian... [23:39:54] 10Cloud-VPS, 06tools-infrastructure-team, 06tools-platform-team, 10Maps: wmf-auto-restart[2162435]: INFO: Could not query the PID(s) of exim4: -11 - https://phabricator.wikimedia.org/T435262#12229179 (10saper) p:05Triage→03Low a:03saper [23:40:08] 10Cloud-VPS, 06tools-infrastructure-team, 06tools-platform-team, 10Maps: wmf-auto-restart[2162435]: INFO: Could not query the PID(s) of exim4: -11 - https://phabricator.wikimedia.org/T435262#12229183 (10saper) 05Open→03In progress