[06:47:33] Hi! I am the admin of the Toolforge tool Paulina (paulina.toolforge.org) [06:47:34] The tool is receiving a lot of automated traffic. According to Toolviews, this is the hit count for the last week: [06:47:36] { [06:47:37]   "end": "2026-07-21", [06:47:39]   "results": { [06:47:40]     "2026-07-14": { [06:47:42]       "paulina": 112288 [06:47:43]     }, [06:47:45]     "2026-07-15": { [06:47:46]       "paulina": 111273 [06:47:48]     }, [06:47:49]     "2026-07-16": { [06:47:51]       "paulina": 122077 [06:47:52]     }, [06:47:54]     "2026-07-17": { [06:47:55]       "paulina": 136513 [06:47:57]     }, [06:47:58]     "2026-07-18": { [06:48:00]       "paulina": 158401 [06:48:01]     }, [06:48:03]     "2026-07-19": { [06:48:04]       "paulina": 149166 [06:48:07]     }, [06:48:09]     "2026-07-20": { [06:48:11]       "paulina": 177978 [06:48:13]     }, [06:48:15]     "2026-07-21": { [06:48:17]       "paulina": 159600 [06:48:19]     } [06:48:21]   }, [06:48:23]   "start": "2026-07-14", [07:21:11] andrewbogott: it did work (and works again). But occasionally some of our WMCS instance loose network connectivity and I just "soft reboot" them from the horizon menu. This fixes the issue most of the time [07:34:05] Hi, it seems the HttpRoute object for all my tools was deleted around 4 hours ago (they all started alerting as down)? Literally just re-created the object and they are all back up now, no deployments on my side happened since yesterday. [07:38:33] Damianz: that sounds quite dangerous, looking [07:39:03] I didn't check all 6 tools but it was the case for at least 4 [07:39:34] acak [07:40:29] Damianz: what are the tools? [07:41:11] cluebotng-trainer, cluebotng-editsets, cluebotng-review, cluebotng [07:42:26] https://cluebotng-monitoring.toolforge.org/d/0403b558-1fa5-4c14-b495-fd8f31fca77f/web-services if you want more exact ish times. manually verified the object was not there yet the pods where all happy and running. [07:42:35] created T432822 to keep track [07:42:35] T432822: [toolforge,prod] httproutes disappeared for 4 tools (at least) - https://phabricator.wikimedia.org/T432822 [07:43:30] These routes were manually created? or through `toolforge webservice`? [07:48:26] 'manually' (programatically but not using `toolforge`) [07:48:59] that's what I mean yes [07:50:28] Damianz: it seems to have happened ~18 UTC yesterday no? [07:51:42] hmm [07:52:01] (by looking at one of the graphs) [07:52:04] https://usercontent.irccloud-cdn.com/file/krfmLpPd/image.png [07:52:14] is that UTC? [07:52:31] yeah you're right - I know exactly what happened there now.... but not why I didn't get spammed with email until early this morning [07:52:33] it's my timezone I think [07:53:35] I think it's in UTC, but yes it started yesterday - I purged the tool accounts for T403735 but didn't run the full deploy (only the components create part) 🙈 [07:53:36] T403735: [jobs-api] use `launcher` also for health-check script commands - https://phabricator.wikimedia.org/T403735 [07:53:53] So that's my fault not something odd in tools... I'll check why alerting didn't start yelling until a few hours ago after coffee [07:55:15] aahhh, okok, I'll close the task them [07:55:16] *then [07:55:47] Sorry about that [07:55:54] np [07:56:22] we are about to merge stuff changing the way httproutes are managed (introducing `--publish` to jobs-api, so it creates them), so I got extra cautious [07:57:20] Once that lands I'll move over - we started managing ingress/httproute directly because running haproxy in webservice as a discovery layer to deal with webservices was never great [07:58:51] Ok, so it did send the email, which was ignored because the deploy was running and the new email was the repeat after 12 hours, so no real mystery [07:58:56] btw. do you rely on extra paths? or just on hostnames? (as in, being able to configure different paths under the same tool would be helpful?) [07:59:33] Nope - it would be nice for some things, but I stayed away from it as it was never supported in toolforge stuff [08:00:17] https://github.com/cluebotng/component-configs/blob/main/fabfile.py#L22 is literally what we apply, which is pretty much directly what the migrated objects (from ingress) where [08:00:47] Outside of that the only odd things left are static files (which could go) and network policies (which there is a ticket open for) [08:01:55] I think those would be covered by allowing `--publish` on the continuous jobs side yep [08:09:18] !log lucaswerkmeister@tools-bastion-15 tools.win-the-kiwi toolforge envvars create TOOL_SECRET_KEY < /proc/sys/kernel/random/uuid [08:09:20] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.win-the-kiwi/SAL [08:09:43] !log lucaswerkmeister@tools-bastion-15 tools.win-the-kiwi toolforge envvars create TOOL_OAUTH__CONSUMER_KEY 321c54d7da6e0ce319dfec8f5eb85119 [08:09:44] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.win-the-kiwi/SAL [08:10:00] !log lucaswerkmeister@tools-bastion-15 tools.win-the-kiwi toolforge envvars create TOOL_OAUTH__CONSUMER_SECRET [08:10:01] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.win-the-kiwi/SAL [08:25:19] !log lucaswerkmeister@tools-bastion-15 tools.win-the-kiwi toolforge build start --use-latest-versions https://gitlab.wikimedia.org/toolforge-repos/win-the-kiwi && webservice --mount=none --health-check-path=/healthz buildservice start [08:25:21] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.win-the-kiwi/SAL [08:58:55] !log lucaswerkmeister@tools-bastion-15 tools.cdnjs refreshed the GitHub token in cdnjs-index/tokenfile (now expires 2026-10-20) [08:58:57] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.cdnjs/SAL [09:58:10] !log lucaswerkmeister@tools-bastion-15 tools.win-the-kiwi toolforge envvars create OAUTH__CHANGE_TAG "OAuth CID: 18794" [09:58:12] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.win-the-kiwi/SAL [10:00:38] !log lucaswerkmeister@tools-bastion-15 tools.win-the-kiwi deployed 9bd2cbb20c [10:00:39] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.win-the-kiwi/SAL [10:26:27] !log lucaswerkmeister@tools-bastion-15 tools.ranker deployed 258b59cd93 (fix README) [10:26:29] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.ranker/SAL [10:35:32] !log lucaswerkmeister@tools-bastion-15 tools.speedpatrolling deployed fc6fda67f8 (fix README) [10:35:34] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.speedpatrolling/SAL [10:37:20] !log lucaswerkmeister@tools-bastion-15 tools.wd-image-positions deployed e33530c9d2 (fix README) [10:37:21] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.wd-image-positions/SAL [10:59:21] !log lucaswerkmeister@tools-bastion-15 tools.win-the-kiwi deployed 02a8568de9 [10:59:23] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.win-the-kiwi/SAL [12:46:35] !log admin add security group rules for all projects for new metricsinfra VIPs T401813 [12:46:41] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [12:46:42] T401813: Migrate metricsinfra project off of Bullseye - https://phabricator.wikimedia.org/T401813 [13:13:34] !log jeanfred@tools-bastion-15 tools.integraality Deploy 116fe90 (Override format_html_snippet in ReferenceColumn) for T428637 [13:13:36] !log jeanfred@tools-bastion-15 tools.integraality Deploy 47fd820 (Refactor: extract ReferenceCheck strategy from ReferenceColumn) for T428637 [13:13:38] !log jeanfred@tools-bastion-15 tools.integraality Deploy eaa22b6 (Add P123/S456 syntax for specific reference property) for T428637 [13:13:38] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.integraality/SAL [13:13:40] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.integraality/SAL [13:13:42] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.integraality/SAL [13:27:15] * dcaro starts thinking that we might want to make dologms log in cloud-feed instead [13:34:44] eh… I sometimes wonder the same, but at the same time, I find it interesting (and somewhat useful) to see when other people work on their tools [13:40:36] Yep, I have a similar back-and-forth between both xd [13:44:43] thank you jelto that's very useful to know. We're tracking this issue (at least, I hope it's just one issue) on https://phabricator.wikimedia.org/T432426 [15:04:19] I created a task on Phabricator for the issue I shared this morning: T432878 [15:06:33] Jorgemet: thanks! that will help following up [15:12:56] Thank you! [18:05:57] !log admin serveraction powercycle on cloudvirt1071 [18:05:59] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [18:14:03] Is it possible to link the Spring Boot Actuator/Prometheus project to grafana wmcloud? [22:24:24] !log logging logging-logstash-04 hard reboot via Horizon (T432254) [22:24:27] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Logging/SAL [22:24:28] T432254: scap on deployment-deploy04 failing to reach logging-logstash-04.logging.eqiad1.wikimedia.cloud - https://phabricator.wikimedia.org/T432254