[00:05:05] I don’t think so, the CLI just “manually” turns the YAML contents into CLI arguments and then feeds them into argparse: https://gitlab.wikimedia.org/repos/cloud/toolforge/jobs-cli/-/blob/c76409aaf9/jobs_cli/cli.py#L1073 [00:05:34] Toolforge Components Service files, on the other hand, seem to have some kind of schema: https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/blob/f33d880516/openapi/tool-config-schema.json [00:06:42] unfortunately that's slightly different from the CLI values in some places, like the mem vs memory thing https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/blob/f33d880516/openapi/tool-config-schema.json#L372 (https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/blob/f33d880516/openapi/tool-config-schema.json#L372) [00:07:45] but the cli.py is fine enough for me, was hoping I could add a test to validate my file against the schema so I couldn't break it, but it's not really worth going to the trouble of creating the schema file myself. I'll just manually check it. Thanks! :) [01:24:37] the docs still say the default CPU/Memory cap is quite low, but the default cap (unless I requested an increase years ago and forgot, which I don't think I did...) seems to be 16 CPU/8GB RAM? [07:28:11] Hi! Accidentally pasted a huge text block into the Toolforge console instead of a command, now the session is completely frozen and blocks login. Please clear or reset all my active sessions. Thanks! (/iluvatar) [07:30:19] type: [07:30:20] `~#` (re @wmtelegram_bot: Hi! Accidentally pasted a huge text block into the Toolforge console instead of a command, now the session is complet...) [07:33:02] are you on Linux? [07:33:17] or mac? [07:33:49] It freezes during key auth before SSH connects, so escape codes don’t work. Please kill my terminal processes on the server side. [07:34:03] @jeremy_b, windows :) [07:34:56] atesz_01: in case it helps the error you were getting is a regression, will be fixed soon T438649 [07:34:56] T438649: [jobs-api,regression] Can't delete one-off jobs anymore - https://phabricator.wikimedia.org/T438649 [07:37:17] I meant on the existing frozen connection. but if it's putty then that's not the way. maybe what I said would still work for ssh in WSL. (re @wmtelegram_bot: It freezes during key auth before SSH connects, so escape codes don’t work. Please kill my terminal processes on the ...) [07:39:50] https://imgur.com/a/3KVlJYi [07:40:32] https://www.kilala.nl/index.php?id=2629 [07:45:08] there's some issue with the storage on the bastion, new sessions seem to be getting stuck getting the last puppet run to show in the MOTD when the users log in [07:45:09] looking [07:51:26] Iluvatar: btw. you have ~20 ssh sessions stuck currently [07:53:21] From just now? Those are my attempts to log in [07:57:57] Iluvatar: probably yep, can you try now? I did a bit of cleanup, but I don't know why it got stuck [07:59:54] No, still can’t log in [08:00:01] @dcaro: [08:01:22] ack [08:01:50] I'm trying to make a snapshot of an instance in Horizon for the project "wikispeech". When I click on "Create Snapshot" under "Actions" I get the error: "Error: Unable to create snapshot." The error details are "Policy doesn't allow os_compute_api:servers:create_image to be performed. (HTTP 403) (Request-ID: [08:01:51] req-03d9bd50-4c5b-4af9-b865-290aa060e8dd)". [08:13:48] I think I'll reboot the bastion, get it unstuck [08:14:09] !log tools rebooting bastion, connections will close/fail for a few seconds [08:14:13] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools/SAL [08:18:33] Iluvatar: can you try now? just freshly restarted the bastion [08:19:04] Thanks! everything works now [08:19:42] I don't know what was the issue though :/, it seemed to have trouble reading the puppet last run from disk [08:24:41] So, it’s not my fault :) [08:25:03] By the way, another question. My bot needs to get data from wiki. I can use SSE (Event Platform) + a series of API requests, or I can use Real-time Wikimedia Enterprise. Which one is better for Toolforge infrastructure in terms of server load and network traffic? Is Enterprise hosted inside the infrastructure, or is it considered external traffic for Toolforge (e.g., if it's on AWS)? [09:03:28] I would say that it depends more on what you are going to do, if you don't plan on doing a lot of traffic, it will be ok either way. Both are treated (more or less) like external networks, but yep, enterprise will have to go through some extra hops to reach, while event platform/api requests just go through our network, so might be able to get higher throughput/less latency [09:10:13] So there is no strict recommendation to stick to Event Platform, and the main concern isn't infrastructure cost, but rather speed and latency, right? My bot needs to fetch all images / templates in articles whenever a new edit is made across all projects. The data volume should be comparable to regular SSE. I've just never used Enterprise before and wanted to make sure I wouldn't accidentally break or overload anything. [09:48:04] Iluvatar: I think it should be ok if you want to try yep, would be nice to hear back how it goes too :) [09:49:27] Ok, thanks! :) [09:56:06] !log admin restart neutron-openvswitch-agent on cloudvirt1073 T438697 [09:56:14] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [09:56:14] T438697: Instances migrated to cloudvirt1073 lost their network following a live migration - https://phabricator.wikimedia.org/T438697 [10:15:37] !log tools reboot tools-k8s-worker-nfs-70 [10:15:40] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools/SAL [12:15:56] !log lucaswerkmeister@tools-bastion-15 tools.wd-image-positions deployed 45d853810e (l10n updates: pl) [12:15:59] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.wd-image-positions/SAL [14:49:00] !log quarry moving quarry.wmcloud.org proxy to k8s-clusterapi-cluster-magnum-system-kube-i98u3-kubeapi load balancer in front of new quarry cluster [14:49:03] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Quarry/SAL [17:41:23] !log admin systemctl restart mariadb@s3.service mariadb@x3.service on clouddb1022; running close to 100% RAM again T438200 [17:41:29] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Admin/SAL [17:41:29] T438200: clouddb memory alert - https://phabricator.wikimedia.org/T438200