[00:25:00] currently there's an optional place to leave an email address if you're willing to do a more in depth interview. but no links to any other domains. (re @wmtelegram_bot: @jeremy_b: I don't know that we have ever tried to talk to Legal about something like that. I would guess the major conc...) [00:26:42] @jeremy_b: That seems like an opt-in that should be fine. [00:27:01] plus it's data you're explicitly providing [05:21:27] ok great. so still have to resolve the top of the terms seem to indicate it needs a statement but 7.2 doesn't say that. [08:43:16] !log tools.stashbot migrating historic irc-* indexes to the new opensearch cluster T435223 [08:43:21] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.stashbot/SAL [08:43:22] T435223: Migrate storage to new OpenSearch v2 Toolforge cluster - https://phabricator.wikimedia.org/T435223 [09:20:22] Hi, I see status is ok but there appears to be data missing in the replica dbs since around 05:50 (bot started failing to find user edit counts around that time). I see https://sal.toolforge.org/log/c95RMqABffdvpiTruvnT which seems likely to be related, but not much info on that ticket at least regarding tools [09:21:16] https://cluebotng-monitoring.toolforge.org/d/da6h9bs/cluebot-ng?orgId=1&from=now-6h&to=now&timezone=utc&refresh=1m&viewPanel=panel-24 is where I'm taking the time from (just got the alert for no contributions in the last hours) [09:27:54] !log tools.stashbot migrate bot with current indexes to new opensearch cluster T435223 [09:27:58] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.stashbot/SAL [09:27:59] T435223: Migrate storage to new OpenSearch v2 Toolforge cluster - https://phabricator.wikimedia.org/T435223 [09:28:49] Damianz: yes, occasional replag on the wiki replicas is expected due to the architecture of the service and it seems like you've found the cause of it this time [09:29:23] !log tools.sal update to read from new opensearch cluster T435223 [09:29:25] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.sal/SAL [09:29:27] !log tools.bash update to read from new opensearch cluster T435223 [09:29:29] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools.bash/SAL [09:40:42] taavi: thanks, I finally found that noted on one of the wikitec pages. It would be good if that could be announced on the mailing list, the linked page (https://wikitech.wikimedia.org/wiki/Map_of_database_maintenance) doesn't list anything [10:46:06] !log tools disable access to elastic cluster as announced on cloud-announce T401818 [10:46:11] Logged the message at https://wikitech.wikimedia.org/wiki/Nova_Resource:Tools/SAL [10:46:12] T401818: Upgrade Toolforge (Elastic|Open)Search cluster off of Bullseye - https://phabricator.wikimedia.org/T401818 [10:48:16] komla: Hi, I don't really understand what the recently created many tasks for "Migrate $project away from Debian Bullseye to Bookworm/Trixie" are for. For progress tracking? To reach out to maintainers? [10:48:31] It sounds like the latter because of sentences like "If you are unable to complete your migration". But you neither subscribed maintainers nor added project tags, so maintainers will never see these tasks. So I'm confused. [14:02:13] I have a Perl script that's been running for years on Toolforge that a reader reported on the weekend wasn't working.  The error message that usually comes up when I try invoking it is "no healthy upstream".  It looks like the script is getting regularly shut down and started up.  (I just tried ia again after an auto-restart and got the message [14:02:14] "The tool you are trying to access is currently receiving more traffic than it can handle. Please try again later. [14:02:14] If this issue persists, you may wish to notify the tool's maintainers about the error."  Is this happening across Toolforge scripts, or does this indicate a problem with my script? [14:12:33] !help [14:12:33] If you don't get a response in 15-30 minutes, please email the cloud@ mailing list -- https://wikitech.wikimedia.org/wiki/Help:Cloud_Services_communication [14:13:14] JohnMarkOckerblo: which tool? [14:13:52] The tool is ftl (Forward to Libraries).  It's invoked from boxes like the "Library resources" box near the bottom of the Louisa May Alcott article. [14:15:34] I'm looking at the error log; it is reporting some unusual input characters (which my script is currently verbose about in its reports to stderr) but that shouldn't make the script malfunction.  The only other reports in the file are the auto-start and auto-stop (and the auto-stops seem to happen not long after the auto-starts). [14:17:00] I don't know if there are other errors it's reporitng in other files.  (I did wonder if there's a dependency that's no longer supported, since I'm not sure how many Perl scripts of this vintage are running on the server) but I would assume that if a dependency was missing it would be reported in the error log. [14:18:39] something seems to have happened on the 19th to make that tool get a lot more traffic than it was getting before https://grafana.wmcloud.org/d/TJuKfnt4z/tool-dashboard?orgId=1&from=now-7d&to=now&timezone=utc&var-cluster=P8433460076D33992&var-namespace=tool-ftl&viewPanel=panel-27 [14:19:57] It does get hit by bots from time to time (as does the version running on my own onlinebooks server).  But I would expect it to work normally if it's not currently under a bot swarm (though maybe it still is?) [14:23:23] something's making the actual process of your tool: [14:23:24] ftl-6798cb7c8d-q82lb 0/1 CrashLoopBackOff 1108 (16s ago) 2d13h [14:23:47] (that 1108 is the number of crashes since the tool was restarted 2d13h ago) [14:25:19] making it crash*, I mean [14:25:29] I'm tail -f'ing the tool's own log; it looks like it's still getting a large number of requests per second when it's up. [14:25:38] > Warning Unhealthy 4m21s (x3324 over 2d13h) kubelet Liveness probe failed: Get "http://192.168.13.154:8000/index.html": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [14:26:18] Does that mean occasionally invocations are not halting? [14:26:55] it means that the system sees that your tool is not responding, and tries to automatically restart it, and has done that a few thousand times over the last two and a half days [14:27:10] basically, the tool is receiving so much traffic that it is locking up the lighttpd process entirely [14:28:53] OK.  If it's purely a traffic issue, would my adding triage attempts to the script help?  (I could for instance try to terminate immediately if there are input characters I would only expect from a bot and not from a legit article forward.  But I don't know if that'd help much in practice; I suspect that bot control is probably more effective at [14:28:54] the HTTP server than in the CGI script.) [14:31:01] (By Toolforge design/prolicy my script doesn't get to see the IP address of the actual invoker of the script, so I can't do anytihng to check that) [14:37:33] essentially, your options are either to increment the compute resources given to the tool (https://wikitech.wikimedia.org/wiki/Help:Toolforge/Web#Runtime_and_cpu_memory_limits), and/or to make the script use as little resources as possible to detect and early reject non-useful traffic [14:39:48] unfortunately, for many tools the only practical options have either been to move the tool behind authentication entirely, or requiring a cookie to access the page (with an error page setting the cookie with javascript and refreshing) or similar [14:42:39] Hmmm.  Ideally it should work when invoked from any Wikipedia article (without requiring being logged into a Wikimedia account), but can refuse requests that are *not* from a Wikipedia article (e.g. just called out of the blue by a bot). [14:43:42] Requiring a cookie might be OK (you already need a cookie to remember what library you prefer to be forwarded to). But I don't know if there's a good way to ensure that random users coming from Wikipedia have a cookie. [14:45:36] I do worry that if the script is getting called 30-50x its normal rate (which the graph you pointed me to implies) simply increasing its compute resources isn't really going to help. [14:50:01] I mean you can check referer [14:50:52] and use that to assume it's a probably-okay request, then do the more intrusive filtering (what taavi describes) for stuff that doesn't have a wikipedia referer header [14:52:23] In my logs, the referer is always toolforge (I was told when I put in the script that it couldn't see what was ultimately invoking it) [14:52:41] o rip [14:57:52] (actually sometimes it's wmflabs.org.  Not sure that helps-- it may reflect the two different behaviors of 'go straight to a library' or 'pick a library to go to first'-- but I'll have a closer look and see if that helps.  I'll also save the URL for the graph and see if the bot swarm goes down eventually.) [14:59:27] I don't recall Toolforge doing anything to hide referrer headers. [14:59:49] I don't believe we do [15:01:48] Maybe I'm misremembering then and it's just the originating IP that gets hidden. [15:02:21] The client IP is certainly not passed through the proxy [15:06:29] OK.  I'l look into the referer issue a bit more with my older logs.  Depending on the traffic, I'll see if it makes sense to simply refuse any request without a referer, if all legit requests coming from Wikipedia itself or my own tool would have one.  (Though do any commonly used privacy features strip out referer info?) [15:11:31] That could be promising if it works; looks like the large majority of logged requests earlier in the summer had no referer info at all.  So maybe I'll put in a settable flag to refuse referer-less requests, at least at times the tool appears to be under a bot swarm. [15:19:04] JohnMarkOckerblo: Referrer-Policy can control such things, but I think from like enwiki you should get the domain as the referrer and not the full page name. [15:19:25] Thta [15:19:58] It's fine if I just get the domain, if I'm using it to triage bot traffic and if most bots don't send a referer at all. [15:20:22] I'll try adding some referer checks and see if that helps.  Thanks! [16:13:00] andre: it is for the remaining projects that don't have a ticket or any activity yet. This is to have a 'space' to work on these projects. I only listed the admins(using an existing template from the past) without subscribing or assigning them to the ticket. [16:13:19] This is based on lessons from the grid migration project in the past. Some admins were not happy that they were subscribed or assigned to tickets without their consent. So I've avoided that this time. [16:13:59] komla: hmmm, I see. But how would they find out that there is a ticket for their product if their product is not tagged on the ticket for their product? [16:14:09] I've added the tag: Cloud VPS deprecation tag to all tickets [16:15:41] andre: that is a challenge and sort of a catch22 [16:16:17] yeah, but they maintain _their_ product and they don't maintain the product called "Cloud VPS deprecation"? [16:16:44] if there's something to fix with your product I'd expect a ticket tagged with their product, so product owners can get aware. That's all I can say :) [16:17:57] doesn't the script I suggested for creating the tasks automatically subscribe the project admins to it? [16:19:07] If a project maintainer doesn't want to be subscribed to urgent things to sort out in their project then I probably have a different understanding what ticket systems are good for, heh [16:27:22] andre: I can go back and subscribe the admins if it's fine. the tickets are not that many, really [16:27:32] komla: subscribed should honestly be fine. I recall, and generally agree, that assigned was the step too far. [16:27:36] taavi: yeah, it does. I ended up using a script I had put together in the past because it required fewer modifications. I wanted the tasks to be subtasks and also not automatically subscribe admins. I did make use of part of that template though. [16:27:58] I can't think of any reason not to subscribe the project admins here. [16:28:00] bd808: yeah, I will fix that. [16:28:06] thanks! [16:28:14] thank you komla [16:46:45] do the newly added subscribers get notifications as they were added after the task was created in the same way as if they'd been there from the start? [16:50:28] You can add me to something if you wanna test [17:10:53] perryprog: you have been subscribed into T435848 [17:11:02] T435848: Please close WMCS project "wikifunctions" - https://phabricator.wikimedia.org/T435848 [17:15:20] taavi: no indication whatsoever, not even in the phabricator notification list. I also have the email notification setting of "A task's subscribers change" set to ignore which should be the default. [17:28:55] o/ [17:29:06] any updates on https://phabricator.wikimedia.org/T432761 ? [17:32:18] @nemoralis, The attached gitlab patch has a discussion awaiting response, I think [17:34:44] * taavi blames that on gitlab's poor notifications [17:39:39] andre: any suggestions to force an update to subscribers of a phab ticket so they get some notice? [17:41:20] bliviero: not sure if I get you. Adding a comment with explicit @usernames? [17:42:01] (I mean, if someone wants something from me, I appreciate an "@aklapper: Can you answer the last comment please?" comment :) [17:48:06] if komla just added people as subscribers to a phab ticket, it appears they may not necessarily get email/notification (if i am reading the conversation correctly above in this channel). in which case, it sounds like using the @ (as you suggest) may be the more effective way to trigger a notification [17:48:34] (that last comment was for andre:) [17:51:14] @nemoralis try now? [17:51:19] (modulo ttl) [17:52:16] bliviero: phab subscribers get emails any time they are added or a ticket is edited. If the users filter those emails there's not much we can do from the phabricator side. That said, adding a /project/ to a phab ticket is not the same as adding individual users as subscribers. [18:01:17] https://www.mediawiki.org/wiki/Phabricator/Help#Receiving_updates_and_notifications [18:03:24] jah, users have to opt in to hearing about project updates but not to hearing about things they are themselves subscribed to. [18:03:35] Is I assume what andre is saying there :) [18:51:27] bliviero andrewbogott: from testing, being added as a subscriber by someone else doesn't seem to give you a notification. (bliviero did read the conversation correctly) [18:52:07] huh... I guess I can believe that [18:52:25] it's not even in the in-phabricator notification thing. It's weird. [18:55:09] bliviero: sorry if I sent you in a circle! [19:02:36] There is only a general "A task's subscribers change" on https://phabricator.wikimedia.org/settings/panel/emailpreferences/, and if you want something explicitly from user @foobar, I recommend to mention "@foobar" in your comment. [19:09:57] chlod: hmm I think techcontribs needs updating to the URL of the new opensearch cluster [20:29:35] Hey all, so I think there's something interfering with git sync upstream on the cloud infra puppet server. Anyone available to have a look? [20:31:53] cwhite: what makes you believe that? [20:32:44] ah, I see [20:33:35] `2026-08-24T20:23:26Z ERROR sync-upstream: Local diffs detected. Commit your changes!` [20:34:35] andrewbogott: reverted your uncommitted networkd restart testing from the cloud-wide server :/ [20:36:08] taavi: I definitely reset --head the puppet repo there as soon as I noticed I was typing in the wrong window and checked after... was it really still there? Or was there some branch/ownership screwup? [20:36:35] andrewbogott: the changes were in there in the working tree, uncommitted [20:36:52] did you forget --hard from the git reset or something? [20:36:53] what the heck [20:36:59] must've :/ [20:37:02] thanks for cleaning up [20:37:20] bit surprised that didn't trigger an alert like the normal rebase failure does [20:37:59] me too [20:38:17] sadly that update failure means that the changes have been making their way to VMs [20:39:08] yeah -- if we merge https://gerrit.wikimedia.org/r/c/operations/puppet/+/1318779 then that'll fix the systemd unit and then we just have the node.d file to cleanup [20:39:22] (if we don't then I can clean up the unit too) [20:39:25] Thanks, y'all! I see changes moving through again :) [20:44:22] sorry for the mess cwhite, glad you noticed it quickly [20:54:28] taavi: does the prometheus scraper delete the metric file after ingesting? [20:54:34] no [20:55:39] ok. So then this only fires on new VMs with metric file at all, right? [20:55:43] https://www.irccloud.com/pastebin/p7Aomo72/ [20:56:26] * andrewbogott thinks "[ ! -f $PROMFILE ]" means 'if the file doesn't exist' but maybe I have my logic backwards... [20:56:32] oh, you're right [20:56:43] ok :) [20:56:58] then we still have the stale file alerting problem i mentioned, unless i'm missing something else? [20:58:34] you mean that's fixed by absenting things off bookworm? [20:59:25] My expectation is that the bit with the ! - f $PROMFILE will only fire on the very first run and then never again. [20:59:39] no, the part where having a file in node.d that doesn't get touched for too long gets confused for a broken exporter script [21:00:19] oh, I see. So "else touch"? [21:00:42] I guess that works [21:01:04] what were you thinking? A second file and mv ? [21:03:20] tbh everything in my mind was way hackier than `touch` :P [21:09:29] new patch is up, assuming that puppet likes my indentations [21:09:56] or puppet linter rather [21:10:09] it does not [23:20:01] TIL MariaDB 10.5 added INSERT…RETURNING: https://mariadb.com/docs/server/reference/sql-statements/data-manipulation/inserting-loading-data/insertreturning [23:20:21] sounds promising – I’ll have to try it out, but I think I should be able to replace this extra query with that: https://gitlab.wikimedia.org/toolforge-repos/quickcategories/-/blob/491d472835/database.py#L788 [23:20:38] anyone wanna hazard a guess how safe it seems to assume that ToolsDB will be MariaDB, not MySQL, for the foreseeable future? 😅 [23:25:59] (pity that INSERT IGNORE … RETURNING doesn’t return anything if the insert fails / is ignored, so I guess to get a result in that case I need to ON DUPLICATE KEY UPDATE column=column (i.e. no-op) instead ^^ [23:26:03] )