[07:15:38] 06Traffic, 06Machine-Learning-Team (Q1 FY2026-27): [draft] CDN caching request for TTS v1 audio files - https://phabricator.wikimedia.org/T434046 (10isarantopoulos) 03NEW [07:20:37] 06Traffic, 06Machine-Learning-Team (Q1 FY2026-27): [draft] CDN caching request for TTS v1 audio files - https://phabricator.wikimedia.org/T434046#12186031 (10isarantopoulos) [09:06:40] 10netops, 06Infrastructure-Foundations, 06SRE: InboundInterfaceErrors alerts firing for Nokia switches on v25.10.1 - https://phabricator.wikimedia.org/T412733#12186371 (10ayounsi) Nokia 26.7.1 release notes have this in "resolved issues": > On the 7220 IXR-Dx, STP packets received on the mgmt0 management por... [10:30:35] 10netops, 06Infrastructure-Foundations, 06SRE: InboundInterfaceErrors alerts firing for Nokia switches on v25.10.1 - https://phabricator.wikimedia.org/T412733#12186604 (10cmooney) For the record does look to be fixed on lswtest-d8-eqiad, which we upgraded a few weeks ago {F97254377 width=400} [11:54:33] 06Traffic, 10Maps, 06SRE: Possibility to allow Wikimedia Maps usage on all Wikibase Cloud instances - https://phabricator.wikimedia.org/T429191#12186892 (10Anton.Kokh) @MSantos There are currently around 800K+ pages using geo coordinates across 140 Wikibases. 90% of these pages are concentrated in 5 of those... [12:03:01] 06Traffic, 06MediaWiki-Media-Platform-Team, 07Wikimedia-Performance-recommendation: Change the webp threshold based on access distribution - https://phabricator.wikimedia.org/T431150#12186897 (10Ladsgroup) 05Open→03Stalled I think we should wait until the patches for {T427465} are merged and deployed. [12:15:56] Hey folks :) Is there anything to do about the cert expiry alert, or y'all already on it? [12:20:43] 06Traffic, 06ServiceOps, 07Epic, 05FY2025-26 KR 5.1, and 2 others: CDN Backend API: recognized centralauth tokens in haproxy - https://phabricator.wikimedia.org/T424831#12186934 (10Tgr) [12:50:47] 10netops, 06Infrastructure-Foundations: ulsfo: upgrade switches to SR-Linux 26.7 - https://phabricator.wikimedia.org/T434068 (10ayounsi) 03NEW [12:55:24] 10netops, 06Infrastructure-Foundations: codfw: upgrade Nokia E/F switches to SR-Linux 26.7 - https://phabricator.wikimedia.org/T434070 (10ayounsi) 03NEW [13:49:05] hello traffic friends - following the etcd switchover yesterday, we're ready to depool etcd client traffic [0] from eqiad and proceed with reimages. that means merging: [13:49:05] * https://gerrit.wikimedia.org/r/c/operations/dns/+/1321090 [13:49:05] * https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321089 [13:49:05] just like codfw last week, the latter will require PyBal restarts in eqiad and the former will (eventually) require liberica restarts. would anyone be able to assist with that, or be available for me to drive and escalate if something goes wrong? [13:49:05] [0] https://wikitech.wikimedia.org/wiki/Etcd/Main_cluster#Depool_etcd_client_traffic_from_a_cluster [13:51:10] swfrench-wmf: yeah happy to. when are you planning to do this? [13:51:23] asking because we have some meetings today so want to make sure we fit this [13:54:19] sukhe: great, thank you! I am ready whenever works for you all. ideally not too late in the day, so that we can get some portion of the reimages done today :) [13:55:44] (i.e., ideally by 19:00 UTC or so, but I totally understand if it needs to wait longer than that) [14:01:34] swfrench-wmf: checking on who can help you and will follow up here [14:02:09] sukhe: sounds good, thanks [14:14:44] swfrench-wmf: starting in ~15 mins or so [14:15:24] sukhe: oh! yeah, that totally works - thank you :) [14:15:39] too early or too late? :P [14:16:35] no that's awesome - just right :) [14:16:45] cool! [14:33:01] cjd91: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321089 [14:33:04] we will merge this first [14:33:15] and then test this by restarting pybal on lvs1020 [14:33:23] and then if that looks good, lvs1019, 18, 17 [14:34:04] and once that is done, we will restart liberica (which we will discuss then) [14:34:15] 10:33:14 < sukhe> and then test this by restarting pybal on lvs1020 [14:34:16] cjd91: sukhe: thank you for your help. I'll move forward with the dns patch in parallel, and will let you know when liberica is ready for restarts [14:34:36] after we do this, we will check the output of pybal and run a command to check for the pools, which we will discuss there [14:48:08] cjd91: let's start with the restart of pybal in lvs1020 [14:48:37] ok [14:52:34] pybal's been restarted on lvs1020 [14:53:11] cjd91: check the journal output [14:53:20] and also the output of a random pool, just to check if things are good [14:53:42] curl localhost:9090/pools/textlb6_443 [14:54:06] or any random pool and see if you get an output basically [14:54:32] and if yes, then proceed with lvs1019-18-17 (any order is fine) [14:57:19] https://www.irccloud.com/pastebin/dGIkrbtg/ [14:57:48] https://www.irccloud.com/pastebin/Yfzw36gV/ [14:57:49] yeah [14:57:59] unfortunately we will have a few reaslservers/healthchecks failing [14:58:22] as long as everything else looks fine and pybal can connect to etcd (which it can given the output above) [14:58:27] we should just move ahead. thanks for checking. [15:01:06] just so I know this in the future, how do I find a random pool? [15:01:34] cjd91: curl localhost:9090/pools/ will list all the pools [15:01:43] you just pick a random one from there basically [15:02:09] thanks [15:02:18] https://www.irccloud.com/pastebin/JeWNttQT/ [15:02:44] https://www.irccloud.com/pastebin/yEaBdNr2/ [15:03:55] +1 [15:09:52] pybal's been restarted on lvs1018 and lvs 1017. I checked the journal and a random pool on each host [15:10:59] cjd91: ok awesome [15:11:15] now we have to restart the Liberica hosts. please check with Scott if you can proceed [15:11:29] restart the _services_ on the liberica hosts [15:12:13] swfrench-wmf: are we good to move on to restarting liberica? [15:13:00] cjd91: so, I still see tcp connections on conf1007 from lvs1017 and 1018. is it possible that puppet did not run before the restart? [15:14:16] that's possible, yes. I'll retry on lvs1017 and lvs1018 [15:15:36] cjd91: if you look in the puppet agent log or at the contents of /etc/pybal/pybal.conf, you should be able to tell. [15:15:56] for example, on lvs1017, the latter still contains references to conf1007 [15:17:16] the same is true of lvs1018 [15:17:35] ... very strange. I wonder if I messed something up in my patch? [15:17:55] oh ... [15:18:12] did we?! [15:18:17] cjd91: sukhe: apologies, it looks like https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321089 was never merged [15:18:28] ohh ok [15:18:36] (sorry, I thought y'all were doing that) [15:18:50] yeah my bad for not communicating it properly [15:18:52] cjd91: please merge that [15:19:17] so, to recap: that gets merged / puppet-merged, then puppet-agent needs run on each host, and then pybal.service can be restarted [15:23:17] swfrench-wmf: sorry about that! merging it now [15:23:33] no problem at all! it was ambiguous who what doing that :) [15:23:40] s/what/was/ [15:23:55] yeah my bad, I should have been clear that it was for cjd91 to do that [15:31:21] I just ran puppet-merge and am about to restart pybal on lvs1020-1017 [15:32:29] cjd91: just to confirm, you've also run the puppet-agent on those hosts? [15:32:38] (i.e., `sudo run-puppet-agent` or the like) [15:33:41] I'm running `puppet agent -tv` on lvs1020 just to make sure I'm doing the right thing, then I'll run `run-puppet-agent` [15:34:27] I'm running `run-puppet-agent` now [15:35:02] * swfrench-wmf thumbs up [15:35:52] as that runs, you should see diffs being applied to pybal.conf. you may also see the icinga check about pybal.conf having been updated more recently than pybal was restarted fire. [15:36:03] from a quick check of pybal.conf https://www.irccloud.com/pastebin/8gfpLFfq/ [15:36:21] nice! [15:36:32] restarting pybal.service (this and the above are on lvs1020) [15:37:44] https://www.irccloud.com/pastebin/Xq6frbOF/ [15:38:03] https://www.irccloud.com/pastebin/JDDnBSJA/ [15:38:26] moving on to lvs1019... [15:38:40] seems to be coming up! it's going to produce a tremendous amount of logs when doing do :) [15:41:27] ... and the "PyBal connections to etcd" check is green now [15:42:52] puppet's been run on lvs1019, and pybal has been restarted [15:42:59] https://www.irccloud.com/pastebin/azfBwvYy/ [15:43:16] https://www.irccloud.com/pastebin/ZD8VEf0a/ [15:44:53] * swfrench-wmf thumbs up [15:45:05] cjd91: yeah, I think +1 to proceed as long as puppet has been run on all hosts [15:45:24] (sorry in between meetings, but you have much better coverage from Scott anyway) [15:47:09] quick check of lvs1018's logs https://www.irccloud.com/pastebin/GQCkPFuU/ [15:47:26] that looks good to me! [15:48:04] https://www.irccloud.com/pastebin/R9kphnNA/ [15:48:13] cjd91: if you happen to have icinga open, you'll see that the "PyBal connections to etcd" is also green for 1018 [15:49:22] swfrench-wmf: noted and TIL! [15:49:26] as a side note, you'll also see that "PyBal connections to etcd" is red for 1017. that's because puppet already ran there (the 30m timer) and expects the connections to be pointing to codfw now [15:50:04] (i.e., the connections are there, they're just still pointed at conf1007, rather than conf2006, which will happen when you restart) [15:56:07] just restarted pybal on lvs1017 [15:56:15] https://www.irccloud.com/pastebin/0lz8yNlW/ [15:56:17] \o/ [15:56:54] https://www.irccloud.com/pastebin/w4ofqNAG/ [15:57:10] LGTM [15:57:19] and the connections to etcd check has resolved [15:58:00] 🎉 [15:58:22] cjd91: sukhe: I need to hop into a meeting in 2m. good to go from my end for liberica restarts. that would be liberia in esams, drmrs, and magru. example cookbook command is in https://wikitech.wikimedia.org/wiki/Etcd/Main_cluster#Move_other_clients [15:58:45] swfrench-wmf: ack and thank you [15:58:56] (i.e., you will need to substitute the alias for each site) [15:59:07] swfrench-wmf: thanks for all the help [15:59:16] thank you both for all the help as well! :D [15:59:21] cjd91: > sudo cookbook sre.loadbalancer.upgrade -t TXXXXXX --seamless --alias liberica-ulsfo --reason 'Clear control-plane connections to etcd' restart [15:59:24] back in 30m [16:00:02] let's adapt this for esams, magru, drmrs and the task ID and I will +1 the command for review [16:09:38] `sudo cookbook sre.loadbalancer.upgrade --query 'C:liberica and (*.esams.wmnet or *.magru.wmnet or *.drmrs.wmnet)' -t T428495 --seamless --reason 'Clear control-plane connections to etcd' restart` [16:09:39] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [16:11:32] cjd91: yeah +1. there is a batch max of 1, so it should prevent multiple sites being affected together [16:14:34] the cookbook didn't like the Class bit, apparently [16:14:42] https://www.irccloud.com/pastebin/9HsG9WSx/ [16:15:02] ok [16:15:22] let's do A:liberica-esams, -magru, -drmrs and so on [16:15:30] ok [16:15:30] perhaps there is also value in doing one by one [16:16:07] starting with esams: `sudo cookbook sre.loadbalancer.upgrade --query 'A:liberica-esams' -t T428495 --seamless --reason 'Clear control-plane connections to etcd' restart` [16:16:07] T428495: Migrate conf* hosts away from bullseye - https://phabricator.wikimedia.org/T428495 [16:16:15] =1 [16:16:17] +1 [16:18:33] esams is done [16:18:44] moving on to drmrs [16:19:23] that was quick [16:19:55] I thought the cookbook took longer but maybe that was when I was staring at it :P [16:20:16] +1 for moving on [16:21:56] drmrs is done https://www.irccloud.com/pastebin/j090JS9a/ [16:22:30] cool [16:23:05] moving on to magru [16:24:19] bblack: regarding timeouts in varnish per backend. Do we differnetiate between appserver-ro and appserver-rw or some modern k8s equivalent of that? Or is all about mw-web/mw-api etc only? [16:24:48] that may be a way to differnetiate between GET and POST [16:24:58] yeah perhaps [16:25:11] it's really at multiple levels, too. [16:25:31] GET 60s, POST 200s I think. https://wikitech.wikimedia.org/wiki/HTTP_timeouts [16:25:40] but I'm guessing Varnish enforces only 200s, not even 60s [16:25:43] it's simpler to talk about the "varnish backend timeout", but really it's ATS and Varnish both with sep timeouts [16:26:15] sukhe: magru is done [16:26:27] cjd91: thanks! swfrench-wmf ^ [16:26:35] hm.. varnish-frontend/text: TTFB 65s. ATS-backend: 180s [16:26:46] ideally, in that layered view (haproxy -> Varnish -> ATS -> envoy -> MW) [16:26:49] that suggests a post request taking 66-200s does not have its response sent to the user [16:27:16] well, and there's a difference, for some of the layers, between TTFB and between-bytes-timeout, too [16:27:17] given that we don't send the first byte until basically everything is done. [16:27:31] neither one is an overall limit [16:27:33] for MW that is, yeah, Swift I imagine streams [16:28:06] anyways, in an ideal world (hah), the timeouts should shrink inwards [16:29:07] they do mostly. 200s < 201s < 202s < 203s [16:29:08] (so that no outer layer has to actually hit its timeout limit, if the next layer inwards is correctly enforcing its own limit. a given layer only has to use its timeout when the next layer down is misbehaving, basically) [16:29:16] some of the laters don't differnatiate get/post, and some are cpu vs wall time [16:30:09] hm.. I guess the varnish ttfb is enforced only on cacheable objects, not pass/hit-for-pass? [16:30:33] ttfb and such would be enforced on all requests [16:30:36] that would explain why people havent complained about POST requests losing their response [16:31:12] even within the CDN stack, we don't have the desirable "decreasing timeouts inwards" property, anymore [16:31:54] for haproxy->varnish->ATS, connect is 3->(3|5)->10 [16:32:18] TTFB haproxy 180s > varnish text 65s > ats 180s [16:32:20] ttfb is 180->(65|35)->180 [16:32:31] yeah, it's weird, hence I'm assuming it's not what it says [16:32:44] it probably is what it says, or at least was last the docs were updated [16:32:45] let me try to produce a post that takes >65s [16:32:58] first_byte_timeout: 65s in puppet that bit checks out [16:36:22] but the categorization of requests (e.g. "all reqs" vs "by-cluster" vs "by-method" vs "by-backend" vs "videoscaler vs api", etc) is different at the different layers, and the kinds of timeouts that can even be enforced (e.g. ttfb vs between-bytes vs whole response, etc)... they all categorically vary across layers too [16:36:40] it's very hard to line it all up in a rational way [16:36:45] but we could do better than we're doing! :) [16:38:16] so I see plenty of parses in logstash that took > 66s or over. [16:38:28] trying to reproduce one in the browser now [16:38:40] cjd91: sukhe: amazing, thank you both very much! (just got out of my meeting) confirmed that no further liberica connections remain :) [16:39:30] https://www.wikidata.org/wiki/Wikidata:Database_reports/Identical_VIAF_ID consistently seems to take more than 60 seconds, so as a GET request, that correctly gets killed proactively by php-excimer within MediaWiki and higher layers never have to kick in. [16:39:58] well, another dimension to consider: at least varnish considers TTFB to be the time to the very first response byte of a request, which is a header line [16:40:23] so if mediawiki were to immediately emit "200 OK" and then stall before sending more, that's no longer ttfb, that's between_bytes [16:40:25] yeah, mw wont' send anything until the page is done parsing, it can't because it could be anything up to that point [16:40:29] or "successive read" [16:40:49] no streaming from php except when proxying binaries from swift or thumbor [16:41:32] status can turn into HTTP 500 the second an uncaught error or db error happens [16:41:53] right [16:42:10] at which point we have to send differen cache-control as well [16:42:16] our CDN's between_bytes could probably already be lower in almost all cases I think [16:43:23] and connect timeout is per-hop, so the "decrease inwards" logic doesn't really apply there [16:43:34] Request served via cp3066 cp3066, Varnish XID 684682586 [16:43:35] Error: 503, Backend fetch failed [16:43:51] yep, as I suspected, MW keeps going but Varnish cuts it off, so this probably is affecting edits and such [16:44:03] (IOW, none of these proxies are capable of stalling the establishment of a client-side connection until they've finished establishing a backend connection) [16:44:05] the POST isn't allowed to finish its 65-180s budget [16:44:34] why is the ttfb 180s at haproxy and ats when varnish enforces 65s? [16:44:45] good question! [16:45:10] envoy is: 203 seconds (appserver) / 65 seconds (api) / 86402.5 seconds (jobrunner, videoscaler) [16:46:05] seems like ATS ttfb should >203s, and every layer out from there should increase [16:46:22] but arguably we don't want to do that now, because it would make some adverse situations worse than they are today [16:46:48] if we already have a ~65s limit, if anything we should reduce in the other two layers to ~match it with slight inwards decreases. [16:47:26] e.g. haproxy 69s, varnish 67s, ats 65s [16:48:15] someone's going to have to do a lot of digging to rationalize that info first I think, though [16:48:29] maybe make a better chart on that wikitech page, that tries to make columns that line up categorically [16:48:53] ack, not everything is stacked though. so jobrunners don't go through varnish for example. [16:49:09] even if ignore jobrunners and videoscalers [16:49:11] and the php process is generally expected to live slightly longer than the web server response, because of DeferredUpdates. [16:49:27] we flush the response first to last byte, but then we don't exit yet. [16:49:33] envoy: 203 seconds (appserver) / 65 seconds (api) -> apache 202 seconds (appserver, api, parsoid) [16:49:41] so that layer has a slightly higher limit again after going down to 60 [16:49:52] parsoid is part of the api limit at envoy? [16:49:58] or the appserver limit? [16:50:04] parsoid as a cluster afaik doesn't exist anymore [16:50:17] we do have a mw-parsoid pod now, but that's for ad-hoc testing. [16:50:19] ah I figured that meant specific API paths for new parsoid [16:50:35] we used to have a separate appserver cluster for Parsoid.js and RESTBase to call into. [16:50:45] now they call mw-api-ext / mw-api-int via rest.php same as other APIs [16:50:47] still, apache should be tighter than envoy. for api, it's looser [16:51:55] and then the GET/POST distinction seem to only exist all the way at the bottom [16:52:04] maybe for some of the other layers it's a possible split, too, I'm not sure [16:52:08] https://codesearch.wmcloud.org/puppet/?q=envoy%3A%3Aupstream_response_timeout&files=mediawiki%7Cmw%7Cphp&excludeFiles=&repos= [16:52:16] I don't see a current source for that 65 envoy number [16:53:17] hieradata/common/profile/tlsproxy/envoy.yaml:profile::tlsproxy::envoy::upstream_response_timeout: 65.0 [16:53:26] not sure exactly what that controls [16:53:39] oh that might be a generic default, right [16:53:51] also: [16:53:52] hieradata/role/common/mediawiki/appserver/api.yaml sets *request* timeout [16:53:54] modules/profile/manifests/tlsproxy/envoy.pp:# @param upstream_response_timeout timeout on a request in seconds. Default: 65 [16:54:10] https://codesearch.wmcloud.org/puppet/?q=envoy%3A%3Aupstream_response_timeout%7Cenvoy%3A%3Arequest_timeout&files=mediawiki%7Cmw%7Cphp&excludeFiles=&repos=operations%2Fpuppet [16:54:38] but this presumably is all separate from the deployment charts which have their own yaml data [16:54:48] what's a request_timeout? the time for client to finish sending an entire request (headers + optional body)? [16:54:56] beats me [16:55:00] it's the only one that sets it [17:06:15] the other golden rule of timeouts (aside from 'decrease inwards, all other things/categorization being equal') is that, given all the layers involved (not just proxy layers, but also when one service calls another as a subrequest, again through potentially multiple proxy layers) [17:06:44] is that ~no service should ever internally "retry" a request that fails or times out. it should just return the failure upstream and move on. [17:07:11] the only exception is the outermost layer of the CDN, which could potentially choose to do that to paper over transient failures deeper in the stack. [17:07:32] any other layer risks multiplying traffic when things go south [17:08:15] if varnish retries once, and ATS retries each of those once, and envoy retries each of those once... you're gaining another power of two at every layer, when something deep is failing all requests temporarily, and everything gets overwhelmed [17:08:51] one little service at the bottom is just failing completely, and as a result one real-user request turns into 256 mediawiki requests or whatever [17:10:08] 06Traffic: SSL certificate expiry alerts fire too early - https://phabricator.wikimedia.org/T434116 (10BCornwall) 03NEW [17:20:53] 06Traffic, 10Maps, 06SRE: Possibility to allow Wikimedia Maps usage on all Wikibase Cloud instances - https://phabricator.wikimedia.org/T429191#12188710 (10ssingh) >>! In T429191#12186892, @Anton.Kokh wrote: > @MSantos There are currently around 800K+ pages using geo coordinates across 140 Wikibases. > 90% o... [17:27:47] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188766 (10ssingh) >>! In T414411#12183764, @RobH wrote: > pre-onsite checks: host mgmt is online and i can connect to it, but I'm having issues on its idrac interface. I want to push new ilom firmw... [17:59:26] 06Traffic, 06Machine-Learning-Team (Q1 FY2026-27): [draft] CDN caching request for TTS v1 audio files - https://phabricator.wikimedia.org/T434046#12188902 (10ssingh) Thanks for creating the task. We are discussing this in the team and will follow up here. [18:13:57] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188950 (10RobH) >>! In T414411#12188766, @ssingh wrote: >>>! In T414411#12183764, @RobH wrote: >> pre-onsite checks: host mgmt is online and i can connect to it, but I'm having issues on its idrac i... [18:23:36] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188983 (10RobH) a:03ssingh @ssingh: This is ready for traffic to take it back over. Currently 'planned' in netbox as its been repaired and provision cookbook successfully run. [18:25:38] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188995 (10ssingh) a:05ssingh→03CDobbins [18:26:19] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12188998 (10ssingh) Thanks Rob. @CDobbins will take care of it from Traffic's end. Appreciate the support here! [20:20:29] 10netops, 06Infrastructure-Foundations, 06SRE: Nokia SR-Linux: Update Homer automation to support v26 - https://phabricator.wikimedia.org/T433105#12189415 (10ayounsi) 05Open→03Resolved a:03ayounsi All done! [23:06:02] 06Traffic, 06Commons, 10MediaWiki-Uploading: Commons' file is inaccessible for some users - https://phabricator.wikimedia.org/T377202#12189833 (10Pppery)