[00:05:00] FIRING: [3x] PurgedHighEventLag: High event process lag with purged on cp7011:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [00:10:00] FIRING: [19x] PurgedHighEventLag: High event process lag with purged on cp7001:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [00:15:00] RESOLVED: [13x] PurgedHighEventLag: High event process lag with purged on cp7001:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [00:22:00] FIRING: [3x] PurgedHighEventLag: High event process lag with purged on cp7001:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [00:27:00] FIRING: [11x] PurgedHighEventLag: High event process lag with purged on cp7001:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [00:32:00] RESOLVED: [23x] PurgedHighEventLag: High event process lag with purged on cp7001:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [00:36:27] 06Traffic, 13Patch-For-Review: Move varnish pseudo-headers to vmod_var variables - https://phabricator.wikimedia.org/T373550#12235004 (10BCornwall) 05In progress→03Resolved What we have is good enough [02:31:00] FIRING: PurgedHighEventLag: High event process lag with purged on cp7016:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://grafana.wikimedia.org/d/RvscY1CZk/purged?var-datasource=magru%20prometheus/ops&var-instance=cp7016 - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [02:36:00] RESOLVED: [25x] PurgedHighEventLag: High event process lag with purged on cp7001:2112 - https://wikitech.wikimedia.org/wiki/Purged#Alerts - https://alerts.wikimedia.org/?q=alertname%3DPurgedHighEventLag [09:36:26] 10Domains, 07HTTPS, 10DNS, 06SRE, 06Traffic-Icebox: Merge Wikipedia subdomains into one, to discourage censorship - https://phabricator.wikimedia.org/T215071#12235881 (10Tgr) Domains are first of all trust boundaries on the modern web, and with our different wikis being written by different communities,... [11:39:50] cswiki and fawiki are now on thumb.wikimedia.org. I wait until Monday and then dewiki (unless objections) [11:40:57] fawiki is fun to look at when can't read the "letters" [11:41:39] How can I see if it's working? [11:41:52] The thumb links stil point to upload [11:42:05] otoh, the reason I picked it is that if people report issues, I can immediately understand, unlike basically every other language. That's also why I'm going dewiki next week. For cswiki, I've asked Martin to keep an eye on it [11:42:19] "The thumb links stil point to upload" use action=purge [11:42:42] fa.wikipedia.org thumbs link to thumb.wm.o for me [11:42:57] for local files, I try https://fa.wikipedia.org/wiki/Special:NewFiles [11:43:21] Ah, so they do, at least some. The frontpage ones doesn't [11:43:53] oh also, https://w.wiki/TecQ [11:44:33] in a day, a lot also go up after cdn caches expire [11:44:57] Is CZ esams or drmrs? IR is drmrs [11:45:31] Semi-unrelated: When you read farsi, do you need to use a bigger font, or it is just because I can't read the letters that they look small [11:45:38] let me check on the cz [11:46:03] CZ is drmrs [11:46:53] If you want a smaller wiki that goes to esams we could do DA [11:46:55] okay, most of the load will be on drmrs, right now the rate is 40 reqs/s which is cute but it'll go up [11:48:26] on the fonts. They are not that small but sometimes they have diacritics that are just hints on how to pronounce it. For example ِ is basically "e" [11:48:44] but you know the words already so they don't write it most of the time [11:50:01] That feels a little like English where if you get the first two and the last two letters correct, you don't really need to bother to much with the spelling of the rest of the word :-) [11:52:05] yeah [11:53:10] okay, if I go with dewiki next week, at max, it'll add 1.7K reqs/sec to text on esams minimum it'll be around 340 reqs/sec. Do you think it's okay or we should reimage an upload host to text in esmas? [11:54:15] based on https://w.wiki/Teeu [11:55:44] I'm honestly not entirely sure, but we can check in on the number tomorrow and Monday and make a qualified guess [12:00:09] Right now it doesn't seem like it make much of a difference to the drmrs nodes [12:01:42] let's look at it on Monday! [12:03:02] I'll check with bblack and see if there's something in particular we need to keep an eye on. Ideally I like to see that we not swapping to much in and out of the cache. [13:29:22] 10netops, 06Infrastructure-Foundations, 06SRE Observability: Prometheus rule evaluation failures (instance titan1001) - https://phabricator.wikimedia.org/T435494#12237036 (10tappof) [13:50:16] 06Traffic, 10Liberica, 10Prod-Kubernetes, 06ServiceOps, and 2 others: Migrate Wikikube k8s apiserver and services to IPIP - https://phabricator.wikimedia.org/T420436#12237166 (10JMeybohm) [14:14:41] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12237291 (10Urbanecm) Cswiki announcement: https://cs.wikipedia.org/wiki/Wikipedie:Pod_l%C3%ADpou_(technika)#c-Martin_Urbanec-20... [14:32:46] 06Traffic, 06Data-Engineering, 06Data-Engineering-Radar: Normalize URI host in Turnilo's webrequest_sampled_live - https://phabricator.wikimedia.org/T434766#12237420 (10ssingh) Adding @Fabfur so that he can comment. [14:47:47] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237508 (10ssingh) [14:48:25] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237509 (10ssingh) Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the maint window of 08:00 UTC. Can you please co... [14:52:52] 10netops, 06Infrastructure-Foundations, 06SRE: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506 (10cmooney) 03NEW p:05Triage→03Medium [14:57:15] 10netops, 06Infrastructure-Foundations, 06SRE: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506#12237562 (10cmooney) Hmm.... I removed the config for both FPCs, and requested they go to "offline", however the system alarms have not cleared: ` cmooney@re0.cr2-eqiad> show ch... [15:04:06] 06Traffic, 06Data-Engineering, 06Data-Engineering-Radar: Normalize URI host in Turnilo's webrequest_sampled_live - https://phabricator.wikimedia.org/T434766#12237590 (10Fabfur) That's very strange because we should've already fixed this (at least for port 443) with T392880 Let me investigate better [15:04:18] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237593 (10RobH) >>! In T435406#12237508, @ssingh wrote: > Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the main... [15:06:38] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237604 (10RobH) [15:07:16] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12237605 (10RobH) a:05ssingh→03RobH [15:32:42] hello traffic friends - yesterday, brett raised a question in another channel about whether something odd is potentially going on with kafka in eqiad, only affecting hosts in magru. [15:32:42] now seeing the intermittent WidespreadPuppetFailure alerts there, I'm starting to suspect that this is a network issue, possibly the result of putting the new HE circuits back to the core DCs in service yesterday? [15:33:23] that's a surprising amount of packet loss: https://www.irccloud.com/pastebin/as7ZSwJn/ [15:37:43] swfrench-wmf: yeah, that explains the purged alerts as well since they were not having a good time talking to kafka [15:37:58] topranks: ^ do we see some data at your end? [15:38:05] precisely yeah - chatting with topranks about potentially draining them as a test [15:38:12] ah ok thank you both [15:38:28] we should consider depooling magru IMO but I leave it to the experts :> [15:44:04] so interestingly we are also seeing some oddness with prometheus hosts in magru https://phabricator.wikimedia.org/T435494 [15:44:52] hnowlan: ah, that kicks in around the same time as the purged lag did [15:45:05] 11:44:18 <+icinga-wm> PROBLEM - Recursive DNS on 2a02:ec80:700:1:195:200:68:4 is CRITICAL: DNS_QUERY CRITICAL - query timed out [15:45:10] one more, that's dns7001 [15:46:16] https://grafana.wikimedia.org/goto/efvrk4v0p2hvka?orgId=default haproxy client TTFB also erratic [15:46:43] I think we should depool magru [15:47:06] that does result in a lot more latency but at least consistently bad latency I guess :[ [15:50:28] swfrench-wmf: thanks for bringing this to our attention [15:50:56] I'd like to point out obviously we don't have 60% packet loss, that would mean the site was offline completly, even a single-digit percentage of loss would likely kill it [15:51:02] but with 10 pings you can get that. [15:51:30] we definitely have some problem here. what's a little bit strange is our BFD across this circuit was clean for the first day after it came up [15:51:49] however since then it has degraded, I see a lot of up/down events. which means packet loss [15:52:43] this is exacerbated seriously by the action of BFD re-routing traffic, i.e. we lose some packets, BFD kicks in and tears all the protocols down, we have some instability, things stabalise, BFD comes back up re-routes back. anyway TL;DR we make the loss worse by responding to it with protocol tear-downs [15:52:52] which likely happened during your 10 pings [15:54:55] thanks, topranks! indeed, that just happened to be the first batch of pings I sent. other subsequent ones consistently saw something lower (i.e., 60% clearly caught "something" and BFD triggering re-routing makes sense) [15:55:16] looks like you maybe drained the links? (I see different mtr results and no more loss) [15:55:53] yes the new link in magru has been costed out for now [15:56:07] I'll do more tests and work with the carrier on it [15:56:54] awesome, thank you very much [15:57:00] sorry for the trouble folks.... probably we got complacent here at the start of this project we were being more conservative with the "bedding in" time, but seeing as they all went so well I was probably a bit over-zealous making this one live [15:59:38] totally reasonable - this one just turned out to be weird :) thanks again for quick response <3 [16:02:08] haproxy client TTFB has stabilized nicely too: https://grafana.wikimedia.org/goto/dfvrlj5u2nqiod?orgId=default [16:09:52] https://usercontent.irccloud-cdn.com/file/0jiq1KA0/grafik.png [16:09:57] caches warming up :D [16:15:59] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12238091 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host dns6002.wikimedia.org with OS trixie [16:27:25] topranks: <3 [16:27:43] Amir1: nice [16:35:27] alright, now I'm back with the question I was going to ask when I noticed the WidespreadPuppetFailure alerts and recalled the weirdness reported last night: [16:35:27] I'd like to roll out a change to an ATS Lua plugin [0]. any concerns or conflicting work if I were to do so in the next 30 minutes or so? [16:35:27] [0] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1304189 [16:36:17] fine by us, thanks! [16:38:21] awesome, thanks - I'll get started in a bit, then [17:00:11] 10netops, 06Infrastructure-Foundations, 06SRE: Power alert for cr2-eqiad old line cards - https://phabricator.wikimedia.org/T435506#12238341 (10cmooney) @VRiley was able to unseat the cards, which means the FPC slots now show as 'empty' rather than 'offline' ` cmooney@re0.cr2-eqiad> show chassis fpc... [17:31:46] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12238522 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host dns6002.wikimedia.org with OS trixie completed: - dns6002 (**WARN**) - Downtimed on Icinga/Ale... [17:40:03] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238568 (10ssingh) @RobH, @cmooney: @SLyngshede-WMF will be handling this from Traffic, just as an FYI. [18:10:53] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238743 (10cmooney) >>! In T435406#12237508, @ssingh wrote: > Thanks, @RobH. We will then aim to depool at `2026-08-26 @ 07:00 UTC` for the m... [18:41:34] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12238851 (10ssingh) >>! In T435406#12238743, @cmooney wrote: >>>! In T435406#12237508, @ssingh wrote: >> Thanks, @RobH. We will then aim to de... [18:49:54] 06Traffic, 06Infrastructure-Foundations, 06ServiceOps, 06SRE, 13Patch-For-Review: Scaling urldownloaders by adding redundancy and load balancing - https://phabricator.wikimedia.org/T429175#12238881 (10ssingh) 05Open→03Resolved a:03ssingh urldownloaders are now behind LVS as a low-traffic servic... [19:35:26] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543 (10cmooney) 03NEW p:05Triage→03High [19:53:06] 06Traffic, 06Commons: Varnish cache purges of file assets sometimes seem to be failing, at least sometimes - https://phabricator.wikimedia.org/T435283#12239059 (10ssingh) p:05Triage→03High [19:53:39] 06Traffic, 06Commons: Varnish cache purges of file assets sometimes seem to be failing, at least sometimes - https://phabricator.wikimedia.org/T435283#12239062 (10ssingh) Thanks for reporting. This has been reported to us in other tasks as well and while they resolved (either after the expiry of the cache or a... [20:19:54] 10netops, 06Infrastructure-Foundations, 06SRE Observability: Prometheus rule evaluation failures (instance titan1001) - https://phabricator.wikimedia.org/T435494#12239145 (10Scott_French) This would likely be T435543. Impact should have resolved around 15:50 when the new circuits were depooled. [20:28:45] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239173 (10cmooney) [20:29:27] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239180 (10cmooney) [20:34:43] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239204 (10cmooney) I've sent a mail to the HE noc (cc'd noc@wikimedia) to raise a ticket about this. Let's see what they say. [21:57:15] 10netops, 06Infrastructure-Foundations, 06SRE: Aug 2026: Performance issues on new HE transports to magru - https://phabricator.wikimedia.org/T435543#12239515 (10cmooney) Ticket ID HE#7216121 [23:23:04] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12239709 (10RobH) 05Open→03Resolved [23:26:46] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12239713 (10RobH)