[07:05:43] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255247 (10SLyngshede-WMF) eqsin depooled: ` slyngshede@cumin1003:~$ sudo cookbook sre.dns.admin depool eqsin -t T435406 -r "eqsin switch mi... [07:16:45] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255289 (10cmooney) [08:14:18] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255465 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b884740c-078e-4c8d-a90b-e1f10dcdef27) set by cmooney@cumin1003 fo... [08:14:56] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255466 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=a5b20745-166d-4473-bbfe-1ba217405d16) set by cmooney@cumin1003 fo... [08:17:49] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255490 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=af2e3dff-6866-4ea4-a3b1-818b0b29d804) set by cmooney@cumin1003 fo... [08:18:04] 06Traffic, 10Internet-Archive, 06SRE, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12255492 (10Fabfur) Thanks for the updates @Cyberpower678! Could you investigate if this is still an issue? If someone from archive.org wants to directly reach us (also mail ch... [08:48:17] 06Traffic, 06SRE Observability, 13Patch-For-Review: Move wikimediastatus.net 301 to ncredir - https://phabricator.wikimedia.org/T419887#12255578 (10MLechvien-WMF) Noted again during yesterday's incident T436001 [10:40:56] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12255975 (10SLyngshede-WMF) Forgot the DNS hosts: ` slyngshede@puppetserver1001:~$ sudo -i confctl select name=dns5003.* set/pooled=no The s... [10:43:18] 06Traffic, 06SRE, 07Wikimedia-Incident: Wikipedia and Wiktionary pages returning "upstream connect error or disconnect/reset before headers" - https://phabricator.wikimedia.org/T436004#12255983 (10toddbradley) Hi, apparently this issue might likely still persists somehow? Apparently I cannot access Wikipedia... [10:50:47] 06Traffic, 10Internet-Archive, 06SRE, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12256016 (10AlexisJazz) >>! In T435743#12255492, @Fabfur wrote: > Thanks for the updates @Cyberpower678! > Could you investigate if this is still an issue? I'm still seeing th... [10:50:59] 06Traffic, 06SRE, 07Wikimedia-Incident: Wikipedia and Wiktionary pages returning "upstream connect error or disconnect/reset before headers" - https://phabricator.wikimedia.org/T436004#12256018 (10Fabfur) Hi @toddbradley , thanks for the report. We've depooled the Singapore DC some hours ago due to a schedul... [11:18:43] 06Traffic: Not all images on pages in wikiprojects are loaded due to "429 Too Many Requests" error - https://phabricator.wikimedia.org/T434205#12256186 (10Aklapper) [12:19:53] 06Traffic, 13Patch-For-Review: Kapow score support on cache hosts - https://phabricator.wikimedia.org/T435799#12256445 (10GGoncalves-WMF) [12:59:37] 06Traffic, 10Internet-Archive, 06SRE, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12256588 (10Fabfur) Can't replicate, I've just archive [[ https://web.archive.org/web/20260826125613/https://en.wikipedia.org/wiki/John_of_Sterngassen | this page ]] (new page,... [13:05:12] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12256611 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=4da7b8be-5c0d-47c2-8661-fbe2c21e768c) set by cmooney@cumin1003 fo... [13:06:21] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12256613 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=f7d13814-60b9-44ed-b58e-b45be17b64e4) set by cmooney@cumin1003 fo... [13:06:39] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12256614 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=a30bcd41-744e-49bf-94f4-505e01bfb698) set by cmooney@cumin1003 fo... [14:23:04] 10netops, 06SRE Observability: Prometheus rule evaluation failures (instance titan1001) - https://phabricator.wikimedia.org/T435494#12257064 (10hnowlan) 05Open→03Resolved a:03hnowlan [14:24:40] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12257080 (10Ladsgroup) Proposal for next week's tech news. > URLs to thumbnails is changing. The domain of URLs... [14:27:50] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12257151 (10MatthewVernon) >>! In T427465#12257080, @Ladsgroup wrote: > Proposal for next week's tech news. >>... [14:36:39] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12257224 (10RobH) [14:37:39] 06Traffic, 06MediaWiki-Media-Platform-Team, 10Thumbor, 13Patch-For-Review: Add `thumb_generated` from X-Analytics to webrequest_sampled_live - https://phabricator.wikimedia.org/T435634#12257251 (10Ladsgroup) [15:21:26] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12257600 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host dns1004.wikimedia.org with OS trixie [16:02:50] 06Traffic, 10Internet-Archive, 06SRE, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12257848 (10Cyberpower678) I'm told everything is working fine on the archive side of things as well. [16:18:53] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12257888 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host dns1004.wikimedia.org with OS trixie completed: - dns1004 (**PASS**) - Downtimed on Icinga/Ale... [16:24:55] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12257936 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b58886a4-bace-4b3a-a42f-d960bc8c4bc4) set by cmooney@cumin1003 fo... [16:25:18] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12257938 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=b0850968-674c-42f7-b6a0-d37d70f8c521) set by cmooney@cumin1003 fo... [16:25:54] 10netops, 06Traffic, 06Infrastructure-Foundations, 10ops-eqsin, 06SRE: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406#12257941 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=1866892e-3c8b-4e8a-b581-189ea17a8c5b) set by cmooney@cumin1003 fo... [16:58:39] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12258158 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host dns1005.wikimedia.org with OS trixie [17:04:42] 06Traffic, 10Internet-Archive, 06SRE, 07Regression: archive.org cannot reach Wikimedia - https://phabricator.wikimedia.org/T435743#12258203 (10ssingh) I have tried to reproduce now with some obscure paths as well but "Save Page" is working fine. @AlexisJazz: can you confirm your workflow here please? Just... [17:21:45] sukhe: any idea why there are still vlan interfaces / IPs defined for eqsin LVS? [17:21:46] https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/refs/heads/production/hieradata/common/lvs/interfaces.yaml#508 [17:22:27] the hosts have no such interfaces configured: [17:22:32] https://www.irccloud.com/pastebin/7AwKTirF/ [17:23:09] topranks: yeah, for the simple reason that we (Traffic) never got to T410411. we wanted to but decided to wait during that time as the Liberica rollout was still fresh [17:23:09] T410411: Cleaning up Puppet and Netbox VLAN sub-ints on edge sites - https://phabricator.wikimedia.org/T410411 [17:23:31] and then we never got around to cleaning these in general, leaving them for Valenti.n to come and clean it up as part of when eqiad/codfw (basically everything) moves to Liberica [17:23:47] but yeah, +1 on getting rid of them; happy to prep a patch. tl;dr is that we didn't clean them up. [17:24:16] but that doesn't make sense... someone DID clean them up. it's not the big mess like in eqiad with all the vlan ints that we don't need still in use [17:24:45] I am scratching my head though - I thought the hiera defs in the above gerrit link would cause them to be created... but seemingly not? [17:25:12] topranks: on the hosts themselves they probably got cleaned up when the LVS boxes were reimaged to Liberica and Liberica never sets up the VLAN interfaces anyway [17:25:24] sukhe: ah yes that explains it [17:25:29] but we never cleaned them up in the file you linked and/or Netbox [17:25:34] 06Traffic, 06Collaboration-Services, 06SRE Observability, 10Wikimedia-Mailing-lists, 07Sustainability (Incident Followup): Consider disabling paging for lists.wikimedia.org - https://phabricator.wikimedia.org/T436049#12258388 (10hnowlan) [17:25:45] ok I suggest - at least for eqsin as I'm cleaning things up - to remove the defs for the three in eqsin [17:26:08] ok thanks, preparing a patch. [17:26:32] I'm prepping a patch to rename the eqsin vlans - so I'll just include it in that (but remove the block rather than rename) [17:26:38] thanks for the info! [17:26:41] ok! thanks [18:17:10] 06Traffic, 13Patch-For-Review: Upgrade Traffic hosts to trixie - https://phabricator.wikimedia.org/T401832#12258604 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host dns1005.wikimedia.org with OS trixie completed: - dns1005 (**WARN**) - Downtimed on Icinga/Ale... [19:04:32] topranks: I have to step out for a bit. please ping brett to repool eqsin, including in checking if things are OK [19:04:45] sukhe: ok thanks! [19:04:47] good luck and thanks for all the work, don't forget to rest tomorrow :) [19:04:57] I keep saying 5 mins away but I'm fairly confident this time :) [19:05:22] $deityspeed [19:07:23] FIRING: PuppetZeroResources: Puppet has failed generate resources on durum5003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetZeroResources [19:07:27] FIRING: [2x] PuppetZeroResources: Puppet has failed generate resources on hcaptcha-proxy5003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetZeroResources [19:08:32] false alarms, I'm assuming :) [19:11:10] brett: not sure [19:11:30] it's likely those VMs just came back online and talked to puppet for the first time in 10 hours [19:11:58] if you've a minute maybe take a look I think they should be ok, I'm checking some stuff on the dns hosts here right now [19:12:18] FIRING: [2x] PuppetZeroResources: Puppet has failed generate resources on hcaptcha-proxy5003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetZeroResources [19:15:52] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12258842 (10STei-WMF) @Ladsgroup @MatthewVernon thank you, I wrote it as: "The domain of URLs is changing from... [19:17:55] brett: so dns5004 seems to be ok, prefixes announced, it is answering queries [19:18:22] dns5003 is not announcing any prefix, the anycast-healthcheck file doesn't have them [19:19:00] perhaps it's not pooled or something? or should it be working and that's a sign of some problem network side perhaps? [19:19:23] both dns5003 and 5004 are depooled [19:22:20] Looks like slyngs depooled it earlier today [19:22:23] *them [19:23:20] ok.... so 5004 is announcing the IPs and answering some queries [19:23:35] https://www.irccloud.com/pastebin/nGwm8dZI/ [19:23:41] seems like he depooled in order to run authdns-update? Not sure why [19:24:19] yes exactly he did it so authdns-update could complete elsewhere [19:24:36] the same host is not answering some internal domains, so I'm somewhat confused [19:24:39] https://www.irccloud.com/pastebin/Phy38XXp/ [19:25:06] (as in I don't know 100% what expected behaviour is when depooled, or why dns5004 is announcing IPs and dns5003 isn't) [19:25:24] brett: can you take a look and repool if it makes sense/things look ok [19:26:04] topranks: I'll try. I'm gonna guess that if 5004 is working properly and 5003 isn't, that 5003 is broken and this isn't a function of depooling [19:27:26] brett: my bad I was checking a non-existant name. hence NXDOMAIN (Doh) [19:27:43] oh, haha, that'll do it! [19:27:49] so we are getting responses from dns5004. not sure what's up with 5003 let me know if you spot anything [19:28:49] brett: sorry I've been at this all day and getting a little confused [19:29:10] both of them are doing the same thing. neither are announcing IPs [19:29:29] I'm not familiar if that's because of the depool or could be a symptom of some other missed problem [19:31:16] I would think that pooling would enable the announcement [19:32:18] RESOLVED: PuppetZeroResources: Puppet has failed generate resources on durum5003:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetZeroResources [19:32:18] RESOLVED: PuppetZeroResources: Puppet has failed generate resources on hcaptcha-proxy5004:9100 - https://puppetboard.wikimedia.org/nodes?status=failed - https://grafana.wikimedia.org/d/yOxVDGvWk/puppet - https://alerts.wikimedia.org/?q=alertname%3DPuppetZeroResources [19:33:17] let's test by depooling another one and checking! [19:35:20] back [19:35:35] what's up with the DNS hosts? [19:35:35] ah, or sukhe could just tell us if depooling has that behavior :) [19:38:13] the eqsin dns hosts are both returning IPs for me? [19:40:59] so it seems safe to depool to me... I'm not sure what the issue is, exactly [19:41:06] *safe to pool [19:41:34] brett: we can come back to this I guess once we figure the liberica issue out [19:41:40] but yeah, pooling should fix it unless there is something else [19:42:23] okay, will pool both dns boxes then [19:42:40] let's wait I guess so that cathal can wait for us to verify changes on the switch [19:42:48] ack [19:44:29] for instance dns5003 does not announce any IP: [19:44:33] https://www.irccloud.com/pastebin/fBSPRqQk/ [19:44:49] topranks: yes, that's expected. brett, let's pool everything please [19:44:51] and cathal can check [19:45:46] er no [19:45:52] hehe, yeah, sorry [19:45:52] brett: no, not DNS pool [19:45:57] haha ok [19:46:04] pool dns500[34] [19:46:17] too many context switches going on [19:46:35] yep [19:48:31] topranks: can you see the IPs now from the dns hosts? [19:48:40] brett has pooled them [19:51:30] FIRING: [2x] LibericaStaleConfig: Liberica instance lvs5005 is running a stale configuration - https://wikitech.wikimedia.org/wiki/Liberica#LibericaStaleConfig - https://alerts.wikimedia.org/?q=alertname%3DLibericaStaleConfig [19:53:02] sukhe: yep looks a whole lot better thanks guys [19:53:04] https://phabricator.wikimedia.org/P96267 [19:53:47] although that hop 2 is giving me an old reverse dns entry [19:53:48] vrrp-gw-522.eqsin.wmnet [19:53:58] yep fixing that [19:54:00] I'm guessing that's just cos they are using the data from before [19:54:01] we need to run authdns-update [19:54:01] cool [19:54:06] running [19:54:12] fine with me I was like god damn not another one I forgot :D [19:54:14] brett: ^ as a step, I just ran it and realized I should have let you [19:54:26] brett: since it was depooled, we have the outdated DNS data there [19:54:58] ah right [19:56:30] RESOLVED: [2x] LibericaStaleConfig: Liberica instance lvs5005 is running a stale configuration - https://wikitech.wikimedia.org/wiki/Liberica#LibericaStaleConfig - https://alerts.wikimedia.org/?q=alertname%3DLibericaStaleConfig [20:27:10] FIRING: VarnishPrometheusExporterDown: Varnish Exporter on instance cp5022:9331 is unreachable - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/000000304/varnish-dc-stats?viewPanel=17 - https://alerts.wikimedia.org/?q=alertname%3DVarnishPrometheusExporterDown [20:27:13] FIRING: HaproxyKafkaExporterDown: HaproxyKafka on cp5022 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://grafana.wikimedia.org/d/d3e4e37c-c1d9-47af-9aad-a08dae2b3fd5/haproxykafka?orgId=1&var-site=eqsin&var-instance=cp5022 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [20:27:15] ok [20:27:36] cjd91: ^ please downtime cp5022 for a day [20:27:43] it's depooled [20:48:52] done [20:49:07] thanks! we will need update its IP in Puppet tomorrow and reimage before we repool [20:49:10] but that's for tomorrow [21:02:13] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqsin, and 2 others: EQSIN:New switch setup/configuration - https://phabricator.wikimedia.org/T418439#12259282 (10cmooney) 05Open→03Resolved Everything is now done here. I'm not sure if I totally followed the plan or not as it was a hectic en... [22:55:30] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12259808 (10Ladsgroup) I made a tiny change to make it clear url of thumbnails are changing. Not all urls. Al... [23:56:48] 06Traffic, 10Beta-Cluster-Infrastructure, 06SRE: Name new CDN servers deployment-cp-(text|upload)0x instead of deployment-cache-(text|upload)0x - https://phabricator.wikimedia.org/T280393#12259942 (10bd808)