[02:22:43] FIRING: HaproxyKafkaSocketDroppedMessages: Sustained high rate of dropped messages from HaproxyKafka - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaSocketDroppedMessages - https://grafana.wikimedia.org/d/d3e4e37c-c1d9-47af-9aad-a08dae2b3fd5/haproxykafka?orgId=1&var-site=esams&var-instance=cp3067&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaSocketDroppedMessages [02:37:43] RESOLVED: HaproxyKafkaSocketDroppedMessages: Sustained high rate of dropped messages from HaproxyKafka - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaSocketDroppedMessages - https://grafana.wikimedia.org/d/d3e4e37c-c1d9-47af-9aad-a08dae2b3fd5/haproxykafka?orgId=1&var-site=esams&var-instance=cp3067&viewPanel=panel-19 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaSocketDroppedMessages [06:22:44] 06Traffic, 06Commons, 10MediaWiki-File-management, 06SRE, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12278778 (10Krinkle) >>! In T435283#12248179, @BBlack wrote: > […] IMHO, the right fix for this is to fix... [06:34:04] 06Traffic, 06MediaWiki-Media-Platform-Team, 10Thumbor: Add `thumb_generated` from X-Analytics to webrequest_sampled_live - https://phabricator.wikimedia.org/T435634#12278793 (10ayounsi) Please sort it alphabetically :) {F101359298} [06:38:36] 06Traffic, 06Commons, 10MediaWiki-File-management, 06SRE, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12278798 (10Bawolff) The main thing i think is weird is the inconsistency of it. I think either choice is... [07:04:13] 10netops, 06Infrastructure-Foundations, 06SRE: Missing series for BGP session_state from eqiad CRs since upgrade to 23.4R2-S8.7 - https://phabricator.wikimedia.org/T435909#12278845 (10cmooney) >>! In T435909#12275570, @ayounsi wrote: > Adding a 3rd netflow host in eqiad shuffled the targets around, it caused... [07:47:14] 06Traffic, 06Data-Engineering, 06Data-Engineering-Radar: Normalize URI host in Turnilo's webrequest_sampled_live - https://phabricator.wikimedia.org/T434766#12278956 (10hashar) Awesome, thank you @Fabfur ! [08:00:40] FIRING: VarnishPrometheusExporterDown: Varnish Exporter on instance cp5022:9331 is unreachable - https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_job_unavailable - https://grafana.wikimedia.org/d/000000304/varnish-dc-stats?viewPanel=17 - https://alerts.wikimedia.org/?q=alertname%3DVarnishPrometheusExporterDown [08:00:43] FIRING: HaproxyKafkaExporterDown: HaproxyKafka on cp5022 is down - https://wikitech.wikimedia.org/wiki/HAProxyKafka#HaproxyKafkaExporterDown - https://grafana.wikimedia.org/d/d3e4e37c-c1d9-47af-9aad-a08dae2b3fd5/haproxykafka?orgId=1&var-site=eqsin&var-instance=cp5022 - https://alerts.wikimedia.org/?q=alertname%3DHaproxyKafkaExporterDown [08:08:44] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12279048 (10ops-monitoring-bot) Icinga downtime and Alertmanager silence (ID=3a675be0-bc5c-44f8-8996-c3f40f00c25d) set by fabfur@cumin1003 for 4:00:00 on 1 host(s) and their... [08:39:15] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12279180 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1003 for host netflow7002.magru.wmnet with OS trixie [11:42:51] 06Traffic, 06Commons, 10MediaWiki-File-management, 06SRE, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12280103 (10Verdy_p) Copy of a message I posted in Wikimedia Commons's Village Pump: == Severe bug in th... [12:02:17] 06Traffic, 06Commons, 10MediaWiki-File-management, 06SRE, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12280178 (10Verdy_p) Note that I also found another problem that this bug has caused, notably a few battl... [12:47:01] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12280451 (10ayounsi) [12:47:56] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12280452 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by ayounsi@cumin1004 for host netflow6001.drmrs.wmnet with OS trixie [13:17:53] 06Traffic, 06MediaWiki-Media-Platform-Team, 10Thumbor: Add `thumb_generated` from X-Analytics to webrequest_sampled_live - https://phabricator.wikimedia.org/T435634#12280592 (10Ladsgroup) I think this needs changing: https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1333087 I will do it soonTM [13:36:57] 10netops, 06Infrastructure-Foundations: Upgrade netflow hosts to Trixie - https://phabricator.wikimedia.org/T424478#12280737 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by ayounsi@cumin1004 for host netflow6001.drmrs.wmnet with OS trixie completed: - netflow6001 (**PASS**) - Downtimed... [14:17:52] 06Traffic, 06DC-Ops, 10ops-drmrs: hw troubleshooting: Memory failure for cp6008.drmrs.wmnet - https://phabricator.wikimedia.org/T431651#12281030 (10RobH) > Dear customer, > > We acknowledge receipt of your case and will proceed to complete your requested task within the timeframe you specified. We will upd... [14:31:05] Hi, I need a little bit of help. I merged https://gerrit.wikimedia.org/r/1333820 an hour ago, and https://airflow-experiment-platform.wikimedia.org/ is responding everywhere except `text-lb.esams.wikimedia.org` [14:33:37] nvm, it has started working (I was getting a placeholder there, but it now returns the correct thing) [14:33:42] atsukoito: curl -s -I --connect-to airflow-experiment-platform.wikimedia.org:443:$(dig +short text-lb.esams.wikimedia.org) https://airflow-experiment-platform.wikimedia.org/ | grep -i x-cache [14:33:46] x-cache: cp3069 miss, cp3069 pass [14:33:47] ah ok! [14:35:13] sukhe: it was `x-cache: cp3069 hit, cp3069 hit/51`, and I wondered when would it expire [14:36:17] atsukoito: it could be that Puppet had not run on the cp host you were reaching, so that's one possibility [14:36:22] but yeah, let us know if it happens again [14:37:10] sukhe: I did `sudo cumin 'A:cp-text_esams' run-puppet-agent`, but I tried accessing the URL before running puppet, so maybe I poisoned the cache [14:37:35] anyways, no worries and thanks! [14:37:57] :)) [14:47:28] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12281197 (10Nux) Hi Lads. and lads ;-) ### Direct link to commons > Also, Since we have deployed to dewiki. O... [15:04:59] jelto: https://gerrit.wikimedia.org/r/plugins/gitiles/operations/puppet/+/a1494c9505578d2e86fded8070f15f256ea58b08 [15:05:03] this requires a pybal restart [15:05:06] want us to take care of it? [15:05:14] not that the changes won't be live until then [15:10:58] The ipip-migration cookbook does a pybal restart afaik. There should be a restart in SAL. [15:11:43] that is correct yep, if you did that [15:11:49] let me force a recheck [15:13:56] Yes I did use the cookbook. Thank you! also have a change for eqiad https://gerrit.wikimedia.org/r/c/operations/puppet/+/1333691 but I'll deploy that tomorrow earliest [15:15:02] jelto: so it seems like for some reason, restart-pybal never kicked in [15:15:15] even though it should [15:15:42] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814 (10cmooney) 03NEW p:05Triage→03Medium [15:16:04] so doing that manually and then will check the cookbook logs later on why it didn't kick in [15:16:19] return self.spicerack.run_cookbook("sre.loadbalancer.restart-pybal", args) [15:16:42] Oh okay, thank you. I'm already away from my computer and can't follow you properly unfortunately. [15:17:02] don't worry at all, we will take care of it [15:17:06] just wanted to let you know it is not live [15:17:10] rest up! [15:33:12] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12281583 (10Ladsgroup) >>! In T427465#12281197, @Nux wrote: > Hi Lads. and lads ;-) > > ### Direct link to com... [15:36:16] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814#12281602 (10Jclark-ctr) @cmooney can these be pulled at anytime or do you want to be online when pulled? [16:37:28] 06Traffic, 06Data-Engineering: Add backend information to webrequests in data lake - https://phabricator.wikimedia.org/T436223#12281866 (10Ahoelzl) a:03GGoncalves-WMF [16:37:47] 06Traffic, 06Data-Engineering: Add backend information to webrequests in data lake - https://phabricator.wikimedia.org/T436223#12281869 (10Ahoelzl) @GGoncalves-WMF can you triage please? [16:51:46] 10netops, 06DC-Ops, 06Infrastructure-Foundations, 10ops-eqiad, 06SRE: Power alert for cr1-eqiad old line cards - https://phabricator.wikimedia.org/T436814#12282002 (10cmooney) >>! In T436814#12281602, @Jclark-ctr wrote: > @cmooney can these be pulled at anytime or do you want to be online when pulled? T... [16:55:32] sukhe: thanks for https://gerrit.wikimedia.org/r/c/operations/puppet/+/1321987, that made the migraiton super smooth on our end [16:55:47] well s/made/will make/, as I'll deploy tomorrow [16:58:10] brouberol: ah, you are migrating stuff to the new urldownloader service, nice! [17:13:06] yep, it's a single URL change in airflow-dags, the rest is auto-generated changes https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/2603 [17:13:53] I am just amazed at the stuff that was depending on urldownloaders, mostly by the change moritz.m has been making since we moved it to behind LVS [17:14:15] changes rather [20:14:12] 06Traffic, 06Commons, 10MediaWiki-File-management, 06SRE, and 2 others: Varnish serving outdated version at original/non-thumb URL of overwritten file upload - https://phabricator.wikimedia.org/T425216#12282999 (10Nux) Interestingly thumbs work. So when the image you upload is big (like 400x400) the thumb... [20:50:38] 06Traffic, 10Citoid, 06MediaWiki-API-Platform-Team, 06ServiceOps, and 3 others: Consider adding a php interface over the citoid api in the Citoid extension - https://phabricator.wikimedia.org/T435916#12283078 (10Scott_French) [20:53:07] Fwiw I’m pushing DE to ensure traffic stays internal when possible. We’ve migrated all internal requests to gitlab.w.o to gitlab.discovery.wmnet:8443 today. I’ll review how we’re connecting to the mw API next, and ensure we go though the mesh [20:56:01] (From our DAGs)