[00:05:25] FIRING: [5x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1002:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [00:10:25] FIRING: [6x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [04:10:25] FIRING: [6x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [08:10:40] FIRING: [6x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [10:10:58] 06Traffic, 06Data-Engineering, 06Data-Engineering-Radar, 13Patch-For-Review: Normalize URI host in Turnilo's webrequest_sampled_live - https://phabricator.wikimedia.org/T434766#12270301 (10Fabfur) We've identified the issue (the Host header was captured and logged before the normalization directive). We've... [10:11:35] This will be actually quite impactful if anyone has a moment: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1332678 [10:14:48] basically half of our thumbs are being served via webp which is quite better in terms of compression (in svg, it's half, in jpg it's around 30%). But we basically didn't advertise this to safari leading to 30% of our traffic still using old larger files. See https://w.wiki/T$PZ [10:15:03] (basically all of svg should been served from webp) [10:17:31] do we have a way to test it afterwards? [10:27:13] with browser stack https://www.browserstack.com/ we could? [10:28:29] I definitely check the graphs [11:47:15] Don't we just need someone with Safari to test? E.g. me? [11:48:23] I don't have Safari 14 and 15, but newer isn't an issue [12:10:40] FIRING: [6x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [12:34:33] Based on dashiki the traffic from safari is like 1% [12:34:40] At most [12:35:57] I want to eventually remove that piece altogether but I also want to show it's negligible traffic. We have switched to webp unconditionally in several places and Noone has complained [12:41:10] (Safari 14 and 15 I mean. Safari overall is around 30%) [13:14:26] yup, based on this: Safari 14 and 15 are 0.2% https://analytics.wikimedia.org/dashboards/browsers/#all-sites-by-browser/browser-family-and-major-hierarchical-view 0.3% if you merge everything below 15 [13:14:34] 15 and below [13:14:34] 10netops, 06Infrastructure-Foundations: Create a 3rd Netflow instance in eqiad/codfw - https://phabricator.wikimedia.org/T436517 (10ayounsi) 03NEW p:05Triage→03High [13:23:02] 06Traffic, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th), 13Patch-For-Review: Turnilo: X-Is-Browser filter doesn't work (always HTTP 500) - https://phabricator.wikimedia.org/T435152#12270798 (10JAllemandou) Hi @CDanis and @Krinkle , I have added a new field `X-Is-Browser score (STR)` to make filt... [14:15:34] 06Traffic, 06Data-Engineering (Q1 FS26/27 July 1st - September 30th): Turnilo: X-Is-Browser filter doesn't work (always HTTP 500) - https://phabricator.wikimedia.org/T435152#12271154 (10CDanis) Looks good to me, thanks Joseph! [14:45:49] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12271335 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage was started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie [15:34:52] 06Traffic, 06Collaboration-Services, 06SRE Observability, 10Wikimedia-Mailing-lists, 07Sustainability (Incident Followup): Consider disabling paging for lists.wikimedia.org - https://phabricator.wikimedia.org/T436049#12271578 (10LSobanski) 05Open→03Stalled Stalling until the ratios change is implemen... [15:39:30] 06Traffic, 06Collaboration-Services, 06SRE Observability, 10Wikimedia-Mailing-lists, 07Sustainability (Incident Followup): Consider disabling paging for lists.wikimedia.org - https://phabricator.wikimedia.org/T436049#12271613 (10LSobanski) p:05Triage→03High [15:52:26] Amir1: Any chance the safari exclusions have to do with bugginess? I know safari is kinda the internet explorer of this age... [15:52:38] I figure not, just crossed my mind [15:53:07] yeah, it was only for versions of 14 and 15 back then but never got fixed. Let me grab you details [15:53:39] brett: see Timo's comment in https://gerrit.wikimedia.org/r/c/operations/puppet/+/1306395/3#message-8b8c1cb606aa97cb63e1ad510e2d4ed6ffb640c0 [15:54:24] basically, it was due to a specific bug but that bug is resolved since version 16 onwards :D [15:54:31] gotcha! [15:55:55] 10netops, 06Infrastructure-Foundations: Consider storing the entire chain in network device's gRPC certificate - https://phabricator.wikimedia.org/T375513#12271748 (10elukey) @ayounsi IIUC in this case the gRPC certificate is not validated by prometheus and other tools? If so it should be sufficient to expose... [15:56:09] unrelated thing: benthos is on cache hosts but this was noop in any host I tried (both PCC and puppet agent after merge) https://gerrit.wikimedia.org/r/c/operations/puppet/+/1328212 still worth mentioning in case thing go sideways [15:58:09] ah, a grep showed where it is hieradata/role/common/syslog/centralserver.yaml [16:02:49] Thanks for the patch, Amir1! Would you like me to roll that out or do you have it? [16:02:49] brett: thank you for the review <3 if you push it, I'd be grateful. If you have time [16:02:49] I look at graphs and cheer [16:02:49] will do! [16:03:05] thanks@ [16:07:09] 06Traffic, 06DC-Ops, 10ops-eqsin, 06SRE, 13Patch-For-Review: cp5022 is unreachable - https://phabricator.wikimedia.org/T414411#12271795 (10ops-monitoring-bot) Cookbook cookbooks.sre.hosts.reimage started by cdobbins@cumin1003 for host cp5022.eqsin.wmnet with OS trixie executed with errors: - cp5022 (... [16:10:40] FIRING: [6x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [16:14:44] 06Traffic, 10Pywikibot, 06SRE, 10Wikidata, and 3 others: Pywikibot reports maxlag retry error on Wikidata - https://phabricator.wikimedia.org/T421642#12271821 (10JeanFred) As I wrote back in T244030 (which was {T242081} back then): >>! In T244030#5842377, @JeanFred wrote: > I was under the impression that... [16:24:26] Amir1: rolled out [16:24:40] \o/ [16:25:55] https://w.wiki/T$t9 [16:27:03] sexy [16:28:37] xD [18:05:30] 06Traffic, 06Data-Engineering, 06Data-Engineering-Radar: Normalize URI host in Turnilo's webrequest_sampled_live - https://phabricator.wikimedia.org/T434766#12272228 (10Fabfur) 05Open→03Resolved a:03Fabfur as fare as I can see on turnilo, after the change above I can't find occurrences of "funny" `... [19:31:12] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12272499 (10Soda) Why has this change been implemented before giving enough time to the community to fix critic... [19:36:40] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12272538 (10Soda) Wikisource's fairly critical OCR tool appears to be broken per T436551 as a result of this. [19:44:48] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12272578 (10Ladsgroup) We have done gradual roll out and given that the old urls still work, the vast majority... [20:07:58] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12272692 (10Ssein) Looking at this T427465, it seems the cause is the thumbnail domain migration from upload.wi... [20:10:40] FIRING: [6x] SystemdUnitFailed: dump_cloud_ip_ranges.service on puppetserver1001:9100 - https://wikitech.wikimedia.org/wiki/Monitoring/check_systemd_state - https://grafana.wikimedia.org/d/g-AaZRFWk/systemd-status - https://alerts.wikimedia.org/?q=alertname%3DSystemdUnitFailed [20:11:52] 06Traffic, 06SRE: Prometheus Ferm MSS export does not work for IPv6 - https://phabricator.wikimedia.org/T433672#12272699 (10CDobbins) I wasn't able to reproduce this on clouddumps1002: ` cdobbins@clouddumps1002:~$ sudo /usr/local/bin/prometheus-ferm-mss -o /tmp/test_run.prom -e 208.80.154.242:2049 -e 208.80.1... [20:12:46] 06Traffic, 06Data-Persistence, 06MediaWiki-Media-Platform-Team, 13Patch-For-Review, 07User-notice: Move thumbnail caching from upload cluster to text - https://phabricator.wikimedia.org/T427465#12272700 (10Soda) >>! In T427465#12272578, @Ladsgroup wrote: > Is there anything beside OCR tool broken? I'll...