[07:43:20] jelto: sorry, didn't mean to fight you over the commit message [07:44:28] np, I just fixed CI, you improved the actual message :D [10:45:37] arnoldokoth: ok to merge your puppet changes? [10:45:47] (gitlab version bump) [10:46:09] taavi: yeah. thank you. [10:46:22] done [10:47:01] ty [14:31:59] Hi, sorry I need a bit history lesson. I thought we used to be on Ceph then switched to Swift, but I can't find it. there are old docs on Ceph but it's just "evaluate": https://wikitech.wikimedia.org/w/index.php?title=Ceph&oldid=74965 My questions are: What it was before swift (NFS?). And if Ceph wasn't picked, is there any detailed reason why? https://wikitech.wikimedia.org/wiki/Obsolete:Media_server/Distributed_File_Storage_choices says [14:31:59] "Eliminated because it is not production code yet (although it looks good). [14:32:36] <_joe_> Amir1: you are talking more than 15 years ago [14:33:11] yeah, I just want to make sure everything is covered [14:33:12] <_joe_> and yes, I think it was NFS before, but I think you might need to ask Mark [14:39:23] all of Commons was served from an Iomega Zipdrive attached to one of the servers [14:45:53] https://diff.wikimedia.org/2012/02/09/scaling-media-storage-at-wikimedia-with-swift/ [14:46:04] > We’ve been able to manage the load for quite a while by using two servers with lots of local storage — (10 and 30TB), [14:46:37] close enough? :P [14:53:49] <_joe_> we definitely used a couple of NetApps for quite a few things IIRC [14:53:56] <_joe_> I thought commons as well [15:37:36] Swift came first, from Labs. this was late 2011-early 2012. we used a very early code of Swift (and actually had support from the authors at the time, SwiftStack) [15:38:07] it was deployed on some really shitty hardware from Dell, the C2100 (C = Cloud) [15:38:15] you would reboot them, and they would lose a bunch of disks [15:38:37] we ended up escalating it to the highest levels of Dell US, and they ended up replacing all the servers for free [15:39:50] but Swift at the time would fire-and-forget rsync, with no periodic auditing, and no way to view how many replicas of an object you have etc. [15:40:33] so these two combined meant that there was a very real risk of losing data, just by rebooting a server [15:43:32] plus Swift was incomplete and had a bunch of bugs (we were early adopters) like code paths that were XXX TODO and poorly thought code paths that caused a few outages (I remember at least one) [15:44:36] finally, there was a change in staff - the person that originally deployed Swift (Ben Hatrshone) left in Aug 2012 - I joined in April 2012, and while I was not replacing Ben, I was the person with the most spare bandwidth and so I picked it up [15:45:43] so I started an... exploration to replace Swift with Ceph. I worked on it for months, Mark helped a lot too. Upstream was immensely helpful, I was talking with Sage (Ceph author) on a daily basis, I think I filed > 50 bugs [15:45:54] and we were in touch with Inktank (eventually acquired by RedHat) as well [15:47:12] but sadly, Ceph was not ready for the use case - all the rage at the time was virtualization and OpenStack, and their focus at the time was RADOS-backed VMs, which was the polar opposite of our use case [15:47:37] we wanted to store billions of tiny thumbnails, while their other users & customers were focusing on storing VM data disks [15:48:02] so sadly, we cut our losses short (way later than I should have, in retrospect - sunk cost fallacy is a bitch) and stayed on Swift [15:49:06] and yes pre-Swift every mw* host would NFS-mount. if you've used NFS, you can imagine the horrors. if not, it's good to be you [15:49:58] but thumbnails were special - there was also a server, in amsterdam(!) called ms6 that was basically an nginx with btrfs(!) that lived A LOT longer than anyone could have ever imagined [15:50:40] Amir1: I think that's all I remember, but if you have more questions I can think about it some more [15:53:33] oh yeah, and Labs was also on a (two-node!) GlusterFS, but that's a story for another day :) [16:00:25] <_joe_> this was horror tales from the past, with paravoid, presented to you by State Farm [16:00:53] <_joe_> (I was resisting pinging you about all this, I remembered you was the one who evaluated ceph :)) [16:08:03] Thank you so much paravoid <3 [17:37:03] !bash paravoid> yes pre-Swift every mw* host would NFS-mount. if you've used NFS, you can imagine the horrors. if not, it's good to be you [17:37:03] Amir1: Stored quip at https://bash.toolforge.org/quip/x74lP6ABDAZZyZXnl63O [17:43:08] we regret to inform you I have indeed used NFS [17:44:16] I've seen horrors second hand mostly :D [18:00:08] I only used NFS in college, and only a few times, socially [18:33:18] haha [18:56:29] I used NFS a ton at various $work in the mid-late 90s and on through... I donno, I might have last really had to deal with it around 2009 or so? [18:58:33] it's the ultimate Leaky Abstraction IMHO (actually, the original even mentions it as an example. I think TCP was the prime example because more people understand that one, but NFS is arguably a much stronger example) [18:59:14] shameless plug for old good stuff: https://www.joelonsoftware.com/2002/11/11/the-law-of-leaky-abstractions/ [19:09:40] topranks: are you still hands-on in eqsin? [19:09:57] or maybe sukhe knows [19:10:42] nothing on https://phabricator.wikimedia.org/T435406 since 16:25 [19:10:58] silence was last extended 4h at 16:23 [19:11:08] ah :) [19:11:12] but verified still depooled so the alerts are-- yeah [19:11:20] hey [19:11:47] but the math doesn't quite math ... like, it's 19:11 now [19:11:51] well, it's not 4h since 16:23 yet, so-- lol [19:11:59] topranks is still working on eqsin. send him some wikilove, he had to work through many small things today [19:12:23] yeah the hosts were downtimed but we may need to extend something extra for the CDN alerts [19:12:36] sorry yeah I probably underestimated the time it'd take on the 4th extension like on the first 3 :P [19:12:53] oh 100%, no shade to topranks here, thanks for the work <3 just making sure we didn't need to do anything for the alerts [19:13:08] sukhe: I propose a small prometheus exporter running on dns boxen that exports the admin_state per site [19:13:11] and it sounds like, just silence [19:13:11] fwiw almost at a situation where things should be ok [19:13:14] rzl: want me to throw something in alertmanager? [19:13:19] then we wouldn't have to do manual silences [19:13:29] the probedown alerts for a PoP site could depend on that [19:13:30] did something page? [19:13:35] swfrench-wmf: nah I got it, only needs one of us [19:13:40] if something is broken I best know [19:13:46] * swfrench-wmf thumbs up [19:13:58] VMs just came back online (I've not checked them all) [19:14:00] topranks: silence expired on the ProbeDown alerts for eqsin [19:14:00] topranks: yeah we got "ProbeDown: Service text:80 has failed probes" [19:14:20] (and also Service the-rest-of-em) [19:14:23] perhaps prometheus is firing alerts it's been trying to send all day and couldn't [19:15:19] topranks: I'm gonna put it in for 24h just to have plenty of wiggle room, we can always remove it when you're done [19:16:41] HEY WE HAVE THAT ALREADY https://grafana-rw.wikimedia.org/goto/ssp8pt?orgId=default [19:17:12] what are we silencing for 24h? [19:17:30] I'd just be concerned we repool the site and there are issues which we don't know about due to some silence [19:18:38] ... no, that's something different [19:18:42] https://alerts.wikimedia.org/#/silences/199f0867-e123-41aa-adb9-f0f0f243364b - I silenced alertname=" ProbeDown"site="eqsin" severity="page" [19:19:03] topranks: the cache nodes there can't reach the backend services in the core DC [19:19:12] but happy to adjust, and agree we should check before repooling anyway to make sure nothing is still firing [19:19:20] the cp nodes? [19:19:43] (at that alerts.wm.o link you can see the 8 alerts currently matching, should be 0 before we repool) [19:20:36] I'm going to resolve in VO too since it won't hear about the silence [19:20:42] !resolve [19:20:42] 8309 (RESOLVED) [3x] ProbeDown sre (probes/service eqsin) [19:21:19] rzl: thanks, I'll silence them all for a few hours, don't want to do 24h and hide a problem [19:25:19] topranks: sure, feel free to edit it [19:26:05] rzl: no that's fine, I guess I need I can delete it after? I need to work out how to do that [19:26:51] argh I guess that deep link doesn't work? I'm trying to find another URL that does [19:27:04] but when you find the silence I promise it has a delete button 🙃 [19:27:23] don't let me distract you from the real stuff you're doing, if it's easier feel free to just ping me to kill it when the time comes [19:29:05] does this work? https://alerts.wikimedia.org/?q=%40silenced_by%3D199f0867-e123-41aa-adb9-f0f0f243364b [19:29:29] (the problem is, when it's time to remove the silence, that search won't match any alerts...) [19:29:31] rzl: i too have had this exact frustration [19:29:41] you have to open the new silence window and "browse" [19:29:49] there is a search box inside there at least [19:30:19] ah yeah -- I put T435406 in the description so you can search for that [19:30:19] T435406: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406 [19:30:27] rzl: ok thanks! [19:30:48] I might ping you when the time comes if that's ok, I should learn this in more detail but now may not be the time [19:30:55] absolutely please do [19:31:03] cdanis: when you said "the cache nodes there can't reach the backend services in the core DC" what was that based on? [19:31:36] on ProbeDown firing, I hadn't actually tested anything [19:31:40] I'm sure it's right just not picked the correct broken endpoint yet I think [19:31:41] ok [19:31:54] I'll keep with the checks then [19:32:34] could you use a hand? [19:34:10] cdanis: yeah I'm not 100% sure how to interpret the failed probe alerts [19:34:20] I guess those are blackbox probes being run on the prometheus VMs and failing? [19:34:21] topranks: I don't know where the exporter actually runs [19:34:28] yeah, I think probably on the prometheus VMs [19:35:47] (back) [19:35:57] rzl: would you be interested in reviewing https://gerrit.wikimedia.org/r/1329644 ? [19:36:00] cdanis: catching up but it can also be that liberica has not pooled the hosts or something? [19:36:03] checking that [19:36:10] cdanis: looking [19:36:40] sukhe: I was just getting to that point of thinking to look at the LVS [19:36:49] yeah liberica has some old IPs [19:36:55] we have do a restart [19:37:00] I probably missed a similar "change the BGP peers" puppet patch for those too [19:37:04] that too yep [19:37:10] so that's two things we need to verify [19:37:12] on it [19:37:54] I grepped the v6 IP earlier in the repo, forgot pybal group is v4-only [19:38:18] sukhe: I just delete profile::liberica::bgp_config: for eqsin/ [19:38:47] yep [19:38:51] that should be it [19:39:11] sukhe@lvs5004:~$ gobgp neighbor [19:39:11] Peer AS Up/Down State |#Received Accepted [19:39:11] 103.102.166.130 14907 03:35:39 Establ | 0 0 [19:39:11] 103.102.166.131 14907 03:35:32 Establ | 0 0 [19:39:33] that's cr[23]-eqsin [19:40:16] https://gerrit.wikimedia.org/r/c/operations/puppet/+/1329647 [19:40:44] yep +1 [19:40:51] I will take care of liberica restarts for you [19:40:56] you can focus on the other stuff [19:41:02] ugh that is a nasty setup we have [19:41:10] what's missing now? [19:41:26] so Liberica can establish BGP? but the routes aren't accepted cos the next-hop is invalid if there are multiple IP hops between [19:41:35] the multihop should be disabled there, was never needed [19:41:36] sigh [19:41:48] LVS should be peering with the switch [19:42:12] sukhe: should be ok when we merge the patch [19:42:35] I'm just commenting on the above where it "establishes" due to "multihop" being allowed, despite the fact it can't work in that scenario [19:43:40] I see ok, so we can come back to this then, I made a note [19:43:56] ok to merge? [19:46:13] topranks: I am merging and testing this [19:48:03] yes please do [19:48:06] thanks [19:49:51] restarting liberica [19:50:01] cdanis: oh thanks for that pointer in the commit message, I was trying to check the template and realizing that my etcdctl is super rusty (to address some other time) [19:50:07] rzl: same [19:50:09] that's why I added it [19:50:21] and you also have to remember to ask for the v2 data also with an env var [19:50:49] yeah I knew I was at least partly eating shit due to v2/v3 [19:51:56] topranks: [19:51:57] sukhe@lvs5004:~$ gobgp neighbor [19:51:57] Peer AS Up/Down State |#Received Accepted [19:51:57] 10.132.0.1 4265005001 00:00:49 Establ | 0 0 [19:54:26] ok [19:55:07] nice yeah it looks up on the switch too [19:55:10] lvs5006 still down [19:55:56] looking a whole lot better [19:56:00] cathal@officepc:/media/scratch2/router_images/nokia/pm2/splits/ai/civit/files$ ping upload-lb.eqsin.wikimedia.org [19:56:00] PING upload-lb.eqsin.wikimedia.org (2001:df2:e500:ed1a::2:b) 56 data bytes [19:56:00] 64 bytes from upload-lb.eqsin.wikimedia.org (2001:df2:e500:ed1a::2:b): icmp_seq=1 ttl=55 time=160 ms [19:56:00] 64 bytes from upload-lb.eqsin.wikimedia.org (2001:df2:e500:ed1a::2:b): icmp_seq=2 ttl=55 time=160 ms [19:56:52] lvs5006 now ok [19:56:55] yes [19:57:09] one more thing, I have 10.132.0.30 in textlb_80 [19:57:16] which I don't know what it is and it should not be there [19:57:33] wmf_geodns_service_pooled{service="text-addrs", datacenter="eqiad"} 1 [19:57:34] wmf_geodns_service_pooled{service="text-addrs", datacenter="eqsin"} 0 [19:58:01] cdanis: very nice <3 [19:58:07] topranks: ok I think I know [19:58:08] cp5022 [19:58:17] I am going to just depool it for now [19:58:19] ah ok [19:58:33] yeah that host is offline, Simon said earlier there was some problem with it? [19:58:44] it's in rack 604 but has a rack 603 IP [19:58:47] yeah we need to move-vlan it [19:58:49] the cookbook [19:58:59] I did it all manually earlier, just need to do regular reimage [19:59:07] ok great thanks [19:59:09] so all resolves are in [19:59:14] brett: are we good on the DNS hosts please? [19:59:15] Simon thought there was some hw fault or something why it didn't happen previous [19:59:15] both pooled? [19:59:27] yea it was a faulty CPU but now fixed [19:59:33] anyway I have left it depooled for now [20:00:03] ah ok, but yeah hence why it wasn't moved previous [20:00:08] regular reimage will fix it now I think [20:00:16] sukhe: Uh, I think so? [20:00:16] topranks: things look OK, I think we should pool and let some traffic flow for a live test. [20:00:33] ProbeDown alerts are all resolved, silence not currently matching anything -- should I delete it? [20:00:34] I'll check getsel on the hosts [20:00:46] brett: cp5022 is depooled so we can do it later [20:00:57] rzl: thanks for confirming, I think yes, if anything fires now it will be unexpected and a real issue [20:01:07] I hope not of course but things look healthy [20:01:15] oh, wait, I thought you were saying the dns hosts were suspected to have hw faults [20:01:21] gotcha [20:01:21] silence deleted [20:03:16] topranks: ok to pool? [20:03:16] ty <3 [20:03:44] sukhe: give me a few mins for a final check [20:03:53] no worries [20:08:31] sukhe: ok I think we're good [20:09:11] brett: want to push the buttons :) [20:09:42] sukhe: dns pool of eqsin? [20:09:46] yep [20:09:50] task ID is T435406 [20:09:51] T435406: eqsin switch migration - 2026-08-26 @ 08:00 UTC - https://phabricator.wikimedia.org/T435406 [20:10:05] aye [20:10:37] heh in fact we've some traffic already as soon as the LVS came up, from the good folk who don't respect DNS TTLs [20:10:45] wmf_geodns_service_pooled{service="upload-addrs", datacenter="eqsin"} 1 [20:11:24] now to figure out how to actually glue the pieces together [20:12:12] topranks: brett: looks good [20:12:12] < x-cache: cp5025 hit, cp5025 hit/625 [20:12:26] cdanis: I would love to get a table of currently-depooled sites on https://grafana.wikimedia.org/d/O_OXJyTVk/home-w-wiki-status too [20:12:35] I know it's always tempting to add everything there but this seems like an actually good candidate [20:12:36] topranks: nice job! <3 [20:12:54] not a table of all sites and their pooled status, just a list that's empty most of the time [20:13:47] and, good stuff everybody and especially topranks [20:14:14] https://grafana.wikimedia.org/goto/smv95d?orgId=default [20:14:22] dns5003 and 4 also picking up traffic nicely [20:15:11] so we decided some other improvements to the DNS cookbook today as well, in light of this work. (not the things cdanis is doing, we did not think of those) [20:15:19] response status NOERROR is still so funny [20:15:20] but for example depooling DNS hosts as part of the cookbook [20:15:29] suggestions welcome on other stuff [20:15:47] basically reduce more toil on what we can do through the cookbook instead of doing manually [20:16:07] yeah, I would have said silence the alerts but cdanis's thing is even better [20:16:22] and, to topranks's point too, I can see the argument for not always coupling them up anyway [20:17:11] (which means there's some room for discussion about exactly how to incorporate the pooled status in the alert -- maybe we still get a ProbeDown but it doesn't page?) [20:17:25] rzl: yeah I think I can make it a warning but not critical [20:17:27] rzl: I guess that point was more because we were right at the end of the maintenance, I was expecting things to be ok at that point [20:17:42] so in a way by then a short ACK and then it resolving it was I was thinking was best [20:17:54] yeah, that makes sense too [20:18:02] oh totally [20:18:31] sorry for stepping on your toes there <3 [20:24:26] nah not at all thanks for the help, sorry for the page :) [20:24:46] sukhe: looks ok to me?? traffic levels ramped up a decent amount [20:25:01] topranks: yep, everything is looking good and stable IMO [20:25:08] rest up topranks! [20:25:18] don't fix stuff that doesn't need to be fixed today please [20:26:38] we have overextended too many people this week already [20:27:52] yep! [20:44:56] gotta log off now but I posted an attempt at the ProbeDown thing [20:47:30] krinkle@stat1011.eqiad.wmnet$ kafkacat .. [20:47:30] bash: kafkacat: command not found [20:47:30] https://sal.toolforge.org/production?p=0&q=stat1011&d= [20:47:30] I guess it got lost in the reimage? [20:48:30] https://gerrit.wikimedia.org/g/operations/puppet/+/7b301c6a275c38149aafc5ac3b78fb2f3f339b91/modules/profile/manifests/analytics/cluster/client.pp#38 [20:48:32] "kcat"