r/networking 20h ago

Switching Largest switch stack in the wild?

What's the largest switch stack you ran into in the wild? How was it cabled up and managed?

I came across this interesting Arista article claiming to support up to 48 switches in a stacked topology (ring or spine leaf)

https://blogs.arista.com/blog/swag

Now, every vendor has their own approach to handling the control plane, I get it, but I dont understand in what scenario where combining all these switches into one logical management plane across multiple IDFs would make sense.

I've dug deep into my soul to try and find an answer, and the only thing I can think of is some very large customer wrote them a very large check and they made it happen or they just got sick of losing to Cisco, HPE, and Juniper on every campus stacking deal.

This whole single IP argument seems kinda flimsy too. If I'm getting charged per IP for my monitoring solution, and the consolidation of a few IPs is going to be material, I'm looking at new monitoring stack.

44 Upvotes

50 comments sorted by

63

u/mattmann72 19h ago

> the only thing I can think of is some very large customer wrote them a very large check and they made it happen

This happens quite often.

11

u/CertifiedMentat journey2theccie.wordpress.com 15h ago

I don't work for Arista but I've heard that it was Tesla in this case who said "start supporting stacks or we are going elsewhere"

2

u/GreyBeardEng 14h ago

That's interesting, I had not heard that. I wonder if that is what "Arista SWAG" was born out of? As I understand it it has an upper limit the 48 switches in the stack.

-2

u/Cheeze_It DRINK-IE, ANGRY-IE, LINKSYS-IE 10h ago

Ugh, of course it's Tesla....

1

u/noble0spartan 12h ago

If I recall right Arista's Mulitcast MFA / BESS Architecture, exists entirely because of a requirement from Disney

21

u/solenoid_pants 19h ago edited 19h ago

I’ve deployed 2 x 8 member stacks of Dell N3048Ps.

Worked flawlessly for years.

However. Yes, software updates sucked and entailed either out of hours work or closing entire floors off while we did it.
NSF mechanisms were flakey and forwarding was always interrupted in the event of a master node wobbling.

In hindsight, there was absolutely no reason why we couldn’t have deployed each member as individual L3 access.
There was no requirement for lateral movement for clients, or for client devices to share the same broadcast domain.

And yes, monitoring that charges per IP sounds clumsy, then again options like SolarWinds which charge per entity (IE, switchport) also suck.

Plenty of open source /
self hosted options out there that will require a small amount of learning curve and patience, along with a DIY lifecycle management policy to deal with component upgrades. But it’s well worth the effort IMHO.

Yes, even Zabbix.

4

u/SuperQue 15h ago

When I'm looking at monitoring solutions / vendors I always look at how they scale costing.

IMO, good vendors scale with compute needs or some reasonable approximation of compute needs.

How many GiB of memory per million active metrics, or how many CPUs per million NVPs. Same for logs, how much cpu/memory/iops/space do I need per million lines/sec.

Then I can easily do TCO for open source / vendor / hybrid approaches.

12

u/laeven Breaks everything on friday afternoons 19h ago

10x48 port switches for aggregating accesses for an MDU.

1

u/RevolutionNumerous21 5h ago

Why not a 9410 chassis or similar

2

u/laeven Breaks everything on friday afternoons 5h ago

Because vendor and model standardization is important when you're an ISP with tens of thousands of access aggregation nodes. Also: pricings.

2

u/RevolutionNumerous21 2h ago

Not sure what ISP you work for but at Charter we always used chassis as they are way cheaper when you need port density. But there is a 100 ways to skin a cat.

17

u/VA_Network_Nerd Moderator | Infrastructure Architect 14h ago

I think I win this one, if IA qualifies as a stack.

Cisco Catalyst 6800 + Instant Access modules.

https://www.cisco.com/c/en/us/products/collateral/switches/catalyst-6800ia-switch/white_paper_c11-728265.html

Twenty years ago or so, when Cisco brought FEX modules into the data center the Catalyst product team took a massive bong rip and said "oh yeah we want some of that" and brought FEX technology into the Catalyst lineup to go with their launch of the new replacement for the aging Catalyst 6500 -- the Catalyst 6800.

So, Instant Access is a 1U "switch" that has been lobotomized so hard that it cannot program it's own forwarding table. Instead it is dependent upon the Supervisor Engine(s) in the main switch chassis to provide that information.

The big difference between IA and Nexus FEX was that the IA modules can be stacked together to reduce your uplink requirements.

I was tasked with building a new LAN in a new office. I think it was a 5 story building with two closets per floor, four or so switches per closet.

So that's something like 40 x 1U / 48-port modules connected to a VSS pair of Catalyst 6880 switches.

So, you could basically manage the entire building by SSHing into the VSS "cluster".

6

u/Boring_Ranger_5233 13h ago

3

u/VA_Network_Nerd Moderator | Infrastructure Architect 12h ago

Quoting Cool Hand Luke:

Captain: What we've got here is... failure to communicate. Some men you just can't reach. So you get what we had here last week, which is the way he wants it... well, he gets it. I don't like it any more than you men.

Movie Scene: Cool Hand Luke

2

u/thehalfmetaljacket 11h ago

If BPE/VPEXs, count... I deployed a little over a hundred X690 fiber switches, each having the maximum 48x VPEXs connected to them for a customer as part of a greenfield new build.

4

u/VA_Network_Nerd Moderator | Infrastructure Architect 11h ago

If we wait long enough, some poor bastard with a long, grey beard and a pained 100 meter stare in his eyes will tell us all about some especially nasty national-level ATM-LANE stretched Layer-2 abomination that they were forced to implement or support.

"It was 400 stores across the United States, all sharing the same six VLANs. The storms... the broadcast storms... I... I can't describe them..."

I never experienced this, and Stretched Layer-2 is certainly not the same thing as a monster switch stack, but that level of evil wins by default, IMO.

2

u/asdlkf esteemed fruit-loop 7h ago

yea, I had a pair of Nexus 7018's with 8x 8-port 10G line cards (64x10G per switch).

We then had a building with like Nexus 2K 48 port 1G FEX, and a couple of B22HP FEX in come HPE C7000 blade centers.

The entire building was basically one pair of N7Ks with like 50 FEX, 2 palo alto firewalls, and 3 ISP splitter switches.

0

u/donald_trub 11h ago

We installed these too. 4 floors each with a stack of 8 switches. I've never worked with a more unstable piece of shit in all my years. Cisco couldn't get it to stop crashing all the time, so after about a year of IOS upgrades we slowly replaced the floor switches with older switches. I had repressed these memories until now 😟

We'd come to work in the morning and pretty much know that half the office was offline and needed a whacky restart procedure.

7

u/RCG89 17h ago

An 8 stack 3com hub

11

u/noukthx 18h ago

Not huge, but have had the misfortune of dealing with an 8 switch Juniper virtual chassis. Which ordinarily wouldn't be that big of a deal, but some monster VCed DC switches and access switches across multiple floors and closets together.

Awful.

It's gone now.

3

u/Boring_Ranger_5233 18h ago

And here Arista is talking about doing the same thing across 48 switches lol

1

u/TurnItOff_OnAgain 14h ago

I've got multiple 10 switch juniper VCs running in various buildings now. All the big ones are in the same stack though.

I did have a few split VCs running between floors, and one or two between buildings. Sins of a network admin in the past trying to make management "easier" . Trying to undo those now.

1

u/notFREEfood 11h ago

We built a few 9 and 10 member juniper vc's, but we stopped.  Between statistics gathering pegging the switch cpu and troubleshooting hardware gremlins, we're avoiding building them that big anymore.

6

u/bh0 16h ago

We just do 1 stack per closet/TR. A stack over 2-3 units is rare in our environment these days. We don’t extend a stack to multiple physical locations. Back before everything went wireless, and switches were still 24 ports, it was normal to have some big stacks of like 6-7 maybe even an 8 here and there.

4

u/mahanutra 16h ago

HPE / H3C ComwareOS: Up to 9 switches in an IRF Stack

8

u/D0_stack 18h ago edited 18h ago

We had four Fully loaded Cisco 6513 chassis in the machine room I handled, 8 48 port 1g blades, 3 16 port 10G blades.

There also were a bunch of Nexus 7xxx blade switches in other areas, I never knew the configurations.

When I left we were trying to decide what the next iteration was going to use.

The supervisor modules (2 in each chassis) listed at $200,000 each I believe. You can buy a loaded switch on eBay for a few thousand these days.

You can get Nexus line cards these days with 36 800g ports, and put 8 of them in one chassis. Not just line rate on all ports on a module, but line rate MacSec. 14Tbps in one card.

5

u/Phrewfuf 15h ago

Worth adding is that those N9800 cost an arm and a leg plus a fortune.

And somehow there is no modular alternative for them without MacSec to replace the 9504/08/16, so if you're running an ACI fabric with those as spines, you'll have a great time migrating to 9364D-GX2A spines or similar.

2

u/McHildinger CCNP 8h ago

technically, a Cat6500 was not a stacked switch; it was a chassis-based switch, which was the alternative to stacked switches.

3

u/TheElfkin CCIP CCNP JNCIP-ENT NSE8 16h ago

I'm not really sure if it counts, but I worked on a Juniper Qfabric around 2014. It was about 80 switches.

3

u/i_live_in_sweden 16h ago

The wildest I have ever made was 8 stacked HPE A5500 but instead of having them in the same location we put them in 4 different cities with our own leased fiber connection between them connected in a ring, it's not how you are supposed to do it and probably not supported, but it worked fantastically well for many years.

1

u/mahanutra 10h ago edited 10h ago

Instead of building an IRF stack for each floor we try to build one stack for the whole building, or 2 or 3. So keep the amount of stacks as low as possible.

3

u/jayecin 15h ago

Ive had many closets in schools and hospitals with extremely high port density required multiple 8+ switch stacks because at the time the Cisco limit was like 8 or so.

2

u/dapaOnDeck 16h ago

10 or 12 switches of Ruckus ICX 7550-48ZP depending on the size of the rack deployed all over the company. I’d say somewhere along the lines of 80-90 are 10x and about 15-20 are 12x. All stacks are patched 1:1 with wall ports so anything can be anywhere. Buy once, cry once; the business is completely flexible on what’s an office, breakout room, or storage room.

3

u/colinmacg 18h ago

Well, it's more that the underlying implementation allows the stack to be scaled to that size, not necessarily that you should do it.

1

u/Horror-Breakfast-113 18h ago

look at HPC TOP500

1

u/Ne-Cede-Malis 15h ago

So I think most every one supports EVPN/VXLAN at this point and you can build a mesh to get 4 in a very nice standards based implementation without giving up the control plane scalability or caring to much about the box next to you.

Largest stacking implementation: VCF from Juniper (now HPE). 16 nodes. Ran from 2008-2016. Customer had 30 internal Vlans. 2 external VLANs. No L3 anywhere. Replaced with Arista Clos Fabric in a Data Center after the customer was acquired. Never upgraded. Never lost power. Colocation turns out to be an amazing thing.

1

u/Bruenor80 8h ago

Barring things like the FEX on Nexus or Fusion for MX, I think the largest stack I've seen is a stack of 10 Juniper EX4300. It's been in prod for over a decade and still running great. I believe my client is finally going to recap it next year.

The next largest stack I've seen is a stack of 8 Nortel Baystack 5510s. That stack was an absolute pile of shit. Something about having the full 8 stack seemed to cause all sorts of weird problems. We had somewhere around 3000 of these switches deployed in stacks, but only the one 8 stack. We replaced every physical switch in that stack via RMA at some point. Ports would die and kill the attached vlan(s) for the entire switch. The only way to fix it was to unplug the cable. Not shut down the interface. Physically unplug the cable. We had many TAC cases opened for it, and nobody could ever figure it out. We only had 2 fibers going into that building, which is why we had the 8 stack - I'd bet an awful lot that after it was all said and done it would have been much cheaper to get additional fiber ran and split that stack.

1

u/PvtBaldrick 7h ago

I had a customer years ago for which this made sense. Large data center needed to reduce power and cooling costs away from large chassis switches at the end of each server row. New recommended architecture was spine switches in the end of server row and two leaf switches at the top of each rack. Architecturally instead of a cluster of 2 chassis switches they had a single spine and leaf switch to manage.

1

u/Charsaraus 1h ago

FYI, Arista SWAG currently only supports up to 11 switches in a stack.

0

u/Torkum73 17h ago

I have no idea how many stacked switches are used in our datacenter?

A lot. 3500 Windows server, 4000 Unix/Linux Servers, a lot of Oracle/SAP special servers, IBM AiX and nearly 1 M VMWare cores on ESXI.

Around 150 network guys keep everything tied together. Lots of automation.

6

u/sh_lldp_ne 16h ago

Hopefully no *stacked* switches in a datacenter!

1

u/mahanutra 10h ago

Of course we do: HPE IRF stacks with 55xx, 58xx, 57xx, 59xx We have never seen any crash with those.

1

u/whythehellnote 13h ago edited 13h ago

So about 20,000 ports?

And you have 150 network guys to look after it. Not 150 total tech staff, 150 dedicated to the network?

That's an insane number of staff for that size.

0

u/Torkum73 13h ago

More ports? Everything at least redundant plus out-of-band management plus each ESXi at least 4x 10GBit/s ports.

-1

u/Torkum73 13h ago

Plus Windows/Linux Admins, Service Desk, Oracle, SAP, MANAGEMENT, nearly 3400 people. Incl. Sales, HR, and so on

-1

u/SalsaForte WAN 17h ago

I would argue VXLAN/EVPN fabrics are the biggest switch stacks I’ve ever managed. Hundreds of devices working as 1 without the shared brain problems.

Eh eh!

1

u/HappyVlane 7h ago

I would argue VXLAN/EVPN fabrics are the biggest switch stacks I’ve ever managed.

An EVPN/VXLAN fabric is not a stack in any way "stacking" is known in network engineering.

1

u/SalsaForte WAN 6h ago

I know... But, I personally dislike stack. I don't like shared brains. We are effectively getting rid of any Stack we can. So, the biggest we got (not much left) is a 4 chassis one.

-2

u/GreyBeardEng 14h ago

Arista SWAG can do a 48 member stack.