
800G Ethernet Switches: Does Your AI Cluster Need the SN5600?
800G switches like the NVIDIA SN5600 pay off in large AI fabrics: 64 ports of 800GbE double the 400G radix, cutting tiers, switch count and hops.
Most AI clusters buying networking today still connect each GPU at 400G, so the question is not whether 800G is faster. It is whether an 800G switch makes your fabric smaller, flatter and simpler to operate. For large clusters it usually does: a switch with 64 ports of 800GbE can serve twice as many 400G endpoints as a 64-port 400G switch, which removes a tier of switching, a layer of optics and a hop of latency at scale. For a cluster of a few racks, a good 400G leaf-spine remains the sensible design. This guide covers the 800G decision specifically, using the NVIDIA Spectrum-4 SN5600 as the reference switch.
What is the NVIDIA Spectrum SN5600?
The NVIDIA Spectrum-4 SN5600 is a 2U Ethernet switch built on the Spectrum-4 ASIC. It has 64 x 800GbE OSFP ports plus one 25GbE SFP28 port, 51.2 Tb/s of switching capacity and a packet rate of 33.3 billion packets per second. Fabric features include RoCEv2, adaptive routing and Spectrum-X. It runs an open network operating system, either Cumulus Linux or SONiC, and has redundant hot-swap power supplies with front-to-back airflow. It is built to serve as a leaf or spine in GPU fabrics rather than as a campus or enterprise access switch.
Do I need 800G switches for an AI cluster?
Work it out from radix, meaning the number of ports per switch at the speed your endpoints use. In a two-tier, non-blocking leaf-spine, each leaf uses half its ports for servers and half for uplinks, so the theoretical maximum is radix squared divided by two. A 32-port 400G switch such as the NVIDIA Spectrum-3 SN4700 tops out at 512 endpoints at 400G in two tiers. A 64-port 400G switch reaches 2,048. An SN5600 with every port split into two 400G links behaves as a 128-port 400G switch, which reaches 8,192. For a 1,024-GPU cluster with one 400G adapter per GPU, that works out to 16 leaves and 8 spines of SN5600 against 32 leaves and 16 spines of 64-port 400G switches, with half as many switch-to-switch links, each running at 800G. Real designs adjust these numbers for rail-optimized layouts and oversubscription, but the ratio holds.
- Choose 800G if the cluster is heading past what 400G-radix switches can build in two tiers, or you plan to grow into that range.
- Choose 800G if your next GPU platforms ship with 800G-class network adapters, so the switch matches the host.
- Choose 800G for the spine of an existing 400G fabric, where fewer, faster uplinks cut switch count and cabling.
- Choose 800G if you want Spectrum-X behavior, with adaptive routing and congestion control tuned for AI traffic across switch and NIC.
Stay at 400G if the cluster is a single rack or a handful of 8-GPU nodes, if one or two 400G switches already cover every port you need, or if your team and optics inventory are standardized on QSFP-DD. Choosing among 400G switches is a separate decision, covered on our 400G Ethernet page and in our guide to 400GbE fabrics for multi-node GPU clusters.
How does an 800G switch compare with 400G switches on paper?
Be careful with headline capacity figures, because vendors do not all count the same way. Some quote traffic in one direction, others count both directions. The reliable comparison is port count multiplied by port speed. The SN5600's 64 ports at 800G are 51.2 Tb/s each way. A 64-port 400G switch is 25.6 Tb/s each way, even where a datasheet quotes 51.2 Tb/s counting both directions, as the Quantum-2 InfiniBand switches do.
- NVIDIA Spectrum-4 SN5600: 64 x 800GbE OSFP, 2U, 33.3 billion packets per second, Cumulus Linux or SONiC.
- NVIDIA Spectrum-3 SN4700: 32 x 400GbE QSFP-DD, 1U, 12.8 Tb/s, 8.4 billion packets per second, 64 MB fully shared buffer, Cumulus Linux or SONiC.
- Cisco Nexus 9364D-GX2A: 64 x 400G QSFP-DD, 2RU, NX-OS or ACI spine, with breakout to 4 x 100G or 2 x 200G per port.
- Cisco Nexus 9332D-GX2B: 32 x 400G QSFP-DD, 1RU, with breakout up to 128 x 100G or 64 x 200G.
800G vs 400G: what changes for optics and cabling?
More than the switch itself. The SN5600 uses OSFP cages, while most 400G data center switches, including the SN4700 and the Cisco models above, use QSFP-DD. The two form factors are not interchangeable, so an existing stock of QSFP-DD optics and cables does not carry over. Our explainer on 100G vs 400G transceivers covers the underlying module types.
- Lane rate: 800G runs eight electrical lanes at 100G each. An 800G port split into two 400G links therefore pairs with 400G endpoints that also use 100G lanes, such as OSFP-based NVIDIA ConnectX-7 400G adapters, not with older 400G optics built on 50G lanes.
- Twin-port optics: 800G OSFP modules are commonly twin-port, carrying two independent 400G links on two fiber connectors, so one switch port cables to two different servers or two different spines.
- Matching standards: both ends of every link must use the same optical standard, single-mode DR variants for longer runs or multimode SR variants for short ones.
- Copper reach: at 100G per lane, passive copper cables are limited to short in-rack runs. Plan optics or active cables for anything longer.
- Power and heat: 800G optics draw more power per module, so check switch airflow direction (front-to-back on the SN5600) against your hot and cold aisles.
- Fiber plant: twin-port optics double the connectors per cage, so count MPO trunks and patch panel capacity, not just switch ports.
What does Spectrum-X add, and which network OS runs on it?
Spectrum-X is NVIDIA's name for an Ethernet fabric built from Spectrum-4 switches and NVIDIA SuperNICs, with adaptive routing and congestion control coordinated between switch and NIC so large collective operations do not pile onto the same paths. The end-to-end features depend on compatible NICs at the hosts. Without them, the SN5600 still works as a high-radix RoCEv2 switch, and you configure lossless behavior with PFC and ECN in the usual way, as described in our RoCE v2 configuration guide. The operating system choice is Cumulus Linux, NVIDIA's own Linux-based NOS, or SONiC, the open-source NOS. Pick the one your network team can automate and support, because that decision outlasts the hardware.
Should I choose 800G Ethernet or 400G InfiniBand?
They are different fabrics serving the same job. The NVIDIA Quantum-2 QM9700 is a 1U InfiniBand switch with 64 NDR 400Gb/s ports on 32 OSFP cages (the same twin-port idea as 800G Ethernet), SHARPv3 in-network collective acceleration and an embedded subnet manager for up to about 2,000 nodes. The QM9790 is the externally managed variant, run from NVIDIA UFM. InfiniBand is the established choice for dedicated training clusters. Ethernet with the SN5600 suits teams that want one network technology across AI, storage and the rest of the data center, or multi-tenant clusters. The trade-offs are covered in our Ethernet vs InfiniBand TCO analysis.
An 800G switch is mostly a radix decision, not a speed decision: buy it when doubling the number of 400G endpoints per switch removes a tier, a layer of optics or a rack of switches from your design.
Nexus Compute supplies NVIDIA Spectrum and Quantum switches, ConnectX adapters and matching optics through authorized channels, with manufacturer warranty, and designs the fabric through our data center networking solutions service. Send us the GPU count, adapters per node and growth plan, and we will return a leaf-spine design with a bill of materials for switches, optics and cabling. Pricing depends on port count, optics and support term, and we quote within 48 business hours via request a quote.
Frequently asked questions
Do I need 800G switches for an AI cluster?
Only when the fabric is large enough that 400G-radix switches would force a third tier or a very large switch count, or when your GPU nodes use 800G-class adapters. A cluster of a few racks on 400G adapters is usually well served by a 400G leaf-spine. 800G is also a strong choice for the spine layer of a growing 400G fabric.
800G vs 400G: what changes for optics and cabling?
800G switches such as the SN5600 use OSFP cages rather than QSFP-DD, run 100G per electrical lane, and commonly use twin-port optics that carry two 400G links per port. Existing QSFP-DD optics do not carry over, both ends of each link must use matching optical standards, and fiber connector counts rise.
What is the NVIDIA Spectrum SN5600?
The SN5600 is a 2U NVIDIA Spectrum-4 Ethernet switch with 64 x 800GbE OSFP ports plus 1 x 25GbE SFP28, 51.2 Tb/s switching capacity and 33.3 billion packets per second. It supports RoCEv2, adaptive routing and Spectrum-X, and runs Cumulus Linux or SONiC.
Can I connect 400G network adapters to an 800G switch?
Yes, by splitting an 800G port into two 400G links with twin-port optics or breakout cables, which is how most current 400G GPU nodes attach to 800G switches. The 400G endpoints need to use 100G lanes, as OSFP ConnectX-7 adapters do. Confirm the supported breakout modes for your network OS release and optics.
Is the SN5600 an InfiniBand switch?
No. The SN5600 is an Ethernet switch. NVIDIA's InfiniBand equivalents are the Quantum-2 switches, such as the QM9700 and QM9790, with 64 NDR 400Gb/s ports each.
Systems covered in this article
Planning a hardware investment?
Tell us what you're trying to build. A procurement specialist will help you specify and quote the right configuration within 48 business hours, no obligation.
