The NVIDIA Switch Field Guide: Models, NVUE CLI, and Protocol Support

A translation guide for network engineers coming from Cisco NX-OS, IOS-XR, or Arista EOS — with model taxonomy, current CLI, and honest protocol coverage.
0. Who This Is For
If you’ve spent your career in configure terminal, write memory, and show ip route, and you’re now being handed a rack of switches that boot into a Linux prompt and want you to type nv set instead of no shutdown, this post is for you.
NVIDIA has, through the Mellanox and Cumulus acquisitions, built a switching business that behaves nothing like the vendors most CCIEs grew up on. Different OS lineage. Different CLI grammar. Different protocol emphasis. And a hard split between two product families — Spectrum (Ethernet) and Quantum (InfiniBand) — that solve overlapping problems with completely different tooling.
There is no shortage of documentation for these platforms. What’s missing is a single map that says: here are the models, here is the CLI, here is what the boxes can and can’t do, here is how it maps to what you already know.
This is that map.
1. The Taxonomy: Two Families, One Company
NVIDIA switching, at a product level, is two lines:
- Spectrum — Ethernet switches, built on the Spectrum ASIC family (currently Spectrum-4). The line most enterprises and modern AI Ethernet fabrics buy.
- Quantum — InfiniBand switches, built on the Quantum ASIC family (currently Quantum-3). Used in HPC and large-scale AI training where InfiniBand’s native lossless behavior and in-network compute matter more than Ethernet ubiquity.
The company also sells BlueField DPUs and ConnectX SmartNICs, which sit at the endpoints — but this post is about the switches.
Historical layers
Understanding why the CLI looks the way it does helps:
| Year | Event | Consequence for the CLI |
|---|---|---|
| 2016 | Mellanox releases Onyx / MLNX-OS | Cisco-like modal CLI on Ethernet switches |
| 2020 | NVIDIA acquires Cumulus Networks | Cumulus Linux inherits Ethernet OS lineage |
| 2020 | NVIDIA acquires Mellanox | Quantum + Spectrum + ConnectX join the portfolio |
| 2021+ | NVUE (“NVIDIA User Experience”) | Unified nv command set across product lines |
| 2024+ | Pure SONiC support on Spectrum-4 | Community OS as first-class option |
The result: today, on any modern NVIDIA switch, you can log in and start typing nv show ... and it works — regardless of whether the underlying platform is a Quantum InfiniBand box or a Spectrum Ethernet box. That single-CLI story is real, and it is genuinely helpful.
2. Spectrum Ethernet Lineup
See front-panel reference diagram — Spectrum-4 flagships, Quantum-3 InfiniBand, and Quantum-2 legacy shown to relative scale (1U : 90px).
Spectrum has gone through four ASIC generations. Only the latest two are worth buying new for AI or greenfield DC builds.
Spectrum-4 (current flagship)
| Model | Ports | Throughput | Role | Notes |
|---|---|---|---|---|
| SN5600 | 64 OSFP at 800GbE + 1 SFP28 at 25GbE | 51.2 Tb/s | Leaf / spine / super-spine | 2U · 160MB shared buffer · Spectrum-X flagship |
| SN5600D | 64x 800GbE OSFP (OCP 21″ chassis) | 51.2 Tb/s | High-density AI fabric | 2U · OCP form factor for hyperscale racks |
| SN5610 | Variant of SN5600 | 51.2 Tb/s | AI fabric | 2U · deployment-specific configuration |
| SN5400 | 64 QSFP-DD at 400GbE + 2 SFP28 at 25GbE | 25.6 Tb/s | Leaf / spine | 2U · workhorse 400G tier |
The flagship SN5600 is the box that shows up in most current AI Ethernet designs. In a rail-optimized SU, one SN5600 per rail can carry all 32 back-end connections for a full SU as leaf, and a pool of SN5600s serves as spine for cross-SU traffic. The Spectrum-4 ASIC delivers sub-microsecond cut-through latency with a fully shared packet buffer — architecturally important because it means every port can absorb bursts against the full buffer, not a sliced share.
Key Spectrum-4 capabilities worth naming for a network engineer’s mental model:
- Ethernet VPN (EVPN) multi-homing and 256-way ECMP — the routing scale is real, matching or exceeding modern Broadcom silicon
- Multi-chassis LAG for active/active L2 multipathing
- Hardware Assisted In Service Software Upgrades (ISSU) — a real property when you cannot take a training run down
- RoCE extensions specific to NVIDIA Spectrum-X, plus nanosecond-level timing precision from switch to host
Spectrum-X: the AI Ethernet reference architecture
Spectrum-X is not a switch — it is a reference architecture that combines Spectrum-4 switches with BlueField-3 SuperNICs and specific tuning of RoCE, adaptive routing, and congestion control. NVIDIA claims a 1.6× performance improvement over traditional Ethernet fabrics for generative AI workloads with Spectrum-X. The mechanisms include:
- Adaptive routing that reacts to congestion in hardware (unlike static ECMP)
- Fine-grained per-packet spraying with reordering at the SuperNIC
- Multi-tenant performance isolation
- Deep telemetry integration for job-level visibility
When you see the phrase “Spectrum-X switch” in a spec sheet, it means a Spectrum-4 switch running the Spectrum-X features — same silicon, different feature enablement and design pattern.
Older Spectrum generations
- Spectrum-3 (SN4000 series) — 100/200/400G, first Mellanox generation with production RoCE at scale. Still deployed but past its refresh point.
- Spectrum-2 (SN3000 series) — 25/50/100/200G. Enterprise DC. Not a modern AI fit.
- Spectrum-1 (SN2000 series) — legacy.
If you’re inheriting an existing fabric, expect Spectrum-3 in most 2020–2023 builds. If you’re designing new, plan on Spectrum-4.
3. Quantum InfiniBand Lineup
The Quantum line is deliberately narrower — InfiniBand is a single-vendor game, and NVIDIA is that vendor.
Quantum-3 (current, “Quantum-X800”)
| Model | Ports (OSFP cages / lanes) | Speed per port | Role | Notes |
|---|---|---|---|---|
| Q3400-RA | 144 XDR ports over 72 OSFP cages | 800 Gb/s (XDR) | Air-cooled leaf/spine | 4U · supports back-compat to NDR/HDR/EDR/FDR/QDR |
| Q3401-RD | 144 XDR ports | 800 Gb/s | DC-power variant | 4U · same silicon, DC power |
| Q3450-LD | 144 XDR ports (co-packaged optics) | 800 Gb/s | Liquid-cooled, CPO | Silicon photonics integrated with the switch ASIC — no pluggable transceivers |
| Q3200-RA | Two independent 36-port switches in a single 2U enclosure | 800 Gb/s | Compact / step-in | Ideal for bridging to legacy Quantum-2 fabrics |
The Q3400-RA is the density workhorse. Its high radix supports a two-level fat-tree topology capable of connecting up to 10,368 ConnectX-8 NICs — for context, that’s roughly enough to interconnect 1,296 eight-GPU nodes in one non-blocking fabric. That’s a real GPU cluster, not a marketing number.
The Q3450-LD is the interesting one architecturally. By integrating silicon photonics directly with the switch ASIC via co-packaged optics, it eliminates pluggable transceivers and improves power and thermal efficiency. This is the direction the industry is heading — once you’re at 800G+ per port, the reach and power draw of pluggables becomes a system-level problem, and CPO is the answer that stops the bleeding.
Quantum-2 (still deployed)
- QM9700 / QM9790 — NDR 400 Gb/s per port, 64 ports, 1U. This is what most 2023–2024 GPU clusters run today. Still perfectly viable; you just don’t buy new ones for greenfield 2026 designs when Quantum-3 is available.
The InfiniBand feature that Ethernet doesn’t have
Every Quantum switch carries the fourth-generation NVIDIA Scalable Hierarchical Aggregation and Reduction Protocol (SHARP), plus adaptive routing, telemetry-based congestion control, and self-healing networking. SHARP is the one that matters most: the switch itself performs the reduction arithmetic for collectives like AllReduce, so instead of every GPU sending its full gradient across the fabric, the switches aggregate on the way up and broadcast the result on the way down. On a large training cluster this can cut collective communication time by a factor of 2 or more.
There is no direct Spectrum-4 equivalent. Ethernet doesn’t do in-network reduction natively — you can get close with in-band programmable pipelines (P4 on some hardware), but SHARP is a first-class InfiniBand feature.
4. The Operating Systems: NVOS, Cumulus Linux, SONiC, and Legacy Onyx
Which OS runs on which platform is a live question because NVIDIA has been actively unifying the story.
| OS | Runs on | Status | CLI |
|---|---|---|---|
| NVOS | Quantum InfiniBand, NVLink switches, newer Spectrum | Active, unified target | NVUE (nv commands) |
| Cumulus Linux | Spectrum Ethernet | Active | NVUE (nv commands) + Linux underneath |
| SONiC | Spectrum-4 (SN5600 family) | Fully supported | SONiC CLI (Cisco-like, sonic-cli) |
| Onyx / MLNX-OS | Older Mellanox Ethernet | Legacy, deprecated | Modal, Cisco-like |
The key convergence point is NVUE — the CLI grammar that spans both NVOS and Cumulus Linux. If you learn nv show and nv set on a Cumulus box, those skills transfer directly to an NVOS Quantum box. This is the single most important thing to understand about the NVIDIA CLI story: one command language, two OS bases, both product families.
The SONiC option
Pure SONiC is fully supported on SN5600 switch systems. This is not a “hobbyist” statement — it means you can run community SONiC on production Spectrum-4 hardware and NVIDIA will honor the hardware warranty. The community open network OS movement has genuine traction here, and if your organization has SONiC operational expertise, this is a real path.
That said, if you’re coming from Cisco/Arista and need to learn something, learn NVUE first. It is the CLI NVIDIA is unifying around, and it maps most naturally to what a modern data center engineer needs to do day-to-day.
5. NVUE: The CLI You Actually Need to Learn
Here is where the mental shift from Cisco/Arista happens. Stay with me.
The four command verbs
NVUE commands all begin with nv and fall into one of four syntax categories: configuration (nv set and nv unset), monitoring (nv show), action commands (nv action), and configuration management (nv config).
That’s it. Four verbs. Every operation on the switch — from checking interface counters to reloading firmware to committing a route-map — collapses into one of those four.
Flat, not modal
The NVUE CLI has a flat structure as opposed to a modal structure — you can run all commands from the primary prompt instead of only in a specific mode.
This is a genuine break from Cisco IOS/NX-OS, where your prompt state (>, #, (config)#, (config-if)#) tells you where you are and what you’re allowed to type. On NVUE, there is no configure terminal. There is no interface Ethernet1/1 mode. You type the whole path in one line:
nv set interface eth1/1 ip address 10.1.1.1/31
That single command targets interface eth1/1, sets its ip address sub-property, and gives it the value. No mode changes. Tab completion walks you through the tree at every step.
Coming from Cisco, this feels alien for about a day and then feels obviously better.
The config commit workflow
This is the other break from Cisco NX-OS habit. On NVOS/Cumulus, config changes are staged in a pending buffer, and only take effect when you commit them:
nv set interface eth1/1 ip address 10.1.1.1/31 # stage the change
nv config diff # review what you've staged
nv config apply # apply pending config to running
nv config save # persist running config to boot
This is closer to Junos than to IOS, and it is the right model. nv config apply is atomic — either all your staged changes take effect together, or none of them do. No more “I typed one command and half the fabric went dark.”
You can also roll back:
nv config apply rev1 # apply a previous revision
nv config apply empty # revert to blank config
The revision system is a real safety net. It exists on NX-OS via checkpointing, but on NVUE it is native to the workflow.
Help, completion, and abbreviation
Tab completion, ? for value-type help, and abbreviation are all supported — nv sh int expands to nv show interface, and ambiguous commands are flagged with the choices.
Practically:
<Tab>at any point shows valid completions?shows the value type and description for a specific option-hor--helpgives context help- Abbreviations work as long as they are unambiguous
If you learn nv show, nv set, nv config apply, and nv config save, you can navigate 80% of daily operations. Add nv unset for removing configuration and nv action for operational triggers (like reboots), and you have the vocabulary.
Practical examples
Set a hostname:
nv set system hostname leaf-01
nv config apply
Configure a routed L3 interface:
nv set interface swp1 type swp
nv set interface swp1 ip address 10.255.0.1/31
nv set interface swp1 link mtu 9216
nv config apply
Bring up BGP:
nv set router bgp autonomous-system 65001
nv set vrf default router bgp router-id 10.255.0.1
nv set vrf default router bgp neighbor 10.255.0.0 remote-as 65002
nv config apply
Enable PFC on priority 3 (RoCEv2 lossless):
nv set qos roce mode lossless
nv set qos roce congestion-control mode ecn
nv set interface swp1 qos pfc mode on
nv config apply
Look at BGP state:
nv show router bgp
nv show vrf default router bgp neighbor
Look at what interfaces are actually doing:
nv show interface
nv show interface swp1 counters
nv show interface swp1 pluggable
Everything is discoverable via Tab completion. There is no memorizing 200 show commands like on NX-OS. You just start typing nv show and let the CLI show you what’s available.
REST API — same schema, different transport
Every NVUE command has an equivalent REST endpoint. For example, setting a hostname maps to PATCH https://<ip>/nvue_v1/system with the hostname field in the body. This means anything you can do at the CLI, you can automate via HTTP — no expect scripts, no screen-scraping. Ansible, Terraform, and custom tooling all have first-class API integration.
6. Protocol Support — Spectrum (Ethernet)
Here is the protocol matrix that matters for design decisions. This is not exhaustive — it is the “will it do what my design needs?” list.
Layer 2
| Protocol | Support | Notes |
|---|---|---|
| VLAN (802.1Q) | Yes | Standard trunking and access |
| LACP (802.3ad) | Yes | Multi-link aggregation, active-active |
| MLAG | Yes | Multi-chassis LAG for active-active L2 to hosts |
| STP / RSTP / MSTP | Yes | Discouraged in DC leaf-spine; use L3 or EVPN |
| LLDP | Yes | Standard neighbor discovery |
| MACsec | Yes (Spectrum-4) | Link encryption at line rate |
Layer 3 / Routing
| Protocol | Support | Notes |
|---|---|---|
| Static routing | Yes | |
| BGP (unicast + multipath) | Yes | FRR under the hood; robust, mature |
| OSPF v2/v3 | Yes | FRR |
| BFD | Yes | Sub-second failure detection |
| ECMP | Yes | Up to 256-way on Spectrum-4 |
| VRR | Yes | Cumulus’s active-active gateway; alternative to VRRP |
| VRF | Yes | Full L3 isolation |
Overlays and virtualization
| Protocol | Support | Notes |
|---|---|---|
| VXLAN (RFC 7348) | Yes | Hardware VTEP at line rate |
| EVPN (BGP) | Yes | Full EVPN-VXLAN control plane |
| EVPN Type-2/3/5 | Yes | MAC/IP, IMET, IP prefix |
| EVPN Multi-homing | Yes | LAG replacement across leaves |
Lossless / RoCE / QoS
| Feature | Support | Notes |
|---|---|---|
| PFC (802.1Qbb) | Yes | Per-priority pause, headroom-aware |
| ECN | Yes | WRED with ECN marking, DCQCN-ready |
| Shared buffer | Yes (Spectrum-4) | Single monolithic buffer across all ports |
| DCQCN | Ecosystem | RNIC-side control loop; switch marks ECN |
| Adaptive Routing (Spectrum-X) | Yes | Hardware-level congestion-aware forwarding |
Telemetry and operations
| Feature | Support | Notes |
|---|---|---|
| gNMI / gRPC | Yes | Streaming telemetry, YANG-modeled |
| WJH (What Just Happened) | Yes (Spectrum-2+) | Per-packet drop reasons — the killer telemetry feature |
| sFlow | Yes | Standard sampling |
| SNMP v2c/v3 | Yes | Legacy but present |
| NetQ | Yes | NVIDIA’s fabric telemetry and validation platform |
On WJH specifically: if you have not used What Just Happened, it is one of the genuinely differentiating features of Spectrum silicon. Instead of just incrementing a drop counter, WJH tells you why the packet was dropped — MTU mismatch, buffer overflow, ACL, TTL, forwarding lookup miss — with the packet header so you can identify the flow. On RoCE fabrics where a single silent drop can stall a training run, this is not a nice-to-have.
7. Protocol Support — Quantum (InfiniBand)
Different world, different protocols. If you have never worked on InfiniBand before, the vocabulary is unfamiliar but the underlying ideas are recognizable.
Native InfiniBand transport
- Reliable Connection (RC) — the transport most GPU workloads use; equivalent to a lossless TCP with RDMA semantics
- Unreliable Datagram (UD) — used for control and multicast
- Extended Reliable Connection (XRC) — scale optimization for very large clusters
Subnet Manager (SM) — the plane you don’t have on Ethernet
InfiniBand fabrics are controlled by a Subnet Manager, a centralized agent that discovers the topology, computes Linear Forwarding Tables (LFTs), and pushes them to every switch. There is no distributed routing protocol like BGP in the fabric — the SM does the work globally.
On modern deployments, the SM runs in NVIDIA UFM (Unified Fabric Manager), which also handles telemetry, alerting, and job-level performance visibility. UFM is not optional at scale — it is the operational plane for the fabric.
SHARP v4 — in-network reduction
The feature that has no Ethernet equivalent. When a collective operation happens (e.g., NCCL AllReduce), the switches themselves perform the reduction arithmetic as data flows up the tree, and broadcast the result back down. Instead of every GPU sending its full gradient to every other GPU, the fabric carries a fraction of the data and the switches do the math.
This is what makes InfiniBand competitive for large-model training even against faster raw Ethernet. Bandwidth saved is training time saved.
Adaptive Routing (AR)
Each switch can hash flows across equal-cost paths in hardware based on real-time congestion, not just static 5-tuple. This addresses the ECMP hash polarization problem structurally rather than through statistical entropy tuning. On Spectrum-X Ethernet you get an analogous feature; on Quantum, it has been production-hardened for a decade.
SHIELD — Self-Healing Networking
Sub-second link failover and topology reconvergence without host visibility. When a link fails, the fabric reroutes internally and the RNIC never sees the transition. On Ethernet you get this via BGP + ECMP + BFD, but it takes tens of milliseconds. SHIELD is faster and simpler because the fabric owns the routing decision.
Telemetry-based Congestion Control (CC)
InfiniBand-native congestion control that uses per-switch telemetry to rate-limit senders. This is the equivalent function to DCQCN on RoCE, but implemented natively rather than as a bolted-on feedback loop.
Router functionality
Quantum-X800 switches feature optional router capabilities, enabling the expansion of InfiniBand clusters to support a large scale of nodes across multiple sites. This lets you stitch multiple InfiniBand subnets together — historically a hard problem, and a common reason people fell back to Ethernet for multi-site fabrics.
8. Practical NVUE Config: A Working Leaf Switch
Here is the vocabulary in action. This is a minimum-viable leaf switch config for a RoCEv2 fabric on a Spectrum-4 SN5600 running Cumulus Linux.
Bring up the physical interfaces
nv set interface swp1-32 type swp
nv set interface swp1-32 link mtu 9216
nv set interface swp1-32 link fec rs
nv config apply
This sets 32 switch ports as generic L3-ready interfaces with jumbo MTU and RS FEC (required at 400G/800G).
Configure L3 routed uplinks to spine
nv set interface swp33 ip address 10.255.0.1/31
nv set interface swp33 link mtu 9216
nv set interface swp34 ip address 10.255.0.3/31
nv set interface swp34 link mtu 9216
nv config apply
Point-to-point /31s to two spine switches. Standard leaf-spine underlay.
BGP underlay
nv set router bgp autonomous-system 65001
nv set vrf default router bgp router-id 10.255.1.1
nv set vrf default router bgp neighbor 10.255.0.0 remote-as external
nv set vrf default router bgp neighbor 10.255.0.2 remote-as external
nv set vrf default router bgp address-family ipv4-unicast redistribute connected
nv config apply
BGP unnumbered would also work here, and NVUE supports it natively. This example uses explicit /31s for clarity.
RoCE lossless class
nv set qos roce mode lossless
nv set qos roce congestion-control mode ecn
nv set qos roce classification trust dscp
nv set qos roce priority 3
nv set interface swp1-34 qos pfc mode on
nv config apply
This is the NVUE equivalent of the Nexus config we walked through in the last post — DSCP-trusted classification, priority 3 for the lossless class, ECN marking for congestion control, PFC enabled on every RoCE-carrying interface.
Show commands for verification
nv show interface swp1 counters
nv show interface swp1 pluggable # optics info
nv show vrf default router bgp neighbor summary
nv show qos roce
nv show qos roce counters
nv show qos pfc
The nv show command tree is discoverable — Tab completion at any level shows what’s queryable.
9. Translating from Cisco / Arista: A Cheat Sheet
For engineers with muscle memory in NX-OS or EOS, here is the direct mapping for the most common daily operations.
| Task | Cisco NX-OS | Arista EOS | NVUE |
|---|---|---|---|
| Enter config mode | configure terminal | configure | (no mode — configure from prompt) |
| Set hostname | hostname leaf-01 | hostname leaf-01 | nv set system hostname leaf-01 |
| Configure L3 interface | interface Eth1/1 → no switchport → ip address ... | interface Ethernet1/1 → no switchport → ip address ... | nv set interface swp1 ip address ... |
| Save config | copy running-config startup-config | write memory | nv config save |
| Apply pending changes | (immediate) | (immediate) | nv config apply |
| Show interfaces | show interface | show interfaces | nv show interface |
| Show BGP neighbors | show ip bgp summary | show ip bgp summary | nv show vrf default router bgp neighbor |
| Show routing table | show ip route | show ip route | nv show vrf default router route |
| Set MTU | mtu 9216 (under interface) | mtu 9216 | nv set interface swp1 link mtu 9216 |
| Enable PFC | priority-flow-control mode on (under intf) | priority-flow-control on | nv set interface swp1 qos pfc mode on |
| Configure BGP AS | router bgp 65001 | router bgp 65001 | nv set router bgp autonomous-system 65001 |
| Reboot | reload | reload now | nv action reboot system |
The mental shift is straightforward once you accept it: Cisco/Arista treat the config as a sequence of commands you type. NVUE treats the config as a tree of objects you set values on. Once that clicks, NVUE feels faster and less error-prone.
10. When Do You Actually Buy NVIDIA vs. Cisco / Arista / Juniper?
An honest answer based on where each vendor wins today.
Buy Spectrum-4 when…
- You are building an AI Ethernet fabric and want Spectrum-X features (adaptive routing, congestion control, SuperNIC integration)
- You want a single-vendor stack from GPU → NIC → switch, and you’re already NVIDIA on the compute side
- You need very deep telemetry per packet (WJH is genuinely differentiating)
- Your team is comfortable with Linux-based network OS and NVUE / Cumulus / SONiC
- You want a multi-OS story on the same silicon (Cumulus, SONiC, and community options)
Buy Quantum when…
- You are building large-scale AI training and can pay for InfiniBand
- You need SHARP for collective offload
- Your workloads are latency-sensitive HPC (weather, molecular dynamics, seismic)
- You want the operational maturity of an InfiniBand ecosystem: UFM, adaptive routing, self-healing
Stay with Cisco Nexus / Arista when…
- You have a large existing NX-OS or EOS operational estate and toolchain
- Your regulatory/procurement environment demands specific vendors
- You need protocols NVIDIA doesn’t emphasize (MPLS-TE, extensive SR-MPLS, complex service-provider features)
- Your team’s expertise is deep and your workload is enterprise DC, not AI training
The honest overlap
For a modern AI backend Ethernet fabric, Spectrum-4 with Spectrum-X features is genuinely competitive with the best Broadcom Tomahawk 4/5-based switches from Arista or Cisco. The choice increasingly comes down to operational preference, tooling investment, and how deep your team wants to go into the NVIDIA stack.
For InfiniBand, there is no meaningful choice — you buy Quantum or you don’t buy InfiniBand.
11. Closing
The learning curve for NVIDIA switches is real but short. If you already understand what a leaf-spine fabric does, what BGP looks like, what PFC and ECN mean, and what a lossless class buys you — the NVUE grammar takes a week to internalize, and the protocol semantics take a day.
The rewards are meaningful: a single CLI across two very different underlying product families, a flat command model that ages better than the modal Cisco tradition, native support for the RoCE and IB features that AI workloads actually need, and per-packet telemetry that turns silent drops into actionable data.
If your career is being pulled toward AI infrastructure — and if you are a network engineer in 2026, it probably is — the NVIDIA switch stack is one of the two ecosystems you have to be conversant in. This post is the map. The lab is where you learn to walk it.
— posted on networkbachelor.com