Short answer: NDR InfiniBand — the fabric that MMA4Z00-NS400 and MMA4Z00-NS optics carry — exists to move data between GPUs across an entire cluster fast enough that the GPUs aren't sitting idle waiting on the network during distributed training. It's the standard scale-out fabric behind NVIDIA's own DGX H100 and H200 SuperPOD reference architecture, and it's what a network architect specs when a training job spans more GPUs than fit in one server. If you're running a handful of GPUs in a single box, you likely don't need it yet. If you're linking multiple GPU servers into one training job, this is almost certainly the fabric NVIDIA's own reference design points you toward.
The Problem This Fabric Actually Solves
Training a large model rarely happens on one server's worth of GPUs anymore. The workload gets split across many GPUs in many servers, and at regular intervals during training — after every batch, in most setups — those GPUs have to synchronize: an operation usually called AllReduce, where every GPU shares its computed gradients with every other GPU before the next step can start. Every GPU in the job waits on that synchronization to finish.
That makes network latency and bandwidth a direct multiplier on training cost, not a background concern. A few extra microseconds of latency per AllReduce, multiplied across thousands of training steps and hundreds of GPUs, adds up to real wall-clock time and real compute spend sitting idle. This is the specific problem NDR InfiniBand is built to minimize: a lossless, low-latency fabric with RDMA (remote direct memory access) that moves data GPU-to-GPU without routing it through the CPU, at 400Gb/s per port.
Where MMA4Z00-NS400 and MMA4Z00-NS Fit in a Real Build
NVIDIA's own DGX SuperPOD reference architecture for H100 and H200 systems specifies NDR InfiniBand as the scale-out fabric, built on exactly this pair of parts:
- ConnectX-7 HCAs in each DGX server, populated with MMA4Z00-NS400 flat-top optics — one port per GPU in a fully non-blocking rail-optimized design
- Quantum-2 switches (QM9700/QM9790), populated with MMA4Z00-NS finned-top twin-port optics on the leaf and spine layers
- A fat-tree or rail-optimized topology, where each GPU's NIC connects to a dedicated "rail" switch, so GPU-to-GPU traffic across servers stays predictable and non-blocking even at full cluster load
This is also why the two parts pair the way they do: one twin-port MMA4Z00-NS switch module breaks its 800G down into two independent 400G links, feeding two ConnectX-7 ports on two different servers. That 1:2 ratio isn't a cost-saving shortcut — it's how NVIDIA's own reference topology is drawn.
Who Is Actually Buying This, and Why
The buyer for switch-side NDR optics specifically is rarely a single-GPU hobbyist — the twin-port Quantum-2 module and the rail-optimized topology it belongs to only make sense once you're wiring multiple servers together. In practice, that's a fairly specific set of builds:
- Enterprise AI teams standing up an on-premises training or fine-tuning cluster on H100 or H200 hardware, following NVIDIA's published SuperPOD design rather than inventing a topology from scratch.
- HPC and research labs extending existing InfiniBand fabric to add GPU nodes, where NDR is often already the standard the rest of the fabric runs on.
- GPU cloud and colocation providers building out H100/H200 capacity for resale, where port count and cost-per-port matter enough that a $250–300 gap per module across a few hundred ports is a real line item.
- Teams repurposing secondary-market H100/H200 hardware — a fast-growing segment as Blackwell-generation systems roll out and displace Hopper-era gear onto the used market, arriving without cabling or optics included.
What these buyers have in common is that the decision is procurement-driven, not discovery-driven: they already know they need NDR InfiniBand because NVIDIA's reference architecture told them so, and the open question is where to source the optics economically and on a lead time that doesn't stall the project — not whether the technology is worth adopting in the first place.
InfiniBand vs. Ethernet: Is NDR Even the Right Fabric?
Not every GPU cluster needs InfiniBand. If your build is inference-serving rather than distributed training, or sits under roughly 64 GPUs, a well-configured 400GbE/RoCE v2 Ethernet fabric is increasingly a legitimate choice — we cover that decision in full in InfiniBand vs. Ethernet for AI Clusters in 2026. NDR earns its cost and complexity specifically at the point where large-scale distributed training's tight AllReduce pattern makes every microsecond of latency count — generally 512+ GPU training runs, or any build following NVIDIA's DGX SuperPOD reference design directly.
Sizing the Optics for Your Build
Once NDR is the right call, sizing the optics order comes down to two numbers: how many GPU/NIC ports you're populating, and the 1:2 switch-to-server ratio above.
| You're populating | Module needed | Ratio |
|---|---|---|
| ConnectX-7 / ConnectX-8 / BlueField-3 / DGX server ports | MMA4Z00-NS400 flat-top (RHS) | 1 per server port |
| Air-cooled Quantum-2 switch ports (QM9700/QM9790) | MMA4Z00-NS finned-top (IHS) | 1 per 2 server ports |
Both are new, coded specifically for their respective platforms, ship from Knoxville, Tennessee, and carry a lifetime warranty backed by us directly. Given where OEM lead times currently sit — see why 800G/400G optics are on 40+ week backorder in 2026 — sourcing the compatible pair is frequently the difference between a fabric that's live this month and one that's waiting on an allocation queue into next year.
FAQ
Do I need NDR InfiniBand for my AI cluster, or is Ethernet enough?
NDR earns its cost at scale: large distributed-training runs (roughly 512+ GPUs) with tight AllReduce synchronization benefit most from InfiniBand's lossless, low-latency design. Smaller clusters, inference-serving workloads, or teams converging onto existing Ethernet infrastructure can often run RoCE v2 on 400GbE instead — see our full InfiniBand vs. Ethernet breakdown.
Is MMA4Z00-NS/NS400 the right optic for a GB200 or Blackwell-generation build?
Check your reference architecture first. NDR (this part pair) is the standard for H100/H200 SuperPOD designs. NVIDIA's newer Blackwell-class scale-out reference architectures are moving to Quantum-X800/XDR at 800G per port with ConnectX-8, which uses a different transceiver family. If you're unsure which generation your build calls for, tell us your platform and we'll confirm before you order.
How many switch-side modules do I need for a given number of GPU servers?
Roughly half as many MMA4Z00-NS (switch-side) modules as MMA4Z00-NS400 (server-side) modules — one switch module's 800G breaks down into two independent 400G links, each feeding a different server.
Where does this fabric run in NVIDIA's own reference design?
NVIDIA's published DGX SuperPOD reference architecture for H100 and H200 systems specifies NDR InfiniBand — ConnectX-7 HCAs and Quantum-2 switches — as the scale-out fabric, in a rail-optimized topology built on this exact part pair.
Resilient Tec stocks and tests the MMA4Z00-NS400 and MMA4Z00-NS against ConnectX-7 and Quantum-2 specifically, ships from Knoxville, Tennessee, and backs both with a lifetime warranty.
Ready to Buy?
- 400G OSFP SR4 Flat-Top (RHS) — MMA4Z00-NS400 / 980-9I51S-00NS00 compatible — $579, the server/NIC-side module for ConnectX-7
- 800G OSFP SR8 Finned-Top (IHS) — MMA4Z00-NS compatible — $829, the switch-side module for Quantum-2
Sizing a full fabric build? Tell us the port count and we'll quote both sides together at the correct 1:2 ratio.
You Might Also Need
- InfiniBand vs. Ethernet for AI Clusters in 2026 — the fabric decision behind this build
- Is a third-party MMA4Z00-NS / NS400 safe? — what to verify before you order
- Flat-top vs finned-top OSFP — confirm which module your ports need
- Why NVIDIA optics have 40+ week lead times in 2026 — plan around current OEM supply
- Network switches — Arista, Cisco, HPE/Aruba and Mellanox, new and tested pre-owned