If your DGX Sparks are cabled to a 400G switch through QSFP-DD breakout cables and the numbers look wrong, the ceiling you are hitting tells you the cause. No link at all means the switch port is still trying to auto-negotiate. About 10 Gb/s means the switch is forwarding in software or RoCE is stalling on flow control. About 100 Gb/s means you are driving only half of each Spark port. Operators who worked through all three report roughly 196–198 Gb/s per node pair and about 23–24 GB/s NCCL all-reduce bus bandwidth across four nodes. The cable is almost never the problem.
Short answer: it's configuration, not cabling
Three ceilings, three causes — and the cable that makes the build possible
A 400G QSFP-DD to 2×200G QSFP56 passive breakout feeds two Sparks from one switch port. Two cables cover a four-node cluster.
|
MCP7H60-W002 — 400G QSFP-DD to 2×200G QSFP56 breakout DAC, 2 m MCP7H60-W002R26 compatible. Two per four-Spark cluster. |
$179.45 · In stock at Resilient Tec Add 2 to cart |
New, lifetime warranty, free standard shipping on all US orders, shipping from Knoxville, Tennessee. Price checked September 28, 2026. Stuck on a build? Send us your switch model and symptoms.
Symptom, Cause, Fix
| What you see | Likely cause | Fix |
|---|---|---|
| Cable detected, "auto-negotiation failed", no link on the legs | QSFP56-DD port won't auto-negotiate 2×200G | Set speed manually on lanes 1 and 5 of each DD port |
| ~10 Gb/s on every transport; single-flow RoCE ~5 Gb/s | No hardware offload (unused breakout sub-ports in the bridge), PFC/trust stalling RoCE, duplicate fabric IPs | Clean up the bridge, disable PFC, give every rail interface a unique IP |
| ~100–110 Gb/s per port with a clean 200G link | Only one of the port's two 100G halves in use | Drive both logical interfaces; test with parallel streams |
| NCCL far below iperf3; logs show Socket transport | NCCL fell back from RDMA to TCP sockets | Remove settings forcing Socket; keep RoCE GID indices stable |
1. No Link: Set the Speed by Hand
MikroTik's QSFP56-DD ports on the CRS804 and CRS812 do not auto-negotiate a 2×200G split. The switch sees the breakout cable but leaves the legs down. Set the speed explicitly on the lane-master interface of each leg — lanes 1 and 5 of each DD port:
/interface ethernet
set qsfp56-dd-1-1 auto-negotiation=no speed=200G-baseCR4
set qsfp56-dd-1-5 auto-negotiation=no speed=200G-baseCR4
Repeat for every DD port carrying a breakout. The links should come up within about ten seconds. Other switch vendors have their own breakout commands, but the principle is the same: configure the split on the switch, and never on the Spark.
2. Stuck Near 10 Gb/s: Get the Switch Out of the Way
A July 2026 write-up on a four-node CRS804 build hit a hard ~10 Gb/s ceiling on every transport, with single-flow RoCE collapsing to about 5 Gb/s and NCCL all-reduce at 0.46 GB/s — while mlxlink showed clean links and the switch reported zero drops. It had three causes, and all three are easy to reproduce:
-
No hardware offload. Unused breakout sub-ports left in the management bridge stopped the switch ASIC from offloading, so the switch CPU forwarded every packet. The fix removed the QSFP56-DD interfaces from the bridge:
/interface/bridge/port/remove [find bridge=bridge interface~"qsfp56-dd"] -
PFC and trust stalling RoCE. Lossless-Ethernet settings stalled the data path. The fix turned them off on the fabric ports:
/interface/ethernet/switch/qos/port/set [find name~"qsfp56-dd"] trust-l3=ignore pfc=disabled, and on each Spark,mlnx_qos -i <interface> --pfc 0,0,0,0,0,0,0,0 -
Duplicate fabric IPs. Several interfaces sharing an address caused ARP confusion and loss. Give every rail interface its own IP, one subnet per rail, and set
arp_ignore,arp_announceandrp_filterto match.
After those three changes the same cluster ran about 109 Gb/s per rail uniformly across every node pair, with the switch CPU at 0%.
3. Stuck Near 100 Gb/s: Use Both Halves of the Port
This one surprises everyone. The GB10 chip gives each network device no more than a PCIe Gen5 x4 link, so NVIDIA runs the Spark's ConnectX-7 across two x4 links per port. Each physical 200G port is really two 100G halves, and Linux shows them as separate interfaces. One stream, or one interface, tops out near 100 Gb/s.
An operator on the NVIDIA developer forums running four Sparks on a CRS812 with 2×200G breakouts saw TCP capped at about 106 Gb/s on a 200G link. What got them to full speed:
- Treat each port as two logical halves and drive traffic on both at once
- MTU 9000 end to end, on the switch and the hosts
- Disable IPv6 on the ConnectX-7 fabric interfaces so RoCE GID indices stay consistent
- Remove leftover
/etc/nccl.confsettings that forced NCCL onto the Socket transport
Result: about 196–198 Gb/s aggregate between node pairs on parallel iperf3 sessions, and NCCL all-reduce around 23.8 GB/s bus bandwidth.
How to Test Properly
-
MTU first:
ping -M do -s 8972between every pair of nodes. If it fails, jumbo frames are broken somewhere in the path. -
Link health: use
ib_write_bw(RDMA) rather than single-stream iperf3, which will under-report on a Spark. - TCP: if you do use iperf3, run parallel streams across both logical interfaces.
-
End to end:
all_reduce_perffrom nccl-tests. Four tuned nodes should land in the low-20s GB/s bus bandwidth.
A note on where these numbers come from: they are results published by the operators cited below, not Resilient Tec bench tests. Your figures will vary with switch firmware, OS image and driver versions.
FAQ
Why is my DGX Spark only getting 100 Gb/s on a 200G link?
Each Spark QSFP port is two 100G halves, each on its own PCIe Gen5 x4 link, and Linux shows them as separate interfaces. A single stream or a single interface tops out near 100 Gb/s. Drive both halves in parallel to reach about 200 Gb/s.
Why won't my breakout cable link on a MikroTik QSFP56-DD port?
The port does not auto-negotiate 2x200G. Turn auto-negotiation off and set speed=200G-baseCR4 on the lane-master interfaces, lanes 1 and 5 of each DD port.
Why is my DGX Spark cluster stuck at 10 Gb/s?
Usually the switch is forwarding in software because unused breakout sub-ports sit in a bridge, PFC or trust settings are stalling RoCE, or fabric interfaces share IP addresses. Fix all three.
Is the breakout cable the bottleneck?
Rarely. A passive QSFP-DD to 2x200G QSFP56 breakout such as the MCP7H60-W002 carries a full 200G per leg; published four-node builds on these cables reach about 198 Gb/s per pair once the switch and hosts are configured.
What NCCL bandwidth should a four-node DGX Spark cluster get?
Published tuned builds report about 23–24 GB/s all-reduce bus bandwidth across four nodes.
Ready to Buy?
New, lifetime warranty backed by Resilient Tec, free standard shipping on all US orders, shipping from Knoxville, Tennessee. Price checked September 28, 2026.
- MCP7H60-W002 — 400G QSFP-DD to 2×200G QSFP56 breakout DAC, 2 m — $179.45; add two for a four-Spark cluster
You Might Also Need
- MikroTik CRS804 + DGX Spark — the port math and parts list
- MCP7H60-W001R30 vs W002R26 vs W003R26 — which breakout to buy
- Can You Build a 4-Node DGX Spark Cluster Without a Switch?
- Lossless Ethernet for RoCE