AI Network Designer Design a network →

128 GPU AI cluster network: GB200 NVL72

Two racks, 16 GPUs more than asked for, and the first build here that needs more than one switch. Whole rails share each leaf rather than taking one apiece.

The build

PlatformNVIDIA GB200 NVL72, 72 × GB200 Blackwell per rack
Compute fabricinfiniband, 72 rails at 400G
Leaf5 × QM9700 (Quantum-2)
Spine3 × QM9700 (Quantum-2)
Corenot required at this scale
Racks2 compute + 1 network, 1 rack per rack, limited by fixed
Power240 kW compute + 13 kW network = 253 kW

How this fabric was sized.

Each rail needs only 2 of a leaf's 32 downlinks, so 16 whole rails share each leaf. That is 5 leaf switches instead of 72, and every rail is still single-hop because no rail is split across leaves. The cost is fault domain: losing one leaf now costs 16 rails rather than one. Set rail packing off to keep strict one-rail-per-leaf isolation.

Asked for 128, built 144. An NVL72 rack is indivisible, so GPU counts round to multiples of 72. 2 racks is the smallest build that meets the target.

Fabrics

Compute

144 endpoints across 72 rails, 2 tiers, realised oversubscription 1:1 (non-blocking).

TierSwitchCountDown/swUp/swPorts used
leaf NVIDIA QM9700 (Quantum-2) 5 32 32 304 / 320
spine NVIDIA QM9700 (Quantum-2) 3 64 0 160 / 192

Storage

104 endpoints across 1 rail, 2 tiers, realised oversubscription 1.91:1.

TierSwitchCountDown/swUp/swPorts used
leaf NVIDIA QM9700 (Quantum-2) 3 42 22 170 / 192
spine NVIDIA QM9790 (Quantum-2, externally managed) 2 64 0 66 / 128

In-band management

72 endpoints across 1 rail, 2 tiers, realised oversubscription 3.67:1.

TierSwitchCountDown/swUp/swPorts used
leaf Cisco Nexus 93600CD-GX 4 22 6 96 / 112
spine Cisco Nexus 93600CD-GX 1 28 0 24 / 28

Out-of-band management

62 endpoints across 1 rail, 2 tiers, realised oversubscription 1:1 (non-blocking).

TierSwitchCountDown/swUp/swPorts used
leaf Cisco Nexus 9348GC-FX3 2 48 4 62 / 96
spine Cisco Nexus 93180YC-FX3 2 48 0 8 / 96

Bill of materials

Every line carries the calculation that produced its quantity. Prices are deliberately absent: the catalog ships without list prices, because inventing them would be worse than leaving them out.

ItemQtyWhy this quantity
Servers
GB200 NVL72 (72 x GB200 Blackwell) 2 128 GPUs requested / 72 GPUs per node = 2 nodes (144 GPUs).
Network adapters
NVIDIA ConnectX-7 400G (NDR)
MCX75310AAS-NEAT
144 2 nodes x 72 compute NICs per node (one per rail).
Compute fabric switches
NVIDIA QM9700 (Quantum-2)
MQM9700-NS2F
11 11 across Compute fabric (rail-optimized) leaf, Compute fabric (rail-optimized) spine, Storage fabric leaf.
NVIDIA QM9790 (Quantum-2, externally managed)
MQM9790-NS2F
2 2 across Storage fabric spine.
Management switches
Cisco Nexus 93180YC-FX3
N9K-C93180YC-FX3
2 2 across Out-of-band management spine.
Cisco Nexus 9348GC-FX3
N9K-C9348GC-FX3
2 2 across Out-of-band management leaf.
Cisco Nexus 93600CD-GX
N9K-C93600CD-GX
5 5 across In-band management leaf, In-band management spine.
Optics — transceivers
25G SFP28 SR
SFP-25G-SR-S
17 8 x 25G OOB switch to OOB aggregation at 30 m. 30 m run needs discrete optics over structured fiber. 25G SFP28 SR at OOB switch uplink and 25G SFP28 SR at OOB aggregation, over LC duplex OM4 multi-mode patch.
800G OSFP 2xSR4 twin-port (NDR)
MMA4Z00-NS
233 160 x 400G Leaf to Spine at 30 m. 30 m run needs discrete optics over structured fiber. 800G OSFP 2xSR4 twin-port (NDR) at Compute leaf and 800G OSFP 2xSR4 twin-port (NDR) at Compute spine, over MPO-12 APC OM4 multi-mode trunk. The Compute leaf module is twin-port, so it terminates 2 links and only 0.5 module is needed per link.
Cables and assemblies
100G QSFP28 passive DAC 75 72 x 100G In-band NIC to In-band switch at 2 m. 2 m run is within the 2 m passive copper limit. Optics are integrated into the assembly, so no separate transceivers are needed.
800G OSFP to 2x400G OSFP splitter DAC
MCP7Y00-N001
128 144 x 400G Compute NIC to Leaf at 2 m. 2 m run is within the 2 m passive copper limit. Optics are integrated into the assembly, so no separate transceivers are needed; one cage breaks out to 2 links.
Cat6a patch lead 64 62 x 1G BMC / device management port to OOB switch at 2 m. 1G over structured copper, 2 m within the 100 m limit for Cat6a patch lead.
Structured fibre
LC duplex OM4 multi-mode patch 9 8 x 25G OOB switch to OOB aggregation at 30 m. 30 m run needs discrete optics over structured fiber. 25G SFP28 SR at OOB switch uplink and 25G SFP28 SR at OOB aggregation, over LC duplex OM4 multi-mode patch.
MPO-12 APC OM4 multi-mode trunk 233 160 x 400G Leaf to Spine at 30 m. 30 m run needs discrete optics over structured fiber. 800G OSFP 2xSR4 twin-port (NDR) at Compute leaf and 800G OSFP 2xSR4 twin-port (NDR) at Compute spine, over MPO-12 APC OM4 multi-mode trunk. The Compute leaf module is twin-port, so it terminates 2 links and only 0.5 module is needed per link.
Racks
Rack 3 2 compute racks at 1 nodes each (fixed-limited) plus 1 network racks.
Power
Rack PDU 6 3 racks x 2 PDUs per rack.

Related