128 GPU AI cluster network: GB200 NVL72
Two racks, 16 GPUs more than asked for, and the first build here that needs more than one switch. Whole rails share each leaf rather than taking one apiece.
- GPUs144
- Racks of GPUs2
- Compute switches8
- Fabric tiers2
- Total racks3
- Total draw253 kW
The build
| Platform | NVIDIA GB200 NVL72, 72 × GB200 Blackwell per rack |
|---|---|
| Compute fabric | infiniband, 72 rails at 400G |
| Leaf | 5 × QM9700 (Quantum-2) |
| Spine | 3 × QM9700 (Quantum-2) |
| Core | not required at this scale |
| Racks | 2 compute + 1 network, 1 rack per rack, limited by fixed |
| Power | 240 kW compute + 13 kW network = 253 kW |
How this fabric was sized.
Each rail needs only 2 of a leaf's 32 downlinks, so 16 whole rails share each leaf. That is 5 leaf switches instead of 72, and every rail is still single-hop because no rail is split across leaves. The cost is fault domain: losing one leaf now costs 16 rails rather than one. Set rail packing off to keep strict one-rail-per-leaf isolation.
Asked for 128, built 144. An NVL72 rack is indivisible, so GPU counts round to multiples of 72. 2 racks is the smallest build that meets the target.
Fabrics
Compute
144 endpoints across 72 rails, 2 tiers, realised oversubscription 1:1 (non-blocking).
| Tier | Switch | Count | Down/sw | Up/sw | Ports used |
|---|---|---|---|---|---|
| leaf | NVIDIA QM9700 (Quantum-2) | 5 | 32 | 32 | 304 / 320 |
| spine | NVIDIA QM9700 (Quantum-2) | 3 | 64 | 0 | 160 / 192 |
Storage
104 endpoints across 1 rail, 2 tiers, realised oversubscription 1.91:1.
| Tier | Switch | Count | Down/sw | Up/sw | Ports used |
|---|---|---|---|---|---|
| leaf | NVIDIA QM9700 (Quantum-2) | 3 | 42 | 22 | 170 / 192 |
| spine | NVIDIA QM9790 (Quantum-2, externally managed) | 2 | 64 | 0 | 66 / 128 |
In-band management
72 endpoints across 1 rail, 2 tiers, realised oversubscription 3.67:1.
| Tier | Switch | Count | Down/sw | Up/sw | Ports used |
|---|---|---|---|---|---|
| leaf | Cisco Nexus 93600CD-GX | 4 | 22 | 6 | 96 / 112 |
| spine | Cisco Nexus 93600CD-GX | 1 | 28 | 0 | 24 / 28 |
Out-of-band management
62 endpoints across 1 rail, 2 tiers, realised oversubscription 1:1 (non-blocking).
| Tier | Switch | Count | Down/sw | Up/sw | Ports used |
|---|---|---|---|---|---|
| leaf | Cisco Nexus 9348GC-FX3 | 2 | 48 | 4 | 62 / 96 |
| spine | Cisco Nexus 93180YC-FX3 | 2 | 48 | 0 | 8 / 96 |
Bill of materials
Every line carries the calculation that produced its quantity. Prices are deliberately absent: the catalog ships without list prices, because inventing them would be worse than leaving them out.
| Item | Qty | Why this quantity |
|---|---|---|
| Servers | ||
| GB200 NVL72 (72 x GB200 Blackwell) | 2 | 128 GPUs requested / 72 GPUs per node = 2 nodes (144 GPUs). |
| Network adapters | ||
NVIDIA ConnectX-7 400G (NDR)MCX75310AAS-NEAT |
144 | 2 nodes x 72 compute NICs per node (one per rail). |
| Compute fabric switches | ||
NVIDIA QM9700 (Quantum-2)MQM9700-NS2F |
11 | 11 across Compute fabric (rail-optimized) leaf, Compute fabric (rail-optimized) spine, Storage fabric leaf. |
NVIDIA QM9790 (Quantum-2, externally managed)MQM9790-NS2F |
2 | 2 across Storage fabric spine. |
| Management switches | ||
Cisco Nexus 93180YC-FX3N9K-C93180YC-FX3 |
2 | 2 across Out-of-band management spine. |
Cisco Nexus 9348GC-FX3N9K-C9348GC-FX3 |
2 | 2 across Out-of-band management leaf. |
Cisco Nexus 93600CD-GXN9K-C93600CD-GX |
5 | 5 across In-band management leaf, In-band management spine. |
| Optics — transceivers | ||
25G SFP28 SRSFP-25G-SR-S |
17 | 8 x 25G OOB switch to OOB aggregation at 30 m. 30 m run needs discrete optics over structured fiber. 25G SFP28 SR at OOB switch uplink and 25G SFP28 SR at OOB aggregation, over LC duplex OM4 multi-mode patch. |
800G OSFP 2xSR4 twin-port (NDR)MMA4Z00-NS |
233 | 160 x 400G Leaf to Spine at 30 m. 30 m run needs discrete optics over structured fiber. 800G OSFP 2xSR4 twin-port (NDR) at Compute leaf and 800G OSFP 2xSR4 twin-port (NDR) at Compute spine, over MPO-12 APC OM4 multi-mode trunk. The Compute leaf module is twin-port, so it terminates 2 links and only 0.5 module is needed per link. |
| Cables and assemblies | ||
| 100G QSFP28 passive DAC | 75 | 72 x 100G In-band NIC to In-band switch at 2 m. 2 m run is within the 2 m passive copper limit. Optics are integrated into the assembly, so no separate transceivers are needed. |
800G OSFP to 2x400G OSFP splitter DACMCP7Y00-N001 |
128 | 144 x 400G Compute NIC to Leaf at 2 m. 2 m run is within the 2 m passive copper limit. Optics are integrated into the assembly, so no separate transceivers are needed; one cage breaks out to 2 links. |
| Cat6a patch lead | 64 | 62 x 1G BMC / device management port to OOB switch at 2 m. 1G over structured copper, 2 m within the 100 m limit for Cat6a patch lead. |
| Structured fibre | ||
| LC duplex OM4 multi-mode patch | 9 | 8 x 25G OOB switch to OOB aggregation at 30 m. 30 m run needs discrete optics over structured fiber. 25G SFP28 SR at OOB switch uplink and 25G SFP28 SR at OOB aggregation, over LC duplex OM4 multi-mode patch. |
| MPO-12 APC OM4 multi-mode trunk | 233 | 160 x 400G Leaf to Spine at 30 m. 30 m run needs discrete optics over structured fiber. 800G OSFP 2xSR4 twin-port (NDR) at Compute leaf and 800G OSFP 2xSR4 twin-port (NDR) at Compute spine, over MPO-12 APC OM4 multi-mode trunk. The Compute leaf module is twin-port, so it terminates 2 links and only 0.5 module is needed per link. |
| Racks | ||
| Rack | 3 | 2 compute racks at 1 nodes each (fixed-limited) plus 1 network racks. |
| Power | ||
| Rack PDU | 6 | 3 racks x 2 PDUs per rack. |