512 GPU AI cluster network: GB200 NVL72
Eight racks, 576 GPUs, and roughly a megawatt of continuous draw before a single switch is powered. Each leaf still carries several whole rails.
- GPUs576
- Racks of GPUs8
- Compute switches27
- Fabric tiers2
- Total racks10
- Total draw999 kW
The build
| Platform | NVIDIA GB200 NVL72, 72 × GB200 Blackwell per rack |
|---|---|
| Compute fabric | infiniband, 72 rails at 400G |
| Leaf | 18 × QM9700 (Quantum-2) |
| Spine | 9 × QM9700 (Quantum-2) |
| Core | not required at this scale |
| Racks | 8 compute + 2 network, 1 rack per rack, limited by fixed |
| Power | 960 kW compute + 39 kW network = 999 kW |
How this fabric was sized.
Each rail needs only 8 of a leaf's 32 downlinks, so 4 whole rails share each leaf. That is 18 leaf switches instead of 72, and every rail is still single-hop because no rail is split across leaves. The cost is fault domain: losing one leaf now costs 4 rails rather than one. Set rail packing off to keep strict one-rail-per-leaf isolation.
Asked for 512, built 576. An NVL72 rack is indivisible, so GPU counts round to multiples of 72. 8 racks is the smallest build that meets the target.
Fabrics
Compute
576 endpoints across 72 rails, 2 tiers, realised oversubscription 1:1 (non-blocking).
| Tier | Switch | Count | Down/sw | Up/sw | Ports used |
|---|---|---|---|---|---|
| leaf | NVIDIA QM9700 (Quantum-2) | 18 | 32 | 32 | 1,152 / 1,152 |
| spine | NVIDIA QM9700 (Quantum-2) | 9 | 64 | 0 | 576 / 576 |
Storage
320 endpoints across 1 rail, 2 tiers, realised oversubscription 1.91:1.
| Tier | Switch | Count | Down/sw | Up/sw | Ports used |
|---|---|---|---|---|---|
| leaf | NVIDIA QM9700 (Quantum-2) | 8 | 42 | 22 | 496 / 512 |
| spine | NVIDIA QM9790 (Quantum-2, externally managed) | 3 | 64 | 0 | 176 / 192 |
In-band management
288 endpoints across 1 rail, 2 tiers, realised oversubscription 3.67:1.
| Tier | Switch | Count | Down/sw | Up/sw | Ports used |
|---|---|---|---|---|---|
| leaf | Cisco Nexus 93600CD-GX | 14 | 22 | 6 | 372 / 392 |
| spine | Cisco Nexus 93600CD-GX | 3 | 28 | 0 | 84 / 84 |
Out-of-band management
224 endpoints across 1 rail, 2 tiers, realised oversubscription 1:1 (non-blocking).
| Tier | Switch | Count | Down/sw | Up/sw | Ports used |
|---|---|---|---|---|---|
| leaf | Cisco Nexus 9348GC-FX3 | 5 | 48 | 4 | 224 / 240 |
| spine | Cisco Nexus 93180YC-FX3 | 2 | 48 | 0 | 20 / 96 |
What the engine flagged
UNEVEN_STRIPING — 22 uplinks per leaf do not divide evenly across 3 spines. Link load will be slightly uneven; consider adjusting the switch count or oversubscription.
Bill of materials
Every line carries the calculation that produced its quantity. Prices are deliberately absent: the catalog ships without list prices, because inventing them would be worse than leaving them out.
| Item | Qty | Why this quantity |
|---|---|---|
| Servers | ||
| GB200 NVL72 (72 x GB200 Blackwell) | 8 | 512 GPUs requested / 72 GPUs per node = 8 nodes (576 GPUs). |
| Network adapters | ||
NVIDIA ConnectX-7 400G (NDR)MCX75310AAS-NEAT |
576 | 8 nodes x 72 compute NICs per node (one per rail). |
| Compute fabric switches | ||
NVIDIA QM9700 (Quantum-2)MQM9700-NS2F |
35 | 35 across Compute fabric (rail-optimized) leaf, Compute fabric (rail-optimized) spine, Storage fabric leaf. |
NVIDIA QM9790 (Quantum-2, externally managed)MQM9790-NS2F |
3 | 3 across Storage fabric spine. |
| Management switches | ||
Cisco Nexus 93180YC-FX3N9K-C93180YC-FX3 |
2 | 2 across Out-of-band management spine. |
Cisco Nexus 9348GC-FX3N9K-C9348GC-FX3 |
5 | 5 across Out-of-band management leaf. |
Cisco Nexus 93600CD-GXN9K-C93600CD-GX |
17 | 17 across In-band management leaf, In-band management spine. |
| Optics — transceivers | ||
25G SFP28 SRSFP-25G-SR-S |
42 | 20 x 25G OOB switch to OOB aggregation at 30 m. 30 m run needs discrete optics over structured fiber. 25G SFP28 SR at OOB switch uplink and 25G SFP28 SR at OOB aggregation, over LC duplex OM4 multi-mode patch. |
800G OSFP 2xSR4 twin-port (NDR)MMA4Z00-NS |
775 | 576 x 400G Leaf to Spine at 30 m. 30 m run needs discrete optics over structured fiber. 800G OSFP 2xSR4 twin-port (NDR) at Compute leaf and 800G OSFP 2xSR4 twin-port (NDR) at Compute spine, over MPO-12 APC OM4 multi-mode trunk. The Compute leaf module is twin-port, so it terminates 2 links and only 0.5 module is needed per link. |
| Cables and assemblies | ||
| 100G QSFP28 passive DAC | 297 | 288 x 100G In-band NIC to In-band switch at 2 m. 2 m run is within the 2 m passive copper limit. Optics are integrated into the assembly, so no separate transceivers are needed. |
800G OSFP to 2x400G OSFP splitter DACMCP7Y00-N001 |
462 | 576 x 400G Compute NIC to Leaf at 2 m. 2 m run is within the 2 m passive copper limit. Optics are integrated into the assembly, so no separate transceivers are needed; one cage breaks out to 2 links. |
| Cat6a patch lead | 231 | 224 x 1G BMC / device management port to OOB switch at 2 m. 1G over structured copper, 2 m within the 100 m limit for Cat6a patch lead. |
| Structured fibre | ||
| LC duplex OM4 multi-mode patch | 21 | 20 x 25G OOB switch to OOB aggregation at 30 m. 30 m run needs discrete optics over structured fiber. 25G SFP28 SR at OOB switch uplink and 25G SFP28 SR at OOB aggregation, over LC duplex OM4 multi-mode patch. |
| MPO-12 APC OM4 multi-mode trunk | 775 | 576 x 400G Leaf to Spine at 30 m. 30 m run needs discrete optics over structured fiber. 800G OSFP 2xSR4 twin-port (NDR) at Compute leaf and 800G OSFP 2xSR4 twin-port (NDR) at Compute spine, over MPO-12 APC OM4 multi-mode trunk. The Compute leaf module is twin-port, so it terminates 2 links and only 0.5 module is needed per link. |
| Racks | ||
| Rack | 10 | 8 compute racks at 1 nodes each (fixed-limited) plus 2 network racks. |
| Power | ||
| Rack PDU | 20 | 10 racks x 2 PDUs per rack. |