AI Network Designer Design a network →

1,024 GPU AI cluster network: DGX H100

Four scalable units: 32 leaf and 16 spine, the second published figure the engine reproduces. Still two tiers.

The build

PlatformNVIDIA DGX H100, 8 × H100 SXM5 per node
Compute fabricinfiniband, 8 rails at 400G
Leaf32 × QM9700 (Quantum-2)
Spine16 × QM9700 (Quantum-2)
Corenot required at this scale
Racks32 compute + 2 network, 4 nodes per rack, limited by fixed
Power1,306 kW compute + 53 kW network = 1,359 kW

Fabrics

Compute

1,024 endpoints across 8 rails, 2 tiers, realised oversubscription 1:1 (non-blocking).

TierSwitchCountDown/swUp/swPorts used
leaf NVIDIA QM9700 (Quantum-2) 32 32 32 2,048 / 2,048
spine NVIDIA QM9700 (Quantum-2) 16 64 0 1,024 / 1,024

Storage

288 endpoints across 1 rail, 2 tiers, realised oversubscription 1.91:1.

TierSwitchCountDown/swUp/swPorts used
leaf NVIDIA QM9700 (Quantum-2) 7 42 22 442 / 448
spine NVIDIA QM9790 (Quantum-2, externally managed) 3 64 0 154 / 192

In-band management

256 endpoints across 1 rail, 2 tiers, realised oversubscription 3.67:1.

TierSwitchCountDown/swUp/swPorts used
leaf Cisco Nexus 93600CD-GX 12 22 6 328 / 336
spine Cisco Nexus 93600CD-GX 3 28 0 72 / 84

Out-of-band management

275 endpoints across 1 rail, 2 tiers, realised oversubscription 1:1 (non-blocking).

TierSwitchCountDown/swUp/swPorts used
leaf Cisco Nexus 9348GC-FX3 6 48 4 275 / 288
spine Cisco Nexus 93180YC-FX3 2 48 0 24 / 96

What the engine flagged

UNEVEN_STRIPING — 22 uplinks per leaf do not divide evenly across 3 spines. Link load will be slightly uneven; consider adjusting the switch count or oversubscription.

Bill of materials

Every line carries the calculation that produced its quantity. Prices are deliberately absent: the catalog ships without list prices, because inventing them would be worse than leaving them out.

ItemQtyWhy this quantity
Servers
DGX H100 (8 x H100 SXM5) 128 1,024 GPUs requested / 8 GPUs per node = 128 nodes (1,024 GPUs).
Network adapters
NVIDIA ConnectX-7 400G (NDR)
MCX75310AAS-NEAT
1,024 128 nodes x 8 compute NICs per node (one per rail).
Compute fabric switches
NVIDIA QM9700 (Quantum-2)
MQM9700-NS2F
55 55 across Compute fabric (rail-optimized) leaf, Compute fabric (rail-optimized) spine, Storage fabric leaf.
NVIDIA QM9790 (Quantum-2, externally managed)
MQM9790-NS2F
3 3 across Storage fabric spine.
Management switches
Cisco Nexus 93180YC-FX3
N9K-C93180YC-FX3
2 2 across Out-of-band management spine.
Cisco Nexus 9348GC-FX3
N9K-C9348GC-FX3
6 6 across Out-of-band management leaf.
Cisco Nexus 93600CD-GX
N9K-C93600CD-GX
15 15 across In-band management leaf, In-band management spine.
Optics — transceivers
25G SFP28 SR
SFP-25G-SR-S
50 24 x 25G OOB switch to OOB aggregation at 30 m. 30 m run needs discrete optics over structured fiber. 25G SFP28 SR at OOB switch uplink and 25G SFP28 SR at OOB aggregation, over LC duplex OM4 multi-mode patch.
800G OSFP 2xSR4 twin-port (NDR)
MMA4Z00-NS
1,214 1024 x 400G Leaf to Spine at 30 m. 30 m run needs discrete optics over structured fiber. 800G OSFP 2xSR4 twin-port (NDR) at Compute leaf and 800G OSFP 2xSR4 twin-port (NDR) at Compute spine, over MPO-12 APC OM4 multi-mode trunk. The Compute leaf module is twin-port, so it terminates 2 links and only 0.5 module is needed per link.
Cables and assemblies
100G QSFP28 passive DAC 264 256 x 100G In-band NIC to In-band switch at 2 m. 2 m run is within the 2 m passive copper limit. Optics are integrated into the assembly, so no separate transceivers are needed.
800G OSFP to 2x400G OSFP splitter DAC
MCP7Y00-N001
676 1024 x 400G Compute NIC to Leaf at 2 m. 2 m run is within the 2 m passive copper limit. Optics are integrated into the assembly, so no separate transceivers are needed; one cage breaks out to 2 links.
Cat6a patch lead 284 275 x 1G BMC / device management port to OOB switch at 2 m. 1G over structured copper, 2 m within the 100 m limit for Cat6a patch lead.
Structured fibre
LC duplex OM4 multi-mode patch 25 24 x 25G OOB switch to OOB aggregation at 30 m. 30 m run needs discrete optics over structured fiber. 25G SFP28 SR at OOB switch uplink and 25G SFP28 SR at OOB aggregation, over LC duplex OM4 multi-mode patch.
MPO-12 APC OM4 multi-mode trunk 1,214 1024 x 400G Leaf to Spine at 30 m. 30 m run needs discrete optics over structured fiber. 800G OSFP 2xSR4 twin-port (NDR) at Compute leaf and 800G OSFP 2xSR4 twin-port (NDR) at Compute spine, over MPO-12 APC OM4 multi-mode trunk. The Compute leaf module is twin-port, so it terminates 2 links and only 0.5 module is needed per link.
Racks
Rack 34 32 compute racks at 4 nodes each (fixed-limited) plus 2 network racks.
Power
Rack PDU 68 34 racks x 2 PDUs per rack.

Related