AI networking reference
The topology rules and the hardware data behind the designer, written out. Everything here is what the tool itself calculates with.
The arithmetic behind an AI fabric is not complicated, but it is unforgiving, and most of the ways it goes wrong are easy to miss: a leaf tier sized as a flat pool, a bond counted as one link, an optic that fits the cage but not the polish. Each page below covers one of them.
Designing the fabric
-
Rail-optimized network design
How rail-optimized fabrics work, why each rail must be sized independently, and the switch-count error that follows from treating the fabric as one flat pool of endpoints.
-
Fat-tree sizing and tier boundaries
The endpoint capacity of two-tier and three-tier fat trees, where the tier boundary actually falls in a rail-optimized fabric, and how port breakout moves it.
-
DGX SuperPOD switch counts
The leaf and spine switch counts published in NVIDIA DGX SuperPOD reference architectures for H100, B200 and B300, and the scalable-unit structure behind them.
-
GB200 and GB300 NVL72 networking
An NVL72 rack carries three separate networks: NVLink inside the rack, ConnectX scale-out between racks, and BlueField-3 for storage and management. Which ones you buy switches for, and how GB200 and GB300 trays differ.
-
InfiniBand against Spectrum-X Ethernet
How choosing InfiniBand or Spectrum-X Ethernet for an AI cluster changes the port arrangement, the link count, the plane structure, and therefore the bill of materials.
-
Optics compatibility for AI fabrics
The six things that must agree before an optical link comes up: form factor, vendor coding, reach, fibre type, fibre connector and ferrule polish — plus heatsink form, which is the seventh.
Hardware specifications
-
Switch specifications
Port counts, cage counts, connector types, breakout, power and rack units for the NVIDIA Quantum, Spectrum and Cisco Nexus switches used in AI cluster fabrics.
-
Server specifications
GPU counts, compute rail counts, storage and management port counts and power draw for NVIDIA DGX, GB200 and GB300 NVL72, Cisco UCS, Supermicro, Dell, HPE and Lenovo AI servers.
Specifications are encoded from public vendor documentation and each entry carries its source and a confidence badge. Anything not confirmed against a current datasheet says so. Check the vendor datasheet before quoting.