NVLink & Scaling Up
Individually, Blackwell GPUs are fast. The real magic is making dozens of them behave as one. That job belongs to 5th-generation NVLink and the NVLink Switch — the interconnect that turns a rack of chips into a single giant GPU.
#Why interconnect is the bottleneck
A trillion-parameter model is too big for any single GPU, so it’s split across many (Module 8 covers how). Those GPUs must constantly swap data — and if the links between them are slow, the expensive compute sits idle. The ordinary PCIe bus simply isn’t fast enough.
#5th-generation NVLink
NVLink is NVIDIA’s dedicated GPU-to-GPU link, and the 5th generation doubles the previous (Hopper) generation. Each link carries 50 GB/s in each direction; with 18 links per GPU that totals 1.8 TB/s of bandwidth (900 GB/s each way) — over 14× a PCIe Gen5 connection.
#The NVLink Switch: 72 GPUs as one
NVLink links GPUs; the NVLink Switch chip ties them into a single high-bandwidth domain. In a 72-GPU domain (NVL72) it moves 130 TB/s of aggregate bandwidth, and adds 4× efficiency with SHARP FP8 in-network reduction. The fabric scales up to 576 GPUs, and lets the GB300 NVL72 reach 9× the throughput of a single eight-GPU system.
The 130 TB/s moving inside one 72-GPU NVLink domain is more data movement than the entire internet handles at once.