Blackwell, Decoded
The Systems · Module 7

GB200 & HGX Servers

The wider Blackwell family8 min readModule 7 of 8

The GB300 is the flagship, but Blackwell ships in a family of systems. The GB200 NVL72 is its predecessor (still a giant leap), and the air-cooled HGX B300 / B200 servers bring Blackwell to standard x86 data centers. Here’s how they compare and where each fits.

#GB200 NVL72 — the predecessor that still stuns

Same rack-scale idea as the GB300: 72 Blackwell GPUs + 36 Grace CPUs in one liquid-cooled NVLink domain. Against Hopper it delivers up to 40× AI-factory output, and on a trillion-parameter GPT-MoE-1.8T model it is transformative:

HGX H1003.5GB200 NVL72116
Relative throughput — about a 30× speedup vs H100, with 25× lower TCO and 25× less energy on the same GPU count.

It’s not just inference. The faster Transformer Engine also gives 4× faster training on GPT-MoE-1.8T versus Hopper, while shrinking the footprint dramatically:

faster LLM training
less rack space
3.5×
lower TCO & energy
30×
faster inference vs H100

Beyond AI: data & simulation

With its tightly-coupled CPU+GPU and the Decompression Engine, GB200 speeds database queries 18× over CPU (7× less energy, 5× lower TCO). In engineering simulation, Cadence SpectreX runs 13× faster than x86, and the Cadence Fidelity CFD solver up to 22× faster.

Sustainable by design

Replacing 65 racks of air-cooled HGX H100 with a single liquid-cooled GB200 NVL72 at equal performance cuts both energy use and TCO by 25× — while using less water for cooling and tolerating warmer ambient air.

#HGX B300 & B200 — Blackwell for x86 servers

Not every data center wants a Grace-CPU rack. The HGX line puts eight Blackwell GPUs on a baseboard that drops into standard x86 infrastructure.

HGX B300

Built for AI reasoning: more AI compute than Hopper, over 2 TB of HBM3e, and ConnectX-8 networking. For training, agentic systems, reasoning and real-time video.

HGX B200

An eight-GPU x86 platform delivering 144 petaFLOPS of AI — 15× the performance and 12× the TCO of HGX H100. Up to 1,000 W per GPU.

Table 3 — HGX B300 vs HGX B200 (8-GPU server)
SpecHGX B300HGX B200
GPU configuration16 Ultra dies / 8 GPUs16 dies / 8 GPUs
FP4 Tensor Core (dense / sparse)108 / 144 PF72 / 144 PF
FP8 / FP6 (dense / sparse)36 / 72 PF36 / 72 PF
Fast memory2.1 TBup to 1.4 TB
Aggregate memory bandwidth62 TB/sup to 62 TB/s
GPU memory (per GPU)270 GB · 7.7 TB/s≤192 GB · 7.7 TB/s
Max TDP (per GPU)1,100 W1,000 W
InterconnectNVLink 5 · PCIe Gen6NVLink 5 · PCIe Gen5

#The software that ties it together

Hardware is half the story. NVIDIA AI Enterprise is the end-to-end software platform — including NIM inference microservices, frameworks and libraries — certified to run on NVIDIA-Certified Systems, giving enterprises the security, support and stability to go from pilot to production.

Continue →