NVL72: the definition
NVL72 is NVIDIA's rack-scale system design in which 72 Blackwell GPUs and 36 Grace CPUs are connected by NVLink into a single high-bandwidth domain. GB200 NVL72 uses Blackwell GPUs and GB300 NVL72 uses Blackwell Ultra GPUs.
The key points
- An NVL72 rack links 72 Blackwell GPUs and 36 Grace CPUs into one NVLink domain with 130 TB/s of aggregate bandwidth.
- The rack is built from 18 compute trays and 9 NVLink switch trays joined by a copper cable spine.
- It is liquid-cooled and draws on the order of 120 to 135 kW, far beyond what most legacy data center racks were built for.
- The large NVLink domain is aimed at models, especially mixture-of-experts models, that are too big to run efficiently inside an eight-GPU server.
What NVL72 means
GB200 NVL72 is a complete rack from NVIDIA's Blackwell generation that contains 72 Blackwell GPUs and 36 Grace CPUs. The GPUs and CPUs are packaged as 36 GB200 superchips, each with one Grace CPU and two GPUs. NVL72 refers to the fact that all 72 GPUs sit in one NVLink domain, so any GPU can exchange data with any other GPU over NVLink rather than over a slower network. [1][3]
In a conventional GPU server, NVLink stops at the edge of the chassis. Hopper-generation HGX systems connect eight GPUs through NVLink 4 switches; anything beyond those eight has to travel over InfiniBand or Ethernet. The NVLink 5 switch generation supports both eight-GPU domains and 72-GPU domains, which is what makes NVL72 possible. [5]
Inside the rack
The rack contains 18 compute trays built on NVIDIA's MGX reference design. Each tray holds two Grace CPUs and four Blackwell GPUs, about 1.7 TB of fast memory and roughly 80 petaFLOPS of AI performance by NVIDIA's count. Nine NVLink switch trays sit between them, each providing 144 NVLink ports at 100 GB/s, so that the nine switch trays fully connect all 18 NVLink ports on every one of the 72 GPUs. [3]
The compute and switch trays are joined by a copper cable spine at the back of the rack. The Register reported that the NVLink backbone uses more than two miles of copper cabling and that NVIDIA chose copper over optical links partly because optics would have added around 20 kW of power. The same report put the rack weight at about 1.36 metric tons. [7]
Supermicro's version of the system lists 18 one-unit compute trays, nine NVLink switch trays, eight one-unit 33 kW power shelves and outer dimensions of 2,236 by 600 by 1,068 mm. Other server makers build their own variants, but the tray counts follow NVIDIA's reference design. [8]
One NVLink domain: memory and bandwidth
NVIDIA lists 130 TB/s of total NVLink bandwidth for the rack. The GPUs together carry about 13.4 TB of HBM3E with 576 TB/s of aggregate memory bandwidth, and the Grace CPUs add about 17 TB of LPDDR5X at 14 TB/s. Because the CPUs connect to their GPUs coherently over NVLink-C2C, NVIDIA describes the rack as offering about 30 TB of unified memory that applications can address. [1][3]
In compute terms, NVIDIA rates GB200 NVL72 at 720 petaFLOPS of dense FP4 (1,440 with sparsity), 720 petaFLOPS of FP8 and 360 petaFLOPS of FP16 or BF16, with the FP8 and FP16 figures quoted with sparsity. The CPU side totals 2,592 Arm Neoverse V2 cores. NVIDIA also says NVLink 5 can extend a single domain to as many as 576 GPUs by linking racks. [1][3]
GB200 NVL72 versus GB300 NVL72
GB300 NVL72 keeps the same shape, 72 GPUs, 36 Grace CPUs and 130 TB/s of NVLink, but uses Blackwell Ultra GPUs. Each Blackwell Ultra GPU has up to 288 GB of HBM3E instead of 192 GB, so NVIDIA lists about 20 TB of GPU memory and 37 TB of total fast memory per rack. Dense FP4 compute rises to about 1.1 exaFLOPS, while FP8 and FP16 figures stay the same as GB200 NVL72. [2][4]
The other visible change is scale-out networking. GB300 NVL72 uses ConnectX-8 SuperNICs that provide 800 Gb/s per GPU and connect to either Quantum-X800 InfiniBand or Spectrum-X Ethernet. That matters when many NVL72 racks are combined into one training cluster, since traffic between racks still leaves the NVLink domain. [2]
Power per GPU also rises. NVIDIA lists Blackwell at up to 1,200 W and Blackwell Ultra at up to 1,400 W, so GB300 racks generally need at least as much power and cooling headroom as GB200 racks. [4]
Power and liquid cooling
An NVL72 rack is a very dense load. The Register reported about 120 kW of compute for NVIDIA's DGX GB200 NVL72 and about 2,700 W for each GB200 superchip. Supermicro's datasheet lists an operating power of 125 to 135 kW per rack. For comparison, NVIDIA rates an air-cooled eight-GPU DGX B200 server at about 14.3 kW maximum. [7][8][6]
At that density the rack is liquid-cooled throughout. Cold plates on the CPUs, GPUs and NVLink switches carry heat into a coolant loop; The Register described a flow of about 2 liters per second with coolant entering at 25 degrees Celsius and leaving at 45 degrees, while small fans still cool lower-power parts. Heat leaves the rack through a coolant distribution unit. Supermicro offers an in-rack unit rated at 250 kW, an in-row unit rated at 1.3 MW and liquid-to-air options for sites without facility water. [7][8]
NVIDIA argues that liquid-cooled GB200 NVL72 racks reduce a data center's energy consumption and increase compute density. The facility side is the hard part: floors have to carry more than a ton per rack, electrical distribution has to deliver well over 100 kW to a single position, and water or heat-rejection capacity has to be planned in. This is why NVL72 capacity tends to appear first in new or heavily retrofitted AI data centers. [1][7]
Why rack-scale matters for large models
NVIDIA explains that mixture-of-experts models spread their computation across many experts and are trained across thousands of GPUs using model parallelism and pipeline parallelism. Splitting a model this way means GPUs exchange data constantly, so the speed of the links between them matters. Inside an NVL72 rack that traffic can stay on NVLink, with all 18 NVLink ports of each of the 72 GPUs connected through the switch trays; in eight-GPU servers, a model that needs more GPUs has to send part of that traffic over the scale-out network instead. [3]
NVIDIA's own benchmark on a 1.8-trillion-parameter mixture-of-experts model illustrates the point. It reports about 150 tokens per second per GPU on GB200 NVL72 against about 3.4 on H100, a 30-fold difference for real-time inference, and 4 times faster training at a scale of 32,000 GPUs. These results combine several changes at once, including FP4, faster GPUs and the larger NVLink domain, and they come from the vendor. [3][1]
A team serving a large mixture-of-experts model with hundreds of billions of parameters might need more memory than one eight-GPU server holds. On eight-GPU servers it would spread the model across two or more nodes connected by InfiniBand. On an NVL72 rack the same model can fit inside one NVLink domain with room left for key-value cache, so expert routing stays on NVLink.
The rack-scale approach mostly pays off for very large models, long contexts and high-concurrency serving. For fine-tuning or serving models that fit comfortably on one to eight GPUs, a standard HGX B200 or H200 node is usually simpler to obtain and to operate, and renting a full NVL72 would leave much of the rack idle.
What comes after NVL72 on Blackwell
NVIDIA's next rack, Vera Rubin NVL72, keeps the 72-GPU domain but moves to sixth-generation NVLink at 3,000 GB/s per GPU and 216 TB/s per rack. On May 31, 2026, NVIDIA said Vera Rubin was ramping into full production and that production shipments would begin in the fall of 2026. As of September 2026, GB200 and GB300 NVL72 are therefore the rack-scale systems most likely to be available from cloud providers. [5][9]
Renting NVL72 capacity
Providers usually sell NVL72 capacity as full racks on reserved contracts, as smaller partitions of a rack, or as managed clusters built from many racks. The unit of sale, contract length and network setup differ more than they do for eight-GPU servers, so it is worth checking exactly which GPUs, how much of the NVLink domain and which interconnect you are getting. On Kovara, the GB200 superchip page shows current listings, the GPU prices view lets you compare rates across providers, and Compare helps you judge whether a rack-scale system or conventional B200 or H200 nodes fit your workload better.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1What does NVL72 stand for?
It refers to 72 GPUs connected through NVLink. Every GPU in the rack can talk to every other GPU at 1.8 TB/s through NVLink switch trays, giving 130 TB/s of total NVLink bandwidth.
2How much power does a GB200 NVL72 rack use?
Roughly 120 kW or more. The Register reported about 120 kW of compute for NVIDIA's rack, and Supermicro's GB200 NVL72 datasheet lists an operating range of 125 to 135 kW with 132 kW of installed power shelves.
3Can a GB200 NVL72 be air-cooled?
The compute and switch trays are liquid-cooled by design. A facility without chilled water can use liquid-to-air heat exchangers, which Supermicro offers, but the rack itself still circulates liquid through cold plates.
4What is the difference between GB200 NVL72 and GB300 NVL72?
Both have 72 GPUs, 36 Grace CPUs and 130 TB/s of NVLink bandwidth. GB300 NVL72 uses Blackwell Ultra GPUs, raising GPU memory from about 13.4 TB to about 20 TB and dense FP4 compute from 720 to about 1,080 petaFLOPS, and it uses 800 Gb/s ConnectX-8 networking per GPU.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · GB200 NVL72 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · GB300 NVL72 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA GB200 NVL72 Delivers Trillion-Parameter LLM Training and Real-Time Inference ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA Technical Blog · Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA · NVLink and NVLink Switch ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · DGX B200 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- The Register · A closer look at Nvidia's 120kW DGX GB200 NVL72 rack system ↗ (opens in a new tab)News report · Checked 29 September 2026
- Supermicro · SuperCluster GB200 NVL72 datasheet ↗ (opens in a new tab)Datasheet · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Vera Rubin Ramps Into Full Production ↗ (opens in a new tab)Press release · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
