NVIDIA Blackwell: the definition
NVIDIA Blackwell is the data center GPU architecture that succeeded Hopper, built from two large dies joined into one GPU and designed around low-precision AI formats such as FP4. It covers the B200 and B300 (Blackwell Ultra) GPUs and the GB200 and GB300 Grace Blackwell superchips.
The key points
- Blackwell GPUs combine two reticle-sized dies with 208 billion transistors in total, joined by a 10 TB/s die-to-die link.
- The second-generation Transformer Engine adds FP4 with fine-grained scaling, doubling peak throughput over FP8 on the same chip.
- B200 carries 192 GB of HBM3E and B300 (Blackwell Ultra) 288 GB, both at 8 TB/s, with NVLink 5 at 1.8 TB/s per GPU.
- GB200 and GB300 pair two Blackwell GPUs with one Grace CPU and are the building block of the 72-GPU NVL72 rack.
What Blackwell is
NVIDIA announced the Blackwell platform on March 18, 2024 as the successor to the Hopper generation that powers the H100 and H200. The launch covered the B200 Tensor Core GPU, the GB200 superchip that attaches two of those GPUs to a Grace CPU, and a fifth generation of NVLink. Among the first cloud providers named were AWS, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure, alongside specialist GPU clouds such as CoreWeave, Crusoe, Lambda and Nebius. [3][2]
Blackwell is a family rather than a single chip. The original Blackwell GPU ships as the B200 in eight-GPU HGX and DGX servers and inside GB200 superchips. Blackwell Ultra, sold as the B300 and inside GB300 superchips, keeps the same basic design but adds memory and low-precision compute. NVIDIA positions both mainly for large language model training and, increasingly, for high-volume inference and reasoning workloads. [2][4]
The dual-die design
The most visible architectural change from Hopper is packaging. A Hopper GPU is a single die with 80 billion transistors made on TSMC 4N. A Blackwell GPU contains 208 billion transistors split across two dies, each close to the reticle limit, the largest area a lithography tool can expose in one shot. Both dies are made on TSMC 4NP, a customised variant of the same process family. [1][2]
The two dies are joined by a die-to-die interface NVIDIA calls NV-HBI, which provides 10 TB/s of bandwidth. NVIDIA describes the result as a unified single GPU: the pair is sold, scheduled and programmed as one device rather than as two separate accelerators. Blackwell Ultra uses the same two-die arrangement; in its full form it exposes 160 streaming multiprocessors and 640 fifth-generation Tensor Cores. [2][1]
The practical reason for going to two dies is that a single die could not grow much further. Moving to a multi-die GPU was the way to more than double transistor count within one process generation, at the cost of a more complex package and a much higher power budget per GPU.
FP4 and the second-generation Transformer Engine
Hopper already supported FP8 Tensor Core math. Blackwell's second-generation Transformer Engine extends this to 4-bit floating point. It relies on what NVIDIA calls micro-tensor scaling: instead of one scale factor for a whole tensor, small blocks of values each get their own scale, which keeps the limited range of 4-bit numbers usable. NVIDIA says the Tensor Cores also add new community-defined microscaling formats. [1][3][2]
NVIDIA's own 4-bit format, NVFP4, applies two levels of scaling: an FP8 (E4M3) scale for every block of 16 values plus an FP32 scale for the tensor. NVIDIA says this gives accuracy close to FP8 while cutting memory use by about 1.8 times compared with FP8. Whether a given model keeps acceptable quality at FP4 still has to be tested case by case. [2]
The throughput gain is large on paper. NVIDIA lists Hopper at 2 petaFLOPS of dense FP8. Blackwell reaches 5 petaFLOPS of dense FP8 and 10 petaFLOPS of dense NVFP4, and Blackwell Ultra raises dense NVFP4 to 15 petaFLOPS. Sparse figures are twice as high for FP8, but they assume structured sparsity in the model, so the dense numbers are the more conservative guide. [2]
Consider a model with 70 billion parameters. Stored at FP16 its weights need roughly 140 GB, at FP8 roughly 70 GB and at a 4-bit format roughly 35 GB plus a small overhead for scale factors. On a 192 GB B200 the 4-bit version leaves most of the memory free for the key-value cache that grows with batch size and context length, which is often what limits inference throughput.
B200 and B300: memory and power
The standard Blackwell GPU carries up to 192 GB of HBM3E with 8 TB/s of bandwidth and a maximum power of up to 1,200 W. Blackwell Ultra moves to eight 12-high HBM3E stacks for 288 GB, keeps bandwidth at 8 TB/s, and allows up to 1,400 W. NVIDIA also doubled the rate of the special function units that compute exponentials in attention layers, from 5 to 10.7 tera-exponentials per second, to speed up long-context attention. [2]
Usable memory in shipping systems can be slightly lower than the chip maximum. NVIDIA's DGX B200 lists 1,440 GB of GPU memory across eight GPUs, which works out to 180 GB per GPU, with 64 TB/s of aggregate HBM3E bandwidth. The same system is rated at about 14.3 kW maximum and uses two Intel Xeon Platinum 8570 host CPUs. On NVIDIA's HGX platform page, an eight-GPU HGX B200 has 1.4 TB of total memory and an HGX B300 has 2.1 TB. [5][4]
Networking also changes between the two. NVIDIA lists 0.8 TB/s of networking bandwidth for an HGX B200 system and 1.6 TB/s for HGX B300, in line with the move from the 400 Gb/s ConnectX-7 adapters listed for DGX B200 to the 800 Gb/s per GPU ConnectX-8 SuperNICs that NVIDIA lists for Blackwell Ultra systems such as GB300 NVL72. Eight-GPU FP8 throughput is the same 72 petaFLOPS with sparsity (36 dense) on both, while dense FP4 rises from 72 to 108 petaFLOPS. [4][5][7]
GB200 and GB300 Grace Blackwell superchips
The GB200 superchip combines one Grace CPU with two Blackwell GPUs on a single board, linked by NVLink chip-to-chip (NVLink-C2C) at 900 GB/s. Grace is NVIDIA's Arm-based server processor with 72 Arm Neoverse V2 cores and LPDDR5X memory. Because the CPU and GPUs are connected coherently, the GPUs can reach the CPU's memory at far higher bandwidth than over PCIe. [3][6]
NVIDIA quotes each GB200 superchip at 372 GB of HBM3E with 16 TB/s of bandwidth across its two GPUs, 3.6 TB/s of NVLink bandwidth and 20 petaFLOPS of dense NVFP4. GB300 follows the same pattern with Blackwell Ultra GPUs, and 36 GB300 superchips form a GB300 NVL72 rack rated at about 1.1 exaFLOPS of dense FP4. [6][2]
Superchips are not sold as stand-alone cards. They are built into rack-scale systems such as the GB200 NVL72 and GB300 NVL72, which connect 72 GPUs in one NVLink domain. For most buyers and renters, choosing GB200 or GB300 therefore means choosing a whole rack or a slice of one, rather than a single server.
NVLink 5 and scale-up
Fifth-generation NVLink doubles per-GPU bandwidth to 1,800 GB/s over 18 links, compared with 900 GB/s for NVLink 4 on Hopper. NVIDIA describes this as about 14 times the bandwidth of PCIe Gen5. In an eight-GPU HGX system the total NVLink bandwidth is 14.4 TB/s. [9][8][4]
The bigger change is the size of the domain. Hopper's NVLink 4 Switch connected eight GPUs. The NVLink 5 Switch supports eight-GPU domains and 72-GPU domains, reaching 130 TB/s of aggregate bandwidth in an NVL72 rack, and NVIDIA says the architecture can scale up to 576 GPUs in a single NVLink domain. A larger domain lets tensor-parallel and expert-parallel work run across many more GPUs without falling back to slower InfiniBand or Ethernet. [9][1][8]
Other additions: decompression, RAS and confidential computing
Blackwell adds a hardware decompression engine that handles formats such as LZ4, Snappy and Deflate, aimed at database and data-analytics pipelines, with 900 GB/s of bidirectional bandwidth to a Grace CPU. A dedicated RAS engine monitors hardware and software health to predict faults, which matters in large clusters where a single failing GPU can stall a training job. NVIDIA also describes Blackwell as the first GPU capable of TEE-I/O confidential computing, with throughput close to unencrypted modes. [1]
How Blackwell differs from Hopper
Putting NVIDIA's figures side by side: Hopper is one die with 80 billion transistors, up to 80 GB of HBM on the H100 or 141 GB of HBM3E on the H200, 3.35 to 4.8 TB/s of memory bandwidth, 900 GB/s NVLink and up to 700 W. Blackwell is two dies with 208 billion transistors, 192 to 288 GB of HBM3E at 8 TB/s, 1.8 TB/s NVLink, FP4 support and 1,200 to 1,400 W. NVIDIA says Blackwell Ultra has 3.6 times the on-package memory of H100. [2]
NVIDIA's headline system claims compare a GB200 NVL72 rack with the same number of H100 GPUs: up to 30 times faster real-time inference on a 1.8-trillion-parameter mixture-of-experts model and up to 4 times faster training. These are vendor benchmarks that depend on FP4, the larger NVLink domain and specific model configurations; per-GPU gains on smaller models in eight-GPU servers are much more modest. [6][8]
For most teams the differences that matter day to day are memory per GPU, which decides how many GPUs a model needs, and power per GPU, which decides whether a facility can host the hardware at all. Blackwell roughly doubles memory bandwidth and more than doubles capacity per GPU relative to the H100, but it also pushes many deployments toward liquid cooling.
Where Blackwell stands in 2026
On May 31, 2026, NVIDIA said its next platform, Vera Rubin, was ramping into full production, with production shipments set to begin in the fall of 2026. Vera Rubin uses sixth-generation NVLink at 3,000 GB/s per GPU. That makes Blackwell and Blackwell Ultra the current mainstream generation as of September 2026, with Rubin systems only beginning to arrive. [10][9]
Renting Blackwell today
Blackwell capacity is offered in three main shapes: single B200 or B300 GPUs carved from HGX servers, full eight-GPU nodes, and GB200 or GB300 NVL72 racks or partitions of them. Availability and pricing vary widely between providers and change quickly as B300 and Rubin capacity comes online. On Kovara you can open the B200 SXM 180GB, B300 SXM and GB200 superchip pages to see current listings, use the GPU prices view to compare hourly rates, and use Compare to set a Blackwell option against an H100 or H200 for the workload you plan to run.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1What is the difference between B200 and B300?
Both use the same two-die, 208-billion-transistor design and 1.8 TB/s NVLink 5. The B300 (Blackwell Ultra) raises HBM3E capacity from 192 GB to 288 GB, raises dense NVFP4 throughput from 10 to 15 petaFLOPS, roughly doubles attention-layer exponential throughput, and has a higher power ceiling of up to 1,400 W versus 1,200 W, according to NVIDIA.
2What is the difference between B200 and GB200?
B200 is the GPU itself, usually installed eight at a time in an HGX or DGX server with x86 host CPUs. GB200 is a superchip that joins two Blackwell GPUs to one Arm-based Grace CPU over a 900 GB/s NVLink chip-to-chip link; 36 of them make up a GB200 NVL72 rack.
3Why does FP4 matter on Blackwell?
Four-bit weights and activations take roughly half the memory of FP8 and let the Tensor Cores run at about twice the FP8 rate. NVIDIA uses fine-grained block scaling, such as its NVFP4 format, to keep accuracy close to FP8 for many inference workloads.
4Is Blackwell still NVIDIA's newest architecture?
As of September 2026, Blackwell and Blackwell Ultra are the most widely deployed current NVIDIA data center parts, but NVIDIA said in May 2026 that its successor, Vera Rubin, was ramping into full production with shipments starting in the fall of 2026.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · Blackwell Architecture ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Blackwell Platform Arrives to Power a New Era of Computing ↗ (opens in a new tab)Press release · Checked 29 September 2026
- NVIDIA · HGX Platform ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · DGX B200 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · GB200 NVL72 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · GB300 NVL72 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA GB200 NVL72 Delivers Trillion-Parameter LLM Training and Real-Time Inference ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA · NVLink and NVLink Switch ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Vera Rubin Ramps Into Full Production ↗ (opens in a new tab)Press release · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
