NVIDIA H100: the definition
The NVIDIA H100 is a data center GPU based on the Hopper architecture, introduced in 2022, that added FP8 Tensor Core math and a Transformer Engine for AI workloads. It ships as an SXM module with 80 GB of HBM3, a PCIe card with 80 GB, and the PCIe-based H100 NVL with 94 GB.
The key points
- The H100 is NVIDIA's Hopper-generation data center GPU, with 80 billion transistors on TSMC 4N and shipping in systems from October 2022.
- Its fourth-generation Tensor Cores and Transformer Engine introduced FP8, roughly doubling peak throughput over 16-bit math.
- The SXM version is the high-end configuration (80 GB HBM3, 3.35 TB/s, 700 W, 900 GB/s NVLink); PCIe and NVL versions trade performance for easier deployment.
- Wide cloud availability from 2023 onward made the H100 the baseline against which newer AI GPUs are measured.
What the H100 is
The H100 is the data center GPU built on NVIDIA's Hopper architecture, the successor to the Ampere-based A100. The full GH100 chip contains 80 billion transistors on an 814 square millimetre die, made on a version of TSMC's 4N process customised for NVIDIA. The full die has 144 streaming multiprocessors (SMs); shipping products enable fewer to improve yield. [2]
NVIDIA announced on September 20, 2022 that Hopper was in full production. Partners including Dell Technologies, HPE, Lenovo and Supermicro planned to roll out the first H100-based products from October 2022, with systems expected to ship in the following weeks, and AWS, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure planned H100 instances starting in 2023. [4]
SXM, PCIe and NVL versions
H100 SXM is a module that mounts on an HGX baseboard with four or eight GPUs. It enables 132 SMs, carries 80 GB of HBM3 with 3.35 TB/s of bandwidth, can be configured up to 700 W and connects to the other GPUs on the board through 900 GB/s of NVLink. It is the version behind most large H100 training clusters. [1][2]
H100 PCIe is a full-height, full-length, dual-slot card with a passive heatsink that relies on server airflow. It enables 114 SMs and uses 80 GB of HBM2e with about 2,000 GB/s of bandwidth, at a default and maximum power of 350 W. Pairs of cards can be joined with NVLink bridges for up to 600 GB/s. The lower power limit, fewer SMs and slower memory make it substantially slower than the SXM version. [3][2]
H100 NVL, introduced in March 2023, is a PCIe card aimed at large language model inference in mainstream servers. It has 94 GB of memory with 3.9 TB/s of bandwidth, runs at 350 to 400 W, and supports 600 GB/s of NVLink between bridged cards. At launch NVIDIA claimed up to 12 times faster GPT-3 inference than the prior-generation A100 at data center scale, a vendor figure rather than a per-card comparison. [1][5]
When a cloud listing just says H100, it usually means the SXM version in an eight-GPU HGX server. PCIe and NVL instances exist but are typically labelled as such. Because the performance gap between SXM and PCIe is large, it is worth checking which one a provider is offering before comparing prices.
Fourth-generation Tensor Cores, FP8 and the Transformer Engine
Hopper's fourth-generation Tensor Cores deliver twice the matrix math rate of the A100 per SM on the same data types, and four times the rate when using the new FP8 format. NVIDIA supports two FP8 variants: E4M3, with four exponent and three mantissa bits, which favours precision, and E5M2, with five exponent and two mantissa bits, which favours range. [2]
The Transformer Engine is the combination of these Tensor Cores with NVIDIA software that chooses, layer by layer, between FP8 and 16-bit precision and handles the scaling and casting between the two formats automatically. The aim is to get most of FP8's speed without the accuracy problems of applying it blindly to every layer. [2]
NVIDIA rates the H100 SXM at 3,958 teraFLOPS of FP8 and 1,979 teraFLOPS of FP16 or BF16, but these Tensor Core figures assume structured sparsity. Dense throughput is half: about 2 petaFLOPS of FP8 and about 1 petaFLOPS of BF16. For comparison, the A100 80GB is rated at 312 teraFLOPS of dense BF16. [1][8][6]
A team fine-tuning a model in BF16 on A100 GPUs that moves to H100 SXM can expect the largest speedup if it also adopts FP8 through the Transformer Engine libraries. Without FP8 the gain comes mainly from higher BF16 throughput and faster memory, which is still large but smaller than the headline figures.
Memory, NVLink 4 and other features
The H100 SXM uses 80 GB of HBM3 with 3.35 TB/s of bandwidth, up from 2,039 GB/s on the A100 80GB SXM. Hopper also has a 50 MB L2 cache and up to 228 KB of configurable shared memory per SM. New programming features include thread block clusters, distributed shared memory that lets SMs read each other's shared memory directly, and the Tensor Memory Accelerator for moving large blocks of data between global and shared memory. [1][2][6]
Fourth-generation NVLink gives each H100 SXM 900 GB/s of GPU-to-GPU bandwidth over 18 links, compared with 600 GB/s on the A100. Inside an HGX H100 board, NVLink 4 switches connect eight GPUs at 7.2 TB/s of aggregate bandwidth. NVIDIA also described an external NVLink Switch System able to connect up to 256 H100 GPUs, although most H100 clusters scale beyond eight GPUs over InfiniBand or Ethernet. [7][2][6]
Other Hopper additions include second-generation Multi-Instance GPU, which can split an H100 SXM into up to seven isolated instances of 10 GB each, confidential computing support, PCIe Gen5 with 128 GB/s of total bandwidth, and DPX instructions that NVIDIA says accelerate dynamic programming algorithms by up to 7 times over the A100. [1][2]
Why the H100 became the reference AI GPU
The H100 reached the major cloud providers from 2023. It arrived with FP8 and the Transformer Engine designed specifically around transformer models, a large increase in bandwidth over the A100, and NVIDIA's CUDA software stack already in place. NVIDIA and others continue to use it as the comparison point: NVIDIA's Blackwell materials, for example, express memory and performance gains relative to the H100. [4][8]
In our view the H100 became the reference point less because of any single specification than because of timing and ubiquity. It was the GPU most AI teams could actually obtain during the years when most of today's training and inference software was written and tuned, so kernels, libraries, benchmarks and cluster designs all default to it.
The H100 in 2026
Two newer lines now sit above the H100. The H200 keeps the Hopper compute design but moves to 141 GB of HBM3E with 4.8 TB/s. The Blackwell B200 offers up to 192 GB of HBM3E at 8 TB/s, 1.8 TB/s NVLink and FP4 support, at up to 1,200 W. See Kovara's H100 vs H200 comparison for the memory difference in detail. [8]
As of September 2026 the H100 is no longer NVIDIA's newest GPU, but it is far from obsolete. For models that fit in 80 GB per GPU, or that shard cleanly across eight GPUs, it often remains a practical choice, particularly where its mature software support matters more than peak throughput. Its main limitation is memory: very large models or long contexts may need more GPUs on H100 than on an H200 or B200.
Renting an H100 today
H100 capacity is available as single GPUs, full eight-GPU HGX nodes and larger InfiniBand clusters, on demand, as spot capacity and on reserved terms. Prices differ widely between providers and between SXM and PCIe versions, and they have moved considerably since 2023. On Kovara you can open the H100 SXM page to see current listings, check the GPU prices view to compare providers, and use Compare to weigh an H100 against an H200 or B200 for the model you plan to run.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1What is the difference between H100 SXM and H100 PCIe?
The SXM version is a module for HGX baseboards with 132 streaming multiprocessors, 80 GB of HBM3 at 3.35 TB/s, up to 700 W and 900 GB/s NVLink. The PCIe card has 114 streaming multiprocessors, 80 GB of HBM2e at 2 TB/s, a 350 W limit and optional NVLink bridges, so it is noticeably slower but fits standard servers.
2What is the H100 NVL?
A dual-slot, air-cooled PCIe version with 94 GB of memory and 3.9 TB/s of bandwidth, running at 350 to 400 W, with 600 GB/s NVLink between bridged cards. NVIDIA introduced it in March 2023 for large language model inference in standard servers.
3What does the Transformer Engine do?
It combines FP8 Tensor Cores with software that decides, per layer, whether to compute in FP8 or 16-bit precision and manages the scaling and conversion between them, so models can use FP8 without manual precision tuning.
4Is the H100 still worth using in 2026?
For many workloads, yes. Newer GPUs such as the H200 and Blackwell B200 offer more memory and throughput, but the H100 runs the same CUDA software stack and is sufficient for many models that fit in 80 GB per GPU or shard cleanly across eight GPUs. Whether it is the cheapest option per unit of work depends on the model and current pricing.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · H100 Tensor Core GPU ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Hopper Architecture In-Depth ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA · H100 Tensor Core GPU PCIe Product Brief ↗ (opens in a new tab)Datasheet · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Hopper in Full Production ↗ (opens in a new tab)Press release · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Launches Inference Platforms for Large Language Models and Generative AI Workloads ↗ (opens in a new tab)Press release · Checked 29 September 2026
- NVIDIA · A100 Tensor Core GPU ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · NVLink and NVLink Switch ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
