The key points
- Each HBM generation has raised bandwidth per stack mainly by increasing the per-pin data rate, and HBM4 doubled the interface to 2,048 bits.
- JEDEC sets the baseline for each generation, but memory makers routinely ship parts that run faster than the published standard.
- Data-center GPUs moved from HBM2 in the A100 40GB to HBM3E in the H200, B200 and MI355X, with HBM4 arriving in the Rubin generation.
- A GPU's total memory bandwidth is roughly the number of stacks multiplied by the speed each stack is run at, which is often below the maximum a memory maker advertises.
What actually changes from one HBM generation to the next
High Bandwidth Memory stacks several DRAM dies on top of one another, links them with through-silicon vias, and exposes a very wide interface to the processor sitting beside it. The first JEDEC HBM standard, JESD235, was published in October 2013, after AMD and SK hynix had worked on the technology together and brought it to the standards body. SK hynix describes itself as having released the first HBM product that same year. [1][8]
Every generation since has pulled on the same four levers: how fast each data pin runs, how many pins the stack exposes, how many DRAM dies are stacked, and how dense each die is. The JEDEC documents for HBM2, HBM3 and HBM4 show all four moving. HBM2 used a 1,024-bit interface split into 8 channels; HBM3 doubled the per-pin rate and moved to 16 channels, and its quoted 819 GB/s at 6.4 Gb/s per pin corresponds to the same 1,024-bit width; HBM4 doubled the interface to 2,048 bits and 32 channels. [3][5][6]
Bandwidth per stack follows from simple arithmetic: interface width in bits, multiplied by the per-pin rate in gigabits per second, divided by eight to convert to bytes. A 1,024-bit HBM3 stack at 6.4 Gb/s per pin gives 819 GB/s, the figure JEDEC quotes. A 2,048-bit HBM4 stack at 8 Gb/s gives roughly 2 TB/s, again matching the standard.
HBM and HBM2: from graphics cards to the first AI accelerators
First-generation HBM appeared in a consumer product in 2015, when AMD launched the Radeon R9 Fury X. Each four-high stack held 1 GB and delivered 128 GB/s, and the card combined four stacks for 4 GB and 512 GB/s. The memory was fast but expensive, and the product did not change the graphics market much. [1][2]
JEDEC published the HBM2 update, JESD235A, in January 2016. It kept the 1,024-bit interface and 8 independent channels, introduced a pseudo-channel mode to improve effective bandwidth, supported 2-high, 4-high and 8-high stacks, and allowed up to 8 GB and 256 GB/s per stack. Samsung began mass-producing HBM2 in January 2016, and in April 2016 NVIDIA launched the Tesla P100 using Samsung HBM2, one of the first data-center accelerators built around stacked memory. [3][2]
HBM2E: taller stacks and the A100
HBM2E is the industry name for the extended HBM2 generation. JEDEC's JESD235B update of December 2018 raised the per-pin rate to 2.4 Gb/s, for up to 307 GB/s per stack, and added 12-high stacks with capacities of up to 24 GB. Memory makers then pushed further: in July 2020 SK hynix began mass production of HBM2E running at 3.6 Gb/s per pin, which it said delivered more than 460 GB/s from an eight-high, 16 GB stack. [4][7]
NVIDIA's A100 straddles the two versions. The 40 GB A100 uses HBM2 and is rated at 1,555 GB/s, while the later 80 GB model moved to HBM2e, reaching 1,935 GB/s in the PCIe card and 2,039 GB/s in the SXM module. That is a useful reminder that the same GPU name can hide different memory generations, with a meaningful gap in bandwidth. [13]
HBM3: double the data rate and the H100 and MI300X
JEDEC published HBM3 as JESD238 in January 2022. It doubled the per-pin data rate compared with HBM2 to as much as 6.4 Gb/s, or 819 GB/s per stack, and doubled the number of independent channels from 8 to 16. It supported 4-high, 8-high and 12-high stacks with a provision for 16-high, dies from 8 Gb to 32 Gb, and capacities from 4 GB to 64 GB per stack. It also added on-die error correction and lowered signalling and core voltages. [5]
The H100 is the best-known HBM3 product. NVIDIA lists the H100 SXM with 80 GB of HBM3 at 3.35 TB/s, and the H100 NVL with 94 GB at 3.9 TB/s. AMD's Instinct MI300X also uses HBM3, arranged as eight stacks for 192 GB of capacity and a peak of 5.3 TB/s, which gave it considerably more memory per accelerator than the H100 at launch. [14][18][19]
HBM3E: the workhorse of the Hopper refresh and Blackwell
HBM3E is an extended, faster form of HBM3 that memory makers began shipping in 2024. In February 2024 Micron announced volume production of an eight-high, 24 GB HBM3E part with pin speeds above 9.2 Gb/s and more than 1.2 TB/s per stack, and said it would appear in NVIDIA's H200. In September 2024 SK hynix began mass-producing 12-high HBM3E with 36 GB per stack at 9.6 Gb/s; to keep the stack the same height as the eight-high version it thinned each DRAM die by 40 percent. [10][8]
NVIDIA describes the H200 as the first GPU with HBM3e, carrying 141 GB at 4.8 TB/s, roughly 1.4 times the H100's bandwidth. Its Blackwell-based DGX B200 system lists 1,440 GB of HBM3e across eight GPUs with 64 TB/s of aggregate bandwidth, which works out to 180 GB and 8 TB/s per GPU. AMD's MI350X and MI355X carry 288 GB of HBM3E at 8 TB/s, with memory sourced from Micron and Samsung. [15][16][20]
HBM4 and HBM4E: a wider interface for the Rubin and MI400 generation
JEDEC released HBM4 as JESD270-4 in April 2025. The standard doubles the interface to 2,048 bits, allows up to 8 Gb/s per pin for as much as 2 TB/s per stack, doubles the independent channels to 32, supports 4-high to 16-high stacks with 24 Gb or 32 Gb dies, and tops out at 64 GB per stack. It also lets a single controller work with either HBM3 or HBM4. [6]
Shipping parts again run well above the baseline. SK hynix said in September 2025 that its finished HBM4 design exceeded 10 Gb/s per pin. Samsung announced commercial HBM4 shipments in February 2026 at a consistent 11.7 Gb/s, or 3.3 TB/s per stack, in 12-high stacks of 24 GB to 36 GB, and in May 2026 began sampling HBM4E at 14 Gb/s and up to 3.6 TB/s per stack, with a 48 GB 12-high version. [9][11][12]
How to read HBM figures on a GPU spec sheet
A GPU's memory bandwidth is roughly the number of HBM stacks multiplied by the bandwidth each stack is run at. The MI300X shows how this works: eight HBM3 stacks at a little over 650 GB/s each give the 5.3 TB/s AMD quotes, below the 819 GB/s per-stack ceiling of the HBM3 standard. Chip designers often run memory below its maximum to manage power, heat and signal integrity across a large package.
Capacity figures deserve the same care. Sold capacity can be lower than the raw capacity of the stacks, as the H200's 141 GB and DGX B200's 180 GB per GPU suggest, and one product name can span two memory generations, as with the 40 GB and 80 GB A100. When comparing offers, the useful questions are which memory generation the exact SKU uses, how much of it is usable, and what bandwidth the vendor actually specifies.
What HBM generations mean for GPU buyers today
In the rental market, memory generation is a practical proxy for what a GPU can do with large models. HBM2e-era A100s remain useful and cheap for smaller models, HBM3 H100s are the default for many teams, and HBM3E parts such as the H200, B200 and MI355X are chosen when a model or a long context needs more memory per GPU. HBM4 systems are only starting to appear, and early capacity is likely to be scarce and expensive. On Kovara you can compare hourly prices across these generations, open a GPU's page to check its exact memory type, capacity and bandwidth, or ask Kova which memory configuration fits your model.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- KitGuru · AMD started to work on HBM technology nearly a decade ago ↗ (opens in a new tab)News report · Checked 29 September 2026
- Korea JoongAng Daily · How HBM split the paths of Samsung, AMD and SK hynix ↗ (opens in a new tab)News report · Checked 29 September 2026
- JEDEC · JEDEC Updates Groundbreaking High Bandwidth Memory (HBM) Standard (JESD235A) ↗ (opens in a new tab)Standards body · Checked 29 September 2026
- TechPowerUp · JEDEC Updates Groundbreaking High Bandwidth Memory (HBM) Standard (JESD235B) ↗ (opens in a new tab)News report · Checked 29 September 2026
- JEDEC · JEDEC Publishes HBM3 Update to High Bandwidth Memory (HBM) Standard ↗ (opens in a new tab)Standards body · Checked 29 September 2026
- JEDEC · JEDEC and Industry Leaders Collaborate to Release JESD270-4 HBM4 Standard ↗ (opens in a new tab)Standards body · Checked 29 September 2026
- SK hynix Newsroom · SK hynix starts mass production of high-speed DRAM HBM2E ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- SK hynix Newsroom · SK hynix Begins Volume Production of the World's First 12-Layer HBM3E ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- SK hynix via PR Newswire · SK hynix Completes World's First HBM4 Development and Readies Mass Production ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- HPCwire · Micron Commences Volume Production of HBM3E Solution to Accelerate the Growth of AI ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- Samsung Newsroom · Samsung Ships Industry-First Commercial HBM4 With Ultimate Performance for AI Computing ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- Samsung Newsroom · Samsung Electronics Begins Shipment of Industry-First HBM4E Samples ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · A100 Tensor Core GPU datasheet ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · H100 Tensor Core GPU ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · H200 Tensor Core GPU ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · DGX B200 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · Vera Rubin NVL72 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD · AMD Instinct MI300X Accelerators ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- Hot Chips 2024 · AMD Instinct MI300X Generative AI Accelerator and Platform Architecture ↗ (opens in a new tab)Conference presentation · Checked 29 September 2026
- AMD · AMD Instinct MI350 Series and Beyond: Accelerating the Future of AI and HPC ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
