The key points
- Peak compute in AI hardware has grown much faster than memory bandwidth, so many workloads now wait on memory rather than on arithmetic.
- Generating tokens with a large language model does little arithmetic per byte read, which makes the decode phase bound by memory bandwidth on today's GPUs.
- The KV cache grows with context length and batch size and often competes with model weights for GPU memory capacity.
- The industry is responding with more and faster HBM, smaller number formats, attention variants that shrink the KV cache, and serving systems that split inference across different hardware.
What the memory wall is
The memory wall is the long-standing observation that processors get faster more quickly than memory can feed them. The idea predates AI accelerators, and it was one of the problems HBM was designed to address when AMD and SK hynix began working on stacked memory in the late 2000s. [1]
For AI hardware, the gap was quantified by Amir Gholami and colleagues in AI and Memory Wall, a paper published in IEEE Micro in 2024. Looking across about 20 years of server hardware, they found that peak compute grew about 3.0 times every two years, while DRAM bandwidth grew about 1.6 times and interconnect bandwidth about 1.4 times over the same interval. They argue that memory bandwidth can become the dominant bottleneck for decoder models, especially when serving them. [2]
Arithmetic intensity and the roofline
A useful way to reason about the problem is the roofline model, proposed by Samuel Williams, Andrew Waterman and David Patterson at Berkeley in 2008. It plots attainable performance against a program's arithmetic, or operational, intensity: the number of floating-point operations it performs for each byte it moves from memory. Below a certain intensity, performance is capped by memory bandwidth; above it, by peak compute. [3]
NVIDIA lists the H100 SXM at 1,979 teraFLOPS of BF16 tensor throughput with sparsity and 3.35 TB/s of memory bandwidth. Dividing one by the other gives roughly 590 operations per byte, or about half that for dense math. A kernel must therefore do several hundred operations for every byte it reads to keep the tensor cores busy. Large matrix multiplications with big batches can reach that; many inference steps cannot.
The H100 figures used in the example come from NVIDIA's own specification table, which notes that the headline tensor numbers assume structured sparsity. [4]
Why LLM token generation is memory-bound
Serving a large language model has two phases. In prefill, the model processes the whole prompt at once, which is dense matrix work and makes good use of compute. In decode, it generates one token at a time, and each step has to read the model's weights and the cached attention state from memory while doing comparatively little arithmetic. The researchers behind the Splitwise system found that the token-generation phase underuses a GPU's compute resources, which is consistent with it being limited by memory. [5][2]
A back-of-the-envelope bound shows the effect. A 70-billion-parameter model stored in 16-bit precision occupies about 140 GB. If generating each token for a single user requires reading all of those weights once, an H200 with 4.8 TB/s of bandwidth can produce at most around 34 tokens per second for that user, however many teraFLOPS it has. Batching many users together shares each weight read across more tokens, which is why serving throughput depends so heavily on how many requests fit in memory at once.
The KV cache: memory capacity becomes the constraint
Transformers keep the keys and values computed for earlier tokens so they do not have to recompute them at every step. This key-value, or KV, cache grows with every token of context and every concurrent request. The authors of vLLM, presented at SOSP 2023, described the KV cache for each request as large and dynamically growing and shrinking, and showed that managing it in fixed-size pages, like virtual memory, cut wasted memory to near zero and raised serving throughput two to four times compared with earlier systems. [6]
To see the scale, take a hypothetical model with 80 layers, 8 key-value heads of dimension 128, and a 16-bit cache. Each token stores 2 × 80 × 8 × 128 × 2 bytes, about 0.33 MB. A single 128,000-token context then needs around 42 GB of KV cache, before any other user is served and before counting the weights.
Model designers have responded by reducing the number of key-value heads. Multi-query attention uses a single key-value head, which speeds up decoding but can reduce quality. Grouped-query attention, published at EMNLP 2023, uses an intermediate number of key-value heads and, according to its authors, achieves quality close to standard multi-head attention at a speed comparable to multi-query attention. [7]
The hardware response: more, faster memory
Accelerator makers have spent much of each new generation's budget on memory. NVIDIA's A100 80GB offered about 2 TB/s; the H100 SXM 3.35 TB/s; the H200 4.8 TB/s with 141 GB; and each GPU in a DGX B200 has 8 TB/s. NVIDIA's Rubin GPU, ramping into production as of September 2026, lists 288 GB of HBM4 and 19.2 TB/s. [8][4][9][10][11]
Even that growth has not caught up with compute. Using the AI and Memory Wall paper's rates, compute roughly tripling every two years against bandwidth growing about 1.6 times, the ratio of FLOPS to bytes per second keeps rising, so the arithmetic intensity a kernel needs to stay compute-bound rises too. More HBM relieves the pressure without removing it.
The software response: fewer bytes per token
If bandwidth is the limit, reading fewer bytes is the most direct fix. The GPTQ method, presented at ICLR 2023, quantised weights to 3 or 4 bits with little loss of accuracy, fitted a 175-billion-parameter model on a single GPU for inference, and reported end-to-end speedups of about 3.25 times over FP16 on an A100. Hardware has followed: NVIDIA's NVFP4 format, supported on Blackwell GPUs, cuts the memory footprint by about 3.5 times relative to FP16 and 1.8 times relative to FP8, and NVIDIA reports accuracy losses of one percent or less on the models it tested. [12][13]
Better memory management helps too. PagedAttention, used in vLLM, lets many more requests share a GPU by removing fragmentation in the KV cache, and it allows cached blocks to be shared within and across requests rather than duplicated. [6]
Disaggregation: different hardware for different phases
Because prefill and decode stress different resources, researchers have proposed running them on different machines. Splitwise reported 1.4 times higher throughput at 20 percent lower cost, or 2.35 times more throughput for the same cost and power, by sending prompt processing and token generation to separately optimised pools. DistServe, presented at OSDI 2024, placed prefill and decoding on different GPUs and served up to 7.4 times more requests within latency targets than the systems it compared against. [5][14]
These ideas are now in commercial products. NVIDIA's open-source Dynamo inference framework splits prefill and decode across nodes and includes a KV block manager that moves cache data to cheaper tiers of memory and storage. NVIDIA's Rubin CPX, announced in September 2025 for the end of 2026, pairs 128 GB of GDDR7 with compute aimed at the context-processing phase, leaving HBM-equipped Rubin GPUs to handle generation. [15][16]
What the memory wall means for GPU buyers
For anyone renting GPUs for inference, the memory wall changes what to compare. Peak TFLOPS is a poor guide to tokens per second; memory bandwidth, memory capacity and the number formats a GPU supports usually matter more. An older GPU with a lot of HBM can outperform a newer one with less for long-context serving, and quantising a model can be worth more than a hardware upgrade. On Kovara you can compare GPUs by memory bandwidth and capacity alongside hourly price, look up a specific GPU's memory specifications, or ask Kova to estimate which GPU fits your model and context length.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- Korea JoongAng Daily · How HBM split the paths of Samsung, AMD and SK hynix ↗ (opens in a new tab)News report · Checked 29 September 2026
- arXiv · Gholami et al., AI and Memory Wall (IEEE Micro, 2024) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- UC Berkeley EECS · Williams, Waterman and Patterson, Roofline: An Insightful Visual Performance Model ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- NVIDIA · H100 Tensor Core GPU ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- arXiv · Patel et al., Splitwise: Efficient generative LLM inference using phase splitting ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- arXiv · Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- arXiv · Ainslie et al., GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (EMNLP 2023) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- NVIDIA · A100 Tensor Core GPU datasheet ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · H200 Tensor Core GPU ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · DGX B200 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · Vera Rubin NVL72 ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- arXiv · Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (ICLR 2023) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- NVIDIA Technical Blog · Introducing NVFP4 for Efficient and Accurate Low-Precision Inference ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- arXiv · Zhong et al., DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving (OSDI 2024) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- NVIDIA Developer · NVIDIA Dynamo ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Unveils Rubin CPX: A New Class of GPU Designed for Massive-Context Inference ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
