The key points
- Capacity determines what can fit; bandwidth affects how quickly data can move.
- Model weights are only one part of the memory budget.
- Evaluate the whole serving configuration, not a headline speedup.
Three budgets, not one
GPU discussions often begin with peak floating-point performance. A more useful first question is what will limit the workload: arithmetic, memory traffic, or the space needed to hold its state. NVIDIA’s performance guide distinguishes compute-limited and memory-limited operations using arithmetic intensity: work performed relative to the data that must move. A high arithmetic rating cannot, by itself, remove a memory bottleneck. [1]
Capacity and bandwidth are different. Capacity is the amount of information memory can hold. Bandwidth is the rate at which information can move to and from that memory. NVIDIA lists the H200 with 141 GB of GPU memory and 4.8 TB/s of memory bandwidth. These describe different properties; neither is a measured tokens-per-second result for your application. [2]
Our interpretation is that these should be separate purchasing questions. A configuration may have enough capacity but inadequate throughput, or excellent theoretical throughput but insufficient room for the intended workload. Treating “more memory” as a single dimension hides that distinction.
The weights are only the beginning
Consider a hypothetical dense model with 70 billion parameters, each stored in a two-byte format. The weights alone occupy about 140 billion bytes, or 140 GB in decimal units. This is multiplication, not a claim about a particular model implementation. It excludes temporary buffers, cached attention state, allocator overhead and, for training, additional state. [3]
That is why comparing 140 GB of calculated weights with a 141 GB specification does not establish that the application will fit on one device. The usable budget must cover the runtime, not just the file containing the parameters. Quantization or distributing the workload may change the answer, but each produces a new configuration to validate rather than a free capacity increase. [3]
Long conversations have an infrastructure cost
During autoregressive inference, a key-value cache retains attention information from earlier tokens so it does not all need to be recomputed. The cache occupies memory alongside model weights. Longer sequences and more concurrent requests increase its footprint; the precise amount depends on the model architecture and representation. Prefill processes the prompt, while decode generates subsequent tokens. Those phases can stress resources differently. [3]
A newer NVIDIA discussion of KV-cache quantization illustrates the trade-off: a smaller representation reduces the stored cache and memory traffic, but the implementation also has to consider accuracy and the serving workload. A result from one model and benchmark is not a general quality or throughput guarantee. [4]
For a product team, the useful question becomes “How many conversations can this configuration serve within our latency target?” rather than “Can it load the model?” A prototype with short prompts and one user does not define the intended production requirement.
Buy the constraint you actually need to remove
Before comparing providers, specify the model version, precision, prompt and output lengths, concurrency and latency objective. Then ask for a test matching that configuration. Compare acceptable completed work, not a throughput figure achieved by allowing a response to become unacceptably slow.
Also distinguish GPU memory from system RAM and storage. NVIDIA’s architecture guide describes a memory hierarchy, not interchangeable buckets of equally fast capacity. A listing with generous CPU RAM is not evidence of an equivalent quantity of local GPU memory. [1]
The conclusion is not that memory always matters more than arithmetic. It is that capacity, bandwidth and computation form a set of constraints. Establish which constraint limits the job, then evaluate the hardware and software combination that removes it. This article is an architectural explanation, not an independent hardware benchmark.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · GPU Performance Background User’s Guide ↗ (opens in a new tab)Manufacturer documentation · Checked 27 September 2026
- NVIDIA · H200 specifications ↗ (opens in a new tab)Manufacturer specifications · Checked 27 September 2026
- NVIDIA · Mastering LLM Techniques: Inference Optimization ↗ (opens in a new tab)Manufacturer technical article · Checked 27 September 2026
- NVIDIA · Optimizing Inference with NVFP4 KV Cache ↗ (opens in a new tab)Manufacturer technical article · Checked 27 September 2026
Prepared with AI assistance. Publication approved by Tommaso Luci. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.