VRAM: the definition
VRAM is the common name for memory used by a GPU to hold working data. Its capacity, bandwidth, allocation behavior and physical relationship to other memory determine which workloads fit and how efficiently data can be accessed.
The key points
- Memory capacity is a fit constraint; bandwidth is a data-movement constraint.
- Model weights are not the entire memory budget.
- Unified addressing does not make all physical memory equally fast.
- Adding devices does not automatically create a single transparent memory pool.
What VRAM means in practice
Video random access memory is the historical expansion of VRAM. In contemporary GPU discussion, the term commonly describes the memory available to an accelerator for working data, even when no video is being rendered. Textures and frame buffers use it in graphics; model parameters, activations and caches use it in machine learning. The label identifies a role in the system, not one universal physical memory technology. [1][2]
A discrete GPU can have dedicated memory connected to its own controllers. Integrated and unified-memory systems may share a physical memory pool with other processors. Those architectures need to be interpreted on their own terms: a shared total must also cover the operating system and other applications. A specification that lists system memory is not automatically a specification for memory exclusively available to one GPU workload. [1][2]
Capacity, bandwidth and latency
Capacity is measured in bytes and determines how much state can be held. Bandwidth is measured in bytes per second across a defined interface. Latency is the delay associated with an access or transfer. A larger memory can solve an out-of-memory problem without proportionally increasing throughput. A faster interface can move an existing working set more quickly without making a larger working set fit. [3]
Imagine two hypothetical devices with equal 48 GB capacity but different sustained memory bandwidth. Both might fit the same model at low concurrency, but one may execute a memory-heavy kernel faster. Conversely, an 80 GB device with the same bandwidth as a 48 GB device could permit a larger cache or batch while leaving a particular small kernel's speed almost unchanged. Capacity and performance interact through the workload; they are not interchangeable specifications.
What occupies memory in an AI workload?
Model weights are only one category. Training can also require gradients, optimizer state, saved activations and temporary buffers. Inference typically avoids the normal training state but still needs runtime allocations and intermediate values. Frameworks may reserve blocks of memory for reuse, so memory reported as reserved can differ from the amount occupied by live tensors. Diagnosing a failure requires understanding the allocator as well as the model. [4][5]
For autoregressive language-model inference, the KV cache retains attention-related state from previous tokens. Its size depends on architecture, representation, sequence length and the number of active sequences. Different cache strategies can change allocation and memory use. The memory required to load a checkpoint therefore does not establish the memory required to serve a production distribution of requests. [6]
A practical budget separates persistent state, variable per-request state, temporary workspace and operating headroom. Record which components are measured and which are estimates. Do not add a fixed universal percentage and call it verified capacity: implementation choices can change the relationship substantially. Headroom is useful, but the appropriate margin should come from the observed workload and failure tolerance.
A transparent sizing calculation
Assume a hypothetical model has 12 billion parameters stored at two bytes each. Weights occupy approximately 24 billion bytes, or 24 GB in decimal units. If its intended serving workload adds 6 GB of KV cache and 4 GB of other working allocations, the modeled total becomes 34 GB before any extra margin. A device with 24 GB cannot satisfy that stated budget just because the weights alone appear to match its label.
For a conventional illustrative attention cache, let L be 32 layers, B be four active sequences, T be 4,096 cached tokens per sequence, H be eight KV heads, D be a head dimension of 128, and S be two bytes per element. The simplified cache size is 2 × L × B × T × H × D × S: 2,147,483,648 bytes, or 2 GiB. The leading two accounts for keys and values. This is not a universal model formula: windowing, compression, sharing and architecture-specific caches can change it.
Decimal GB and binary GiB are different units. One GB is one billion bytes; one GiB is 1,073,741,824 bytes. The same byte count can therefore be displayed as different-looking numbers. A discrepancy between a label and a tool's display may reflect units, reserved resources or another allocation boundary. Compare bytes and definitions before concluding that memory is missing. [9]
Quantization, batching and offloading
Quantization changes how some numerical values are represented. A smaller representation can reduce weight or cache storage, but metadata, scaling values and temporary dequantization buffers may remain. Weight-only quantization does not automatically reduce every activation or cache. Quality, supported kernels and hardware behavior must be evaluated together. A four-bit label is not proof that total application memory becomes exactly one quarter of a sixteen-bit configuration. [7]
Reducing batch size or concurrent sequences can lower variable memory use. Offloading can place selected data in host memory or another storage tier, but movement across slower paths can affect latency and throughput. Unified memory mechanisms can simplify addressing and migration; they do not abolish the physical cost of moving data between devices or memory tiers. [6][2]
For a hypothetical service, lowering concurrency from sixteen requests to eight might make the cache fit, while reducing the maximum served throughput. Offloading a layer might make a model runnable but too slow for an interactive target. These changes can be valuable, but they alter the service being purchased. State the new performance and concurrency assumptions instead of reporting only that the model now loads.
Why GPU memory does not simply add up
In data-parallel training, workers commonly hold model replicas while processing different data. More devices can increase aggregate processing capacity without removing a per-device model-fit requirement. Sharding and model-parallel strategies distribute different state, but introduce communication and placement requirements. The relevant memory total depends on the parallelization strategy, not just the sum of capacities on a purchase order. [8]
Consider a hypothetical model requiring 60 GB of local state. Two 40 GB devices do not solve the problem merely by being in the same server. A supported partition of the state may solve it, but replicated buffers and uneven layer sizes can still limit placement. Interconnect bandwidth then becomes part of the performance question. Memory fit and communication are coupled once a workload crosses device boundaries.
How to read a memory specification
Check whether capacity is per device, per hardware partition, per node or shared across the system. Identify the memory technology, published bandwidth basis and any partitioning. Then test with the exact model, precision, sequence distribution and concurrency. Record peak allocated and reserved memory during the whole workflow, including startup, warm-up and checkpoint operations where applicable.
The most common mistake is asking how much memory a model needs as though there were one permanent answer. The better question is how much memory this implementation needs for this workload at this quality and latency objective. A capacity figure is a starting constraint; the runtime experiment establishes whether the proposed configuration is usable.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1. Why can a model that fits at startup fail later?
Caches, activations, temporary allocations or concurrency may grow during execution. Startup weight storage is only part of peak memory demand.
2. Does weight quantization reduce the KV cache automatically?
No. The cache may use a different representation and strategy. Weight, activation and cache quantization are separate choices.
3. What is the difference between 24 GB and 24 GiB?
They use different units: 24 GB is 24 billion bytes; 24 GiB is 24 × 1,073,741,824 bytes. Compare the same unit before evaluating a budget.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- Intel · What is a GPU? ↗ (opens in a new tab)Manufacturer explainer · Checked 28 September 2026
- NVIDIA · CUDA programming model ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- NVIDIA · GPU Performance Background ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- PyTorch · Optimizing model parameters ↗ (opens in a new tab)Framework tutorial · Checked 28 September 2026
- PyTorch · CUDA memory management ↗ (opens in a new tab)Framework documentation · Checked 28 September 2026
- Hugging Face · Cache strategies ↗ (opens in a new tab)Framework documentation · Checked 28 September 2026
- Hugging Face · Quantization overview ↗ (opens in a new tab)Framework documentation · Checked 28 September 2026
- PyTorch · Distributed Data Parallel ↗ (opens in a new tab)Framework tutorial · Checked 28 September 2026
- NIST · Prefixes for binary multiples ↗ (opens in a new tab)Measurement reference · Checked 28 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.