GPU: the definition
A graphics processing unit is a programmable processor designed to execute large amounts of parallel work. It accelerates suitable parts of a program; it is not a replacement for the entire computer, its CPU, memory, storage or software.
The key points
- GPUs trade architectural resources toward throughput across many operations, rather than making every individual operation finish faster.
- A GPU core count is not directly comparable to a CPU core count or to another vendor's differently defined execution units.
- Memory capacity, memory bandwidth, arithmetic throughput and communication are separate constraints.
- The useful performance measure is completed work at an acceptable quality and latency, not a hardware specification alone.
What does GPU mean?
GPU stands for graphics processing unit. Its name reflects its origins in rendering images, but the modern device is also a programmable parallel processor. Graphics, scientific simulation and machine learning all contain work that can be divided into many operations on different pieces of data. A GPU provides the execution machinery for those operations. The application must still express the work in a form supported by its hardware and software. [1]
A GPU is not precisely the same thing as a graphics card. A card is an assembly that can contain a processor, memory, power-delivery components, connectors and cooling hardware. Some GPUs are integrated with other processors; others are discrete devices on cards or server modules. Data-center accelerators do not need monitor outputs to be useful. Conversely, a product being called a graphics card does not establish whether a particular AI framework supports it. [1]
The best mental model is a division of responsibility. A program has a flow of instructions and data. Some stages benefit from a CPU's general-purpose execution, while selected stages can benefit from GPU acceleration. The resulting system is heterogeneous: different processors cooperate. When someone says a model runs on a GPU, that normally describes where its main tensor operations execute, not a claim that no CPU or operating system is involved. [3]
Why parallelism changes the architecture
Parallelism means performing work concurrently instead of completing every operation in one strictly sequential chain. The opportunity depends on dependencies. If the result of step two requires step one's output, those steps cannot simply be performed simultaneously. If many output elements can be computed independently from existing inputs, the program has data parallelism. The amount of usable parallelism is a property of the algorithm and its implementation, not something a larger processor can create automatically. [5]
Consider an illustrative task that increases the brightness of one million independent pixels. Each pixel can be processed using the same rule without waiting for a neighboring pixel. Now contrast that with following a linked list, where the address of the next item is obtained from the current one. More execution lanes can help the first task much more directly than the second. Real rendering and AI systems contain mixtures of regular, irregular and dependent work rather than one universal pattern.
Throughput and latency are different objectives. Throughput measures how much work finishes per unit of time. Latency measures how long one operation, request or job takes from a defined start to finish. A system may increase throughput by processing a larger batch while increasing the waiting time for an individual request. This is why a server generating many tokens each second across hundreds of users is not necessarily giving each user a fast response. [12]
CPU versus GPU is therefore not a contest between a smart chip and a large collection of unintelligent chips. CPUs also execute parallel instructions and often contain multiple cores. GPUs also have scheduling, caches and control logic. Their architectural balances differ. Workload size, branching, memory access, supported operations and the cost of moving data determine which combination is useful. [2]
What is inside a GPU?
A GPU contains groups of execution resources, not one undifferentiated sea of independent processors. NVIDIA describes its main programmable execution groups as streaming multiprocessors, or SMs. An SM combines instruction scheduling, registers, arithmetic resources and local storage resources. AMD uses its own organization and terminology, including compute units. Exact arrangements change between architectures, so a term learned for one family should not be assumed to define every other family. [3][6]
Arithmetic resources perform operations such as addition and multiplication. Some resources specialize in matrix operations, while graphics-focused devices may also contain specialized ray-tracing, texture or media engines. A program only benefits from a specialized unit when its work is mapped to operations that the unit supports. Merely owning a device with an advertised accelerator does not cause every part of a Python program to use it. [7][1]
Core counts require particular care. A CPU core is generally a relatively complete instruction-execution engine. A marketed GPU core can describe an arithmetic execution resource inside a larger scheduling group. Dividing one count by the other does not produce a meaningful performance ratio. Likewise, multiplying a GPU's arithmetic-core count by its clock is insufficient unless the operation type, issue rate and architectural execution rules are specified. [2][4]
Threads, warps and the SIMT model
A GPU kernel is a function launched for execution on the device. In CUDA, the programmer organizes its threads into blocks and grids. Threads use their indices to choose which data to process. The hardware schedules the work subject to resource limits. The number of logical threads launched is not the number of physical arithmetic units, and launching a million threads does not mean a million instructions execute in the same instant. [3]
CUDA groups threads into warps of 32 threads. This is NVIDIA's execution model, not a universal definition for every GPU. Under single-instruction, multiple-thread execution, threads run the same kernel but can have different data and control paths. When threads in a group follow different branches, execution resources may spend time on different active subsets. Branch divergence can therefore reduce efficiency, although its impact depends on the actual control flow and hardware. [3]
As an illustrative comparison, imagine a group of workers receiving the same instruction to process their assigned item. That is efficient when the items need similar work. If some items require ten additional steps and others do not, some workers may be inactive while those steps run. The analogy explains divergence, but not the complete hardware: real schedulers interleave many groups, and modern architectures have more sophisticated control than a literal classroom following one teacher.
Occupancy describes the proportion of supported resident warps that are active on an SM, subject to the architecture's limits. Registers per thread, shared memory per block and block size can limit residency. More occupancy can help hide waiting, but maximum occupancy is not synonymous with maximum application performance. A kernel with better data reuse can outperform another kernel with more resident work. Profiling is needed to identify the limiting resource. [5]
The memory hierarchy
Execution resources need a place to obtain inputs and store results. Registers hold thread-local working values. Shared memory in CUDA is explicitly managed storage that cooperating threads can use. Caches retain recently used data automatically according to hardware rules. Device memory, often called VRAM in everyday discussion, provides the much larger working capacity. These levels differ in capacity, access behavior and performance; they are not interchangeable just because each stores bits. [4]
The physical memory technology may be GDDR or high-bandwidth memory, HBM. HBM stacks memory dies and uses a wide interface close to the processor package. GPU memory is also distinct from a server's CPU RAM and from SSD storage. Loading a model from an SSD into host memory and then into device-accessible memory crosses boundaries. A server advertising a large quantity of system RAM is not thereby advertising the same quantity of dedicated GPU memory. [8][3]
Capacity answers how much state can fit. Bandwidth answers how quickly bytes can move through a particular interface. Latency answers how long a particular access takes. It is possible to have enough capacity but insufficient bandwidth, or high bandwidth but insufficient capacity. A description such as fast memory compresses three different questions into one phrase. Good infrastructure reasoning separates them before trying to compare devices. [4]
Suppose a hypothetical calculation reads 100 GB of data exactly once through a memory interface delivering a sustained 2 TB/s. Using decimal units, the transfer component alone needs at least 0.05 seconds. This is a simplified lower bound, not an application prediction: repeated reads, writes, cache hits, arithmetic and synchronization change the result. If the data fits in memory but must be reread repeatedly, capacity has solved only one of the constraints.
Why AI uses GPUs
Neural networks represent many of their computations using tensors: arrays with a defined shape and numerical representation. A matrix is a two-dimensional example. Multiplying matrices involves many multiply-and-accumulate operations that can be organized into reusable tiles. This structure gives GPUs useful parallel work and opportunities to reuse values. It does not imply that every neural-network operation has the same performance characteristics as a large matrix multiplication. [7]
Tensor cores are specialized NVIDIA execution resources for supported matrix operations. Supported formats and operation shapes depend on the architecture. Software libraries can select kernels that use them. Lower-precision representations can reduce data movement and increase arithmetic throughput, but numerical behavior must remain suitable for the task. A high low-precision performance rating should not be substituted for a high-precision scientific requirement. [7][9]
Training computes outputs, evaluates a loss, obtains gradients and updates model parameters. Inference uses a trained model to produce outputs without the normal training update loop. Training can require gradients, optimizer state and saved activations in addition to weights. Inference has its own state and temporary allocations. For autoregressive language models, the key-value cache grows with the served sequence and concurrency under the chosen cache strategy. [10][11]
A model that loads successfully has not necessarily met the application's serving requirements. A single short request can consume far less memory and communication than a concurrent workload with long prompts and long outputs. The specification should include model version, numerical format, prompt and output distributions, batching policy and latency targets. That description turns an abstract demand for a powerful GPU into a workload that can be measured. [12]
A worked memory-sizing example
Assume a hypothetical dense model has 20 billion stored parameters and uses two bytes per parameter. Weight storage alone is 40 billion bytes: 40 GB in decimal units, or about 37.3 GiB. This estimate does not include the KV cache, temporary tensors, runtime allocations or allocator overhead. It also says nothing about how fast the model will run. The calculation is useful because its assumptions and exclusions are visible.
Now assume the same hypothetical serving configuration needs another 12 GB for caches and other working memory at its intended concurrency. The estimated requirement becomes 52 GB before any additional safety margin. A 48 GB device would not satisfy that stated budget. Reducing concurrency, changing the representation, using supported offloading or distributing the model could change the requirement, but each change creates a different configuration to evaluate.
Two 24 GB GPUs do not automatically become a single transparent 48 GB device. A framework must distribute data or model state explicitly, and each device retains its own constraints. Communication and replicated state can prevent the full sum from being available for the intended purpose. Aggregate memory is useful only when the software can use its placement and the connections between devices support the resulting traffic. [13][14]
FLOPS, utilization and the roofline idea
FLOPS measures floating-point operations per second. It must be tied to a format and operation basis: FP64, FP32, lower-precision matrix arithmetic and sparsity-assisted arithmetic are different measurements. A fused multiply-add is commonly counted as two operations. Theoretical peak throughput assumes an execution pattern that keeps the relevant resources supplied; it does not include every delay in a complete application. [4][7]
Arithmetic intensity is the amount of arithmetic performed per byte moved through a specified memory boundary. In a simplified roofline model, achievable arithmetic throughput is bounded by the smaller of peak compute throughput and memory bandwidth multiplied by arithmetic intensity. This is a useful reasoning tool, not a substitute for profiling. Different memory levels, operation types and kernels can have different effective ceilings. [4]
Imagine a hypothetical device with 100 trillion relevant operations per second and 2 trillion bytes per second of memory bandwidth. For an operation with an intensity of ten operations per byte, the bandwidth-derived ceiling is 20 trillion operations per second. Doubling peak arithmetic alone would leave that simplified ceiling unchanged. Increasing reuse, so that more work is performed for each byte fetched, could matter more than purchasing a device with more arithmetic units.
A utilization percentage is also incomplete without its definition. Time with an active kernel is not necessarily the percentage of peak arithmetic throughput achieved. A kernel may be waiting for memory or doing small amounts of useful work while still being counted as active. Assess kernel timings, memory behavior, communication and the application's output together. Optimizing one dashboard number can otherwise hide the real bottleneck. [5]
The GPU is part of a system
Before computation, input data must be available. CPU preprocessing, storage reads, decompression and host-to-device transfers can determine when a GPU starts useful work. PCI Express commonly connects a discrete accelerator to the host system. Specialized links can connect accelerators with one another. A fast processor attached to an underprovisioned input pipeline can spend expensive time waiting. [5]
As a hypothetical end-to-end example, suppose a job spends 30 seconds loading and preparing data and 70 seconds on GPU computation. Making the GPU stage twice as fast reduces total time to 65 seconds, not 50. The overall speedup is about 1.54 times. Even eliminating GPU compute time entirely would leave the 30-second preparation stage. This is the practical consequence of optimizing only part of a workflow.
Multi-GPU workloads add communication. Data-parallel training, for example, needs coordination between workers; other parallelization strategies exchange model activations or partitioned state. NVLink and supported switches address certain scale-up connections, while scale-out networks connect larger systems. The presence of eight GPUs in one offering does not establish their topology, link generation or sharing conditions. [13][14]
Distributed performance is not obtained by multiplying a single-device score by the device count. Some work must synchronize, some state is duplicated, and network paths can be oversubscribed. Measure the intended number of devices with the intended software and batch arrangement. A benchmark from another topology can be informative about a component while remaining insufficient evidence for the full configuration. [13]
Drivers, frameworks and portability
A driver allows the operating system and applications to interact with the device. A programming platform provides interfaces and tooling for accelerator work. Frameworks such as PyTorch sit at another layer, translating high-level operations into supported implementations. CUDA is NVIDIA's platform; it is not the generic name for all GPU computing. AMD's ROCm and other ecosystems have their own supported hardware and software combinations. [3][6]
Compatibility is therefore a stack, not a checkbox. The accelerator architecture, driver, runtime, framework build and any custom extensions must work together. A familiar framework API can make application code portable while a custom kernel or particular operation still requires porting. Containers help package user-space dependencies, but they do not manufacture compatible hardware or remove the need for a working host driver. [16]
We recommend recording an exact environment for any meaningful comparison: device configuration, driver and library versions, model checkpoint, numerical format, input distribution and test procedure. Reproducibility is not bureaucracy; it is how another person distinguishes an architectural improvement from a changed benchmark. Without that context, a speed claim is difficult to apply to a purchasing decision.
Whole GPUs, partitions and cloud offerings
A cloud offering may expose a whole GPU, a hardware-supported partition, a time-shared device or a higher-level inference service. These are not identical products. NVIDIA Multi-Instance GPU, on supported devices and configurations, partitions resources into instances with defined compute and memory allocations. Support and partition profiles depend on the device. A partition should not be described as possessing the full parent GPU's resources. [15]
For an illustrative purchasing comparison, consider a whole-device offer and a smaller partition bearing the same parent model name. Equal hourly prices would not establish equal value. The accessible memory, execution resources, isolation, networking and permitted software could differ. The correct unit of comparison is the resource configuration actually supplied, not the most impressive specification associated with its family name.
Hardware choice also has operational boundaries. Power limits, cooling conditions and sustained clocks can affect the device's behavior. A technical specification does not prove current inventory, deployment readiness or a contractual performance promise. Those are separate sourcing questions. This chapter explains architecture; it is neither a live availability statement nor an independently measured ranking of providers. [5]
How to evaluate a GPU for a workload
Begin with correctness: required operations, numerical behavior, supported software and the memory needed at the intended scale. Then define acceptable output, such as a completed simulation with a stated tolerance or an inference service meeting a latency and quality target. After that, measure throughput, sustained behavior and total job time. Finally, compare the cost of obtaining that acceptable work, including preparation, data transfer and idle allocations.
Suppose configuration A costs a hypothetical $2 per hour and completes a job in four hours. Configuration B costs $3 per hour and finishes in two. Compute-only cost is $8 for A and $6 for B. That does not automatically make B the better purchase: provisioning delays, additional charges, software compatibility and successful completion still matter. But it shows why an hourly-price ranking and a completed-work-cost ranking can disagree.
Common mistakes follow from collapsing distinctions: more VRAM is not the same as more bandwidth; more arithmetic units do not guarantee faster serial code; a larger batch can improve throughput without improving interactive latency; a published maximum is not sustained performance; and an aggregate memory total is not a promise that every program can use it as one pool. Keep these questions separate and most GPU specifications become much easier to interpret.
The central lesson is that a GPU is a highly capable component inside a particular computing system. Its value comes from matching parallel execution, memory, communication and software to a real task. The strongest question is not which GPU is universally best, but which tested configuration meets this workload's constraints with an acceptable cost and operating risk.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1. Why might doubling arithmetic throughput barely change a program's runtime?
The program may be limited by memory movement, communication, CPU work or storage. Improving a resource that is not the bottleneck does not proportionally improve the complete job.
2. Do two 24 GB GPUs guarantee that a 48 GB workload fits?
No. Software must distribute the workload, individual devices have local limits, and duplicated state or communication buffers can consume part of the aggregate memory.
3. What is missing from the statement that a GPU achieves 1,000 TFLOPS?
The numerical format, operation and sparsity basis, whether the figure is theoretical or measured, and the workload and software conditions. It is not a universal application speed.
4. How do you distinguish latency from throughput in a chatbot?
Latency describes a user's waiting time, such as time to first token or between output tokens. Throughput counts accepted work across the service per second. Batching can improve one while worsening the other.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- Intel · What is a GPU? ↗ (opens in a new tab)Manufacturer explainer · Checked 28 September 2026
- Intel · CPU versus GPU ↗ (opens in a new tab)Manufacturer explainer · Checked 28 September 2026
- NVIDIA · CUDA programming model ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- NVIDIA · GPU Performance Background ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- NVIDIA · CUDA Best Practices Guide ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- AMD · MI250 microarchitecture ↗ (opens in a new tab)Architecture documentation · Checked 28 September 2026
- NVIDIA · Matrix Multiplication Background ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- Micron · High-bandwidth memory ↗ (opens in a new tab)Manufacturer documentation · Checked 28 September 2026
- PyTorch · Automatic mixed precision examples ↗ (opens in a new tab)Framework documentation · Checked 28 September 2026
- PyTorch · Optimizing model parameters ↗ (opens in a new tab)Framework tutorial · Checked 28 September 2026
- Hugging Face · Cache strategies ↗ (opens in a new tab)Framework documentation · Checked 28 September 2026
- NVIDIA · LLM inference optimization ↗ (opens in a new tab)Manufacturer technical article · Checked 28 September 2026
- PyTorch · Distributed Data Parallel ↗ (opens in a new tab)Framework tutorial · Checked 28 September 2026
- NVIDIA · NVLink and NVLink Switch ↗ (opens in a new tab)Manufacturer documentation · Checked 28 September 2026
- NVIDIA · Multi-Instance GPU introduction ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- NVIDIA · Container Toolkit overview ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.