Glossary
The language of compute.
51 terms from GPUs and networking to data centers and buying models, explained without jargon.
Hardware
A100NVIDIA’s Ampere-generation data center GPU, the predecessor to the H100.AcceleratorA chip designed to speed up specific workloads such as AI. GPUs are the best-known accelerators; others include TPUs and custom AI chips.B200An NVIDIA data center GPU from the Blackwell generation, the successor architecture to Hopper.FLOPSFloating-point operations per second: a measure of raw computing throughput, often quoted in teraflops or petaflops.GPUA graphics processing unit: a chip built to run many calculations in parallel. GPUs are the main hardware for training and running AI models.H100NVIDIA’s Hopper-generation data center GPU, widely used for AI training and inference. It comes in several form factors, including SXM and PCIe.H200An NVIDIA Hopper-generation GPU that pairs the H100’s compute design with larger and faster memory.HBMHigh Bandwidth Memory: stacked memory placed next to the GPU chip so data can move to and from it very quickly.Memory bandwidthHow much data a GPU can read from or write to its memory each second, usually measured in terabytes per second.MI300XAMD’s Instinct MI300X data center accelerator for AI and high-performance computing.PCIePCI Express: the standard connection between a GPU and the rest of a server. PCIe GPUs plug in as cards.SXMNVIDIA’s socketed GPU form factor, mounted directly on a server board and designed for NVLink and higher power.Tensor CoreSpecialised units inside NVIDIA GPUs that perform the matrix maths at the heart of AI far faster than general-purpose cores.VRAMThe memory on a GPU, measured in gigabytes. It holds the model weights, working data and caches while the GPU runs.
Networking
InfiniBandA high-performance networking technology with very low latency, widely used to connect servers in AI training clusters.NVLinkNVIDIA’s high-speed link that connects GPUs directly to each other inside a server, much faster than the standard PCIe bus.NVSwitchA switch chip that lets every GPU in an NVLink-connected server talk to every other GPU at full speed.RoCERDMA over Converged Ethernet: a way to get InfiniBand-style direct memory transfers over Ethernet networks.
AI workloads
Fine-tuningFurther training an existing model on a smaller, specific dataset to adapt it to a task.InferenceRunning a trained model to produce answers or predictions for users.KV cacheMemory that stores intermediate results of attention while a language model generates text, so they don’t have to be recomputed.MIGMulti-Instance GPU: an NVIDIA feature that splits one GPU into several isolated smaller GPUs.Precision (FP8, BF16, FP16)The number format used for calculations. Lower precision such as FP8 uses less memory and runs faster, at some cost in numerical accuracy.TrainingTeaching a model by processing large amounts of data and adjusting its parameters. It needs many GPUs working together for days or weeks.
Data centers
ColocationRenting space, power and cooling in someone else’s data center to house your own hardware.Data residencyThe requirement that data is stored and processed in a specific country or region.Liquid coolingRemoving heat with liquid instead of air, either at the chip or across the rack. It handles much higher power densities than air.Megawatt (MW)A unit of power equal to one million watts. Data center capacity is usually described in megawatts of power available.PUEPower Usage Effectiveness: total facility energy divided by the energy used by IT equipment. A value closer to 1 means less energy is spent on overheads like cooling.Rack densityHow much power each server rack can supply and cool, measured in kilowatts per rack.RegionA geographic area where a provider offers capacity, such as a city, metro or country.TDPThermal design power: the amount of heat a chip is designed to produce, and roughly the power it draws under load, in watts.
Buying compute
Bare metalRenting a whole physical server with no virtualisation layer between you and the hardware.ClusterA group of servers connected by a fast network so they can work on one job together.Committed termThe length of time a buyer agrees to pay for capacity, such as one, three or twelve months.EgressData leaving a provider’s network, which some providers charge for per gigabyte.GPU-hourOne GPU used for one hour. It is the standard unit for comparing compute prices.HyperscalerA very large cloud provider, such as the major public clouds, operating data centers at global scale.NeocloudA newer cloud provider that specialises in GPU and AI infrastructure rather than general-purpose cloud services.NodeOne server in a cluster. A GPU node typically holds several GPUs, commonly eight.On-demandRenting compute by the hour or minute with no long commitment. You can start and stop when you like, subject to availability.Reserved capacityCommitting to use compute for a fixed term, often months, in exchange for a lower price or guaranteed availability.SLAService level agreement: a provider’s written commitment on things like uptime, with remedies if it is missed.Spot (interruptible)Spare capacity offered at a discount that the provider can take back at short notice.
Kovara
Availability statusWhether an offering is Available, Limited or Unavailable, or unknown when the source does not say.Compute IndexKovara’s benchmark for compute prices and availability, calculated from independent provider contributions with a published methodology.Confidence gradeKovara’s A–D rating of how well a fact is supported: A highly verified, B strong evidence, C incomplete evidence, D estimate or unverified.FixingAn official published benchmark value from the Compute Index, recorded with the methodology version used and never silently rewritten.Indicative priceA live estimate that has not been published as an official fixing. It is labelled as indicative so it is never mistaken for one.Normalised priceA price converted to a common unit, such as US dollars per GPU-hour, so that different offers can be compared.Verified ProviderA provider whose identity a person at Kovara has reviewed. The status is never granted automatically.
Missing a term?
Ask Kova to explain anything about chips, networking or data centers, or tell us what to add.
