The key points
- NVIDIA’s compute lineage starts with the 2006 G80, the first GPU programmable in C through CUDA, and continued through Fermi, Kepler and Maxwell.
- Pascal’s P100 (2016) brought HBM2 and NVLink, and Volta’s V100 (2017) introduced the first Tensor Cores, units built for deep-learning matrix maths.
- Ampere added TF32, sparsity and GPU partitioning; Hopper added FP8 and the Transformer Engine; Blackwell joined two dies in one GPU and added FP4.
- Rubin, announced in January 2026, was ramping into full production at major cloud providers as of NVIDIA’s August 2026 results.
How to read this timeline
NVIDIA typically ships each GPU architecture in several chips: a large flagship for data centers and smaller dies for workstations and gaming. This guide follows the data-center line, giving for each generation the year it was introduced, its flagship compute part and the capability that most distinguished it. Where NVIDIA’s own documentation gives a figure, such as a transistor count or memory bandwidth, it is quoted as NVIDIA states it.
A useful way to see the pattern: early generations made GPUs programmable and reliable enough for scientific computing; the middle generations added memory bandwidth and fast links between GPUs; and every generation since Volta has been shaped mainly by the demands of neural networks, adding lower-precision number formats and specialised matrix units.
Tesla and Fermi (2006–2010): making GPUs programmable
G80, introduced in November 2006 in the GeForce 8800, Quadro FX 5600 and Tesla C870, was NVIDIA’s first unified graphics and compute architecture. It replaced separate vertex and pixel pipelines with one array of processors, introduced the single-instruction, multiple-thread execution model and shared memory, and was the first GPU that could be programmed in C through CUDA. This guide groups G80 and its 2008 successor under the Tesla name, which NVIDIA used for the compute products built on them. [1]
GT200, launched in June 2008 in products including the Tesla T10, raised the core count from 128 to 240, doubled the register file and added double-precision floating point for scientific work. [1]
Fermi, launched in 2010, was aimed squarely at high-performance computing. Its first chip had 3.0 billion transistors and up to 512 CUDA cores, and NVIDIA described it as the biggest architectural leap since G80. Key additions were error-correcting (ECC) memory, a true cache hierarchy with a configurable L1 cache and a 768 KB shared L2, full C++ support and roughly eight times the peak double-precision throughput of GT200. [1][2]
Kepler and Maxwell (2012–2014): scale and efficiency
Kepler’s flagship compute chip, GK110, was unveiled at NVIDIA’s GPU Technology Conference in 2012 for the Tesla K20, with deliveries planned for the fourth quarter of that year. Built on a 28-nanometre process, it had 7.1 billion transistors and 2,880 CUDA cores in its full form, and delivered more than 1 teraflop of double-precision throughput. [3][2]
Kepler’s key new capabilities concerned how work reached the GPU. Dynamic Parallelism let GPU code launch new work itself without returning to the CPU. Hyper-Q allowed 32 simultaneous hardware-managed connections from the host, compared with one on Fermi. GPUDirect RDMA let network adapters and storage devices read and write GPU memory directly, an early step toward the multi-GPU clusters that dominate today. [2]
Maxwell focused on energy efficiency. In September 2014 NVIDIA said its second-generation Maxwell chip delivered about 40 per cent more performance per CUDA core and roughly twice the efficiency of Kepler’s GK104, through a redesigned streaming multiprocessor, more shared memory and a larger L2 cache. Its data-center flagship was the Tesla M40, based on the GM200 chip. [4][5]
Pascal and Volta (2016–2017): the deep-learning pivot
Pascal’s Tesla P100, announced in April 2016, was built on a 16-nanometre FinFET process with 15.3 billion transistors. It was the first GPU accelerator to use HBM2 stacked memory, reaching 720 GB per second, three times the bandwidth of the Tesla M40. It also introduced NVLink, NVIDIA’s high-speed GPU-to-GPU link, at up to 160 GB per second, native half-precision (FP16) arithmetic useful for deep learning, and page faulting for unified CPU-GPU memory. [5]
Volta’s Tesla V100, announced on 10 May 2017, introduced Tensor Cores. Each performs a four-by-four matrix multiply and accumulate in one operation, taking FP16 inputs and accumulating in FP16 or FP32. The V100 carried 640 of them, eight per streaming multiprocessor, for a peak of 125 Tensor teraflops. It had 21.1 billion transistors on an 815 square-millimetre die made on TSMC’s 12-nanometre FFN process, 16 GB of HBM2 at 900 GB per second and second-generation NVLink at 300 GB per second. [6][7]
Tensor Cores mark the point at which NVIDIA’s data-center GPUs stopped being graphics chips adapted for computing and became processors designed around neural networks. Every flagship since has devoted a growing share of its silicon and its marketing to them.
Turing and Ampere (2018–2020): inference and flexibility
Turing, unveiled in September 2018, is best known for adding RT Cores that accelerate real-time ray tracing. For data centers its significance lay in its Tensor Cores, which added INT8 and INT4 modes for inference. The Tesla T4, announced on 12 September 2018, packed 320 Turing Tensor Cores into a 75-watt PCIe card. [8][9]
Ampere’s A100, introduced on 14 May 2020, had 54.2 billion transistors on an 826 square-millimetre die made on TSMC’s 7-nanometre N7 process. Its third-generation Tensor Cores added TensorFloat-32 (TF32), a format that gives deep-learning frameworks an easy way to accelerate 32-bit floating-point work, and support for 2:4 structured sparsity, which can double throughput on suitably pruned models. [10]
A100 also introduced Multi-Instance GPU (MIG), which partitions one GPU into as many as seven isolated instances, so that several users or jobs can securely share one GPU. The launch model had 40 GB of HBM2 at about 1.56 TB per second and third-generation NVLink at 600 GB per second. [10]
Hopper and Blackwell (2022–2024): the transformer era
Hopper’s H100, announced on 22 March 2022, has 80 billion transistors on an 814 square-millimetre die built on a TSMC 4N process customised for NVIDIA. Its headline feature is the Transformer Engine, which uses the new 8-bit FP8 format where precision allows and switches to 16-bit where it does not, to speed up the transformer models behind large language models. The H100 SXM5 was the first GPU with HBM3, at more than 3 TB per second, and fourth-generation NVLink raised GPU-to-GPU bandwidth to 900 GB per second. [11]
Blackwell, announced on 18 March 2024, broke the single-die model. Its GPU combines two reticle-sized dies, built on a custom TSMC 4NP process, into one unified GPU with 208 billion transistors, connected by a 10 TB per second chip-to-chip link. A second-generation Transformer Engine added 4-bit floating point (FP4) for inference, and fifth-generation NVLink provides 1.8 TB per second per GPU. [12]
Blackwell’s flagship product is as much a system as a chip. The GB200 NVL72 is a liquid-cooled rack joining 36 Grace Blackwell Superchips, 72 GPUs and 36 Grace CPUs in all, into one NVLink domain. NVIDIA claimed up to 30 times the large-language-model inference performance of an equivalent H100 system. [12]
Rubin (2026): the newest generation
On 5 January 2026 NVIDIA announced the Rubin platform, made up of six new chips: the Vera CPU, the Rubin GPU, the NVLink 6 switch, the ConnectX-9 SuperNIC, the BlueField-4 DPU and the Spectrum-6 Ethernet switch. NVIDIA rates the Rubin GPU at 50 petaflops of NVFP4 compute for inference, with HBM4 memory and a third-generation Transformer Engine. Compared with Blackwell, NVIDIA claims up to a tenfold reduction in inference cost per token and a fourfold reduction in the number of GPUs needed to train mixture-of-experts models. [13]
NVLink 6 provides 3.6 TB per second per GPU, and the Vera Rubin NVL72 rack combines 72 Rubin GPUs and 36 Vera CPUs. NVIDIA said Rubin-based products would reach partners in the second half of 2026. In its results for the quarter ended 26 July 2026, the company said Vera Rubin was ramping into full production, with racks running at Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. [13][14]
NVIDIA’s performance claims for Rubin are vendor figures. As of September 2026, independent measurements across a range of real workloads are still limited, and buyers should treat launch claims as upper bounds.
Choosing between generations today
Several generations are rented side by side in today’s cloud market, from Turing-era T4 cards to A100, H100 and Blackwell systems. The newest parts offer the most compute and the lowest-precision formats, but older generations can be cheaper per job if a workload does not need FP8 or FP4, fits in their memory, or runs on a single GPU. Memory capacity and bandwidth, which grew from 16 GB at 720 GB per second on the P100 to HBM3 at more than 3 TB per second on the H100, often matter as much as raw teraflops. On Kovara you can compare prices for each of these GPUs across providers, open a GPU page to check its memory and interconnect, or ask Kova which generation is the most cost-effective for a given model.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · Fermi Compute Architecture Whitepaper ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · Kepler GK110/GK210 Architecture Whitepaper ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- Tom’s Hardware · Nvidia Debuts GK110-based 7.1 Billion Transistor Super GPU ↗ (opens in a new tab)News report · Checked 29 September 2026
- NVIDIA Technical Blog · Maxwell: The Most Advanced CUDA GPU Ever Made ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · Inside Pascal: NVIDIA’s Newest Computing Platform ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · Tesla V100 GPU Architecture Whitepaper ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Launches Revolutionary Volta GPU Platform ↗ (opens in a new tab)Company press release · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Turing Architecture In-Depth ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Newsroom · New NVIDIA Data Center Inference Platform to Fuel Next Wave of AI-Powered Services ↗ (opens in a new tab)Company press release · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Ampere Architecture In-Depth ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Hopper Architecture In-Depth ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Blackwell Platform Arrives to Power a New Era of Computing ↗ (opens in a new tab)Company press release · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA Kicks Off the Next Generation of AI With Rubin ↗ (opens in a new tab)Company press release · Checked 29 September 2026
- NVIDIA Newsroom · Financial Results for Second Quarter Fiscal 2027 ↗ (opens in a new tab)Company press release · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
