TPU: the definition
A Tensor Processing Unit (TPU) is an application-specific integrated circuit designed by Google to accelerate the matrix arithmetic at the heart of neural networks. TPUs are used inside Google and are rented to customers through Google Cloud as individual chips, slices and pods.
The key points
- TPUs are Google-designed ASICs built around large systolic matrix units, first deployed in Google data centers in 2015 for inference.
- Chips are grouped into slices and pods over a dedicated inter-chip interconnect, which is what lets TPUs scale to thousands of chips per job.
- As of September 2026, TPU7x (Ironwood) is generally available and the eighth generation splits into TPU 8t for training and TPU 8i for inference.
- TPUs are only rentable on Google Cloud and reward matrix-heavy models written for JAX or PyTorch through the XLA compiler.
What a TPU is
A Tensor Processing Unit is a custom chip Google designed to speed up machine learning. Google's documentation describes TPUs as application-specific integrated circuits, meaning the silicon is specialised for one class of work rather than being a general-purpose processor. Every program that runs on a TPU is compiled by XLA, a just-in-time compiler that breaks large matrix multiplications into tiles sized for the hardware. [3]
Google recommends TPUs for models dominated by matrix computations, for models that train for weeks or months, and for recommendation systems with very large embedding tables. The same guidance lists cases where TPUs are a poor fit: code with frequent branching, workloads made mostly of element-wise operations, anything that needs high-precision arithmetic, and training loops that rely on custom operations. For those, Google points users to GPUs or CPUs. [3]
A TPU is best understood as a trade. By giving up flexibility that a GPU keeps, the chip can spend more of its area and power on multiply-accumulate units and on-chip memory. Whether that trade pays off depends almost entirely on how closely a workload matches what the chip was built to do.
Systolic arrays: the engine inside every TPU
The core of a TPU is the matrix multiply unit, or MXU, which is built as a systolic array. In a systolic array, a grid of simple arithmetic cells passes operands to its neighbours in lockstep, so each value read from memory is reused many times as it flows across the grid. Google's first-generation TPU used a 256 by 256 array of 8-bit integer multipliers, 65,536 in total, and the company's description of the chip explains that this short, neighbour-to-neighbour wiring saves both energy and memory traffic. [2][1]
Modern Cloud TPUs organise compute into TensorCores. Each TensorCore contains MXUs, a vector unit and a scalar unit. Google documents 128 by 128 MXUs on earlier Cloud TPU generations and 256 by 256 MXUs on v6e (Trillium) and TPU7x (Ironwood). By default the MXUs multiply in bfloat16 and accumulate results in FP32. [4][5]
Some TPUs also include SparseCores, dataflow processors aimed at the irregular memory lookups of embedding-heavy recommendation models. According to Google Cloud documentation, v5p and TPU7x chips each carry four SparseCores and v6e chips carry two. The TPU v4 paper reported that SparseCores sped up embedding-heavy models by roughly 5 to 7 times while using about 5 percent of die area and power. [4][7]
Picture multiplying a 1,024 by 1,024 weight matrix by a batch of activations on a chip with 128 by 128 MXUs. XLA splits the weights into 64 tiles of 128 by 128, streams activations through each tile, and sums the partial results. If a model's layer widths are not multiples of the tile size, some cells sit idle, which is one reason TPU tuning guides pay attention to tensor shapes.
From TPU v1 to TPU v4
Google's 2017 ISCA paper describes the first TPU, which had been running inference in Google data centers since 2015. It offered a peak of 92 trillion 8-bit operations per second, carried 28 MiB of on-chip memory and was attached to host servers over PCIe. Compared with the Intel Haswell CPUs and NVIDIA K80 GPUs in the same data centers, the authors measured roughly 15 to 30 times higher performance on production neural networks and 30 to 80 times better performance per watt. [1]
The first TPU was inference-only and had deliberately simple control logic: no caches, branch prediction or out-of-order execution. Google's own write-up says the design prioritised predictable latency, which mattered for user-facing services such as search and translation. [2][1]
Later generations were offered to cloud customers for training as well as serving. Google Cloud documents TPU v3 as having two TensorCores per chip, 32 GiB of HBM2 and 123 teraflops of bfloat16 compute per chip, with 1,024 chips in a pod joined in a 2D torus. The TPU v4 paper, published in 2023, describes a 4,096-chip supercomputer deployed since 2020 that used optical circuit switches to reconfigure its interconnect, which the authors said accounted for under 5 percent of system cost. [6][7]
v5e, v5p, Trillium and Ironwood
The v5 generation split into two products. TPU v5e is a cost-oriented chip with one TensorCore, 16 GB of HBM and 197 teraflops of bfloat16 compute, deployed in 256-chip pods on a 2D torus. TPU v5p is the high-end part: two TensorCores and four SparseCores per chip, 95 GiB of HBM at about 2.8 TB per second, and 8,960 chips per pod connected in a 3D torus, with single jobs of up to 6,144 chips. [8][9]
TPU v6e, branded Trillium, became generally available in December 2024. Each chip has one TensorCore with 256 by 256 MXUs, 32 GB of HBM and 918 teraflops of bfloat16 compute, and a full pod holds 256 chips. Google's release notes credit Trillium with more than 4 times the training performance and up to 3 times the inference throughput of the v5 generation. [10][12]
TPU7x is the first product in the seventh-generation Ironwood family. It entered preview on 24 November 2025 and became generally available on 31 March 2026. Each chip combines two chiplets for a total of two TensorCores, four SparseCores and 192 GB of HBM with about 7.4 TB per second of bandwidth. Google lists peak compute of 2,307 teraflops in bfloat16 and 4,614 teraflops in FP8, and a full pod of 9,216 chips. [12][11]
The eighth generation: TPU 8t and TPU 8i
On 22 April 2026 Google announced its eighth TPU generation as two separate chips. TPU 8t targets large-scale pre-training: it keeps SparseCores and a 3D torus, scales to 9,600 chips in a superpod, carries 216 GB of HBM per chip and adds native FP4 arithmetic. TPU 8i targets serving and reasoning workloads: it carries 288 GB of HBM and 384 MB of on-chip SRAM per chip, adds a Collectives Acceleration Engine, and uses a new topology Google calls Boardfly with up to 1,152 chips per pod. [13][14]
Google claims TPU 8t improves training price-performance by 2.7 times over Ironwood and that TPU 8i improves inference price-performance by 80 percent, with both delivering about twice the performance per watt. These are vendor figures from the launch material. At announcement, Google invited customers to register interest and said the chips would be available later in 2026. General availability was not confirmed in the sources reviewed for this article as of September 2026. [13][14]
Splitting one generation into training and inference chips is a meaningful signal. It suggests that at hyperscale, the memory and networking needs of serving large models with long contexts have diverged enough from pre-training to justify separate silicon.
Pods, slices and the inter-chip interconnect
What sets TPUs apart from a rack of accelerator cards is the dedicated network between chips. Google defines a pod as a contiguous set of TPUs joined by a specialised network, and a slice as the subset of a pod you actually allocate, linked by high-speed inter-chip interconnect (ICI). Slices are described by 2D or 3D topologies, and Multislice connects several slices over the ordinary data-center network when a job needs more chips than one slice provides. [4]
ICI bandwidth has grown with each generation. Google documents 400 GB per second per chip on v5e, 800 GB per second on v6e and 1,200 GB per second bidirectional on both v5p and TPU7x. On TPU7x, slice topologies range from a single 4-chip host at 2x2x1 up to 8x16x16, or 2,048 chips, and each virtual machine exposes 4 chips. [8][10][9][11]
How to rent TPUs on Google Cloud
TPUs are available only through Google Cloud. Google lists Compute Engine, Google Kubernetes Engine and its managed AI platform as the access paths, and notes that the older Cloud TPU API is deprecated. Framework support covers JAX, PyTorch and TensorFlow in general, although TPU7x drops TensorFlow and supports only JAX and PyTorch. [3][11]
Google offers several consumption models. On-demand capacity is pay-as-you-go without preemption. Spot VMs are cheaper but can be reclaimed at any time. Flex-start lets a job obtain capacity for up to 7 days without a prior reservation, and calendar-mode reservations secure capacity for 1 to 90 days from a chosen start date; both apply to TPU7x, v6e and v5p. Longer reservations of a year or more come with committed-use discounts. [15]
A team fine-tuning a mid-sized open model might start on a single v6e host with 8 chips to validate its JAX or PyTorch code under XLA, then move to a reserved v5p or TPU7x slice once the training recipe is stable. Spot capacity suits checkpointed experiments, while production serving usually needs on-demand or reserved capacity.
Where TPUs fit when choosing compute today
TPUs are a strong option for teams already on Google Cloud whose models are large, matrix-heavy and written for JAX or PyTorch with XLA in mind. They are a weaker fit for code that depends on hand-written CUDA kernels or for buyers who want to move workloads freely between clouds and neoclouds. We suggest treating TPUs as one line in a broader comparison: use Kovara to see current GPU prices and provider options, then weigh them against a TPU quote for the same job, measuring cost per trained token or per served request rather than cost per chip-hour.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1Can I buy a TPU and install it in my own server?
Data-center TPUs are not sold as standalone cards. Customers access them through Google Cloud, using Compute Engine, Google Kubernetes Engine or Google's managed AI platform, and pay for capacity through on-demand, Spot, flex-start or reservation options.
2Which frameworks run on TPUs?
Google documents JAX, PyTorch and TensorFlow support for Cloud TPU, with all code compiled by the XLA compiler. The newest generation, TPU7x (Ironwood), supports JAX and PyTorch but not TensorFlow.
3What is the difference between a TPU slice and a TPU pod?
A pod is the full group of TPU chips joined by Google's dedicated inter-chip interconnect. A slice is a subset of a pod that you actually reserve and run on, described by a topology such as 2x2x1 or 4x4x4. Multislice joins several slices over the data-center network.
4What is the newest TPU generation as of September 2026?
TPU7x, the first Ironwood product, became generally available on Google Cloud on 31 March 2026. Google announced the eighth generation, TPU 8t for training and TPU 8i for inference, on 22 April 2026, with availability expected later in 2026.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- arXiv · In-Datacenter Performance Analysis of a Tensor Processing Unit (Jouppi et al., ISCA 2017) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- Google Cloud Blog · An in-depth look at Google's first Tensor Processing Unit (TPU) ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- Google Cloud · Introduction to Cloud TPU ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud · TPU architecture ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud · The bfloat16 numerical format ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud · TPU v3 ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- arXiv · TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning (Jouppi et al., 2023) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- Google Cloud · TPU v5e ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud · TPU v5p ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud · TPU v6e ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud · TPU7x (Ironwood) ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud · Cloud TPU release notes ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Google Cloud Blog · TPU 8t and TPU 8i technical deep dive ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- DatacenterDynamics · Google unveils eighth-generation TPUs, two dedicated training and inference chips ↗ (opens in a new tab)News report · Checked 29 September 2026
- Google Cloud · Cloud TPU consumption options ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
