The key points
- GPUs are programmable parallel processors with a broad software ecosystem, while AI ASICs specialise the silicon for neural-network math to gain efficiency.
- Specialisation can deliver large efficiency gains on matching workloads, as Google's first TPU showed, but narrows the range of code that runs well.
- Most custom accelerators are tied to one cloud or one company: TPUs to Google Cloud, Trainium to AWS, Maia to Azure and MTIA to Meta's own data centers.
- For most buyers the practical question is not which chip is fastest but which platform offers the best cost per result for a specific model and team.
Two approaches to AI hardware
A graphics processing unit is a programmable parallel processor. NVIDIA describes CUDA as a parallel computing platform and programming model, and it is used across deep learning, scientific computing and high-performance computing, which is why the same GPU can run model training, inference and simulation. An application-specific integrated circuit, or ASIC, is designed for a narrower job. Google describes its TPUs this way: custom ASICs built to accelerate machine learning workloads, with every program compiled by the XLA compiler. [1][2]
The distinction has blurred. Modern GPUs contain specialised matrix units, and NVIDIA's Hopper generation added FP8 tensor core arithmetic that NVIDIA says doubles throughput compared with FP16 or BF16. Meanwhile, the newest ASICs include vector units, programmable kernel languages and support for mainstream frameworks. Both camps are converging on dense low-precision matrix math surrounded by as much fast memory and interconnect bandwidth as possible. [3][2][4]
It is more useful to think of a spectrum than a binary. At one end sit fully programmable GPUs; in the middle, cloud accelerators such as TPU and Trainium that run standard frameworks through a compiler; at the far end, architectures such as wafer-scale engines or dataflow inference chips that reorganise the whole system around one class of model.
Why specialisation can pay off
The clearest early evidence came from Google's first TPU. In the 2017 ISCA paper describing it, Google reported that the chip ran production inference 15 to 30 times faster than contemporary Haswell CPUs and K80 GPUs, and delivered 30 to 80 times better performance per watt. The design removed features general-purpose processors rely on, such as caches and out-of-order execution, in favour of a deterministic execution model with a large 8-bit matrix unit. [5]
Those gains came with limits. The same paper noted that the first TPU was often held back by memory bandwidth, and Google's current guidance still says TPUs are a poor fit for workloads with frequent branching, many element-wise operations, high-precision arithmetic or custom operations inside the training loop. Specialisation helps most when a workload looks like the one the chip was designed around. [5][2]
Today's hyperscaler chips are marketed on price-performance rather than raw speed. AWS says Trn2 instances offer 30 to 40 percent better price-performance than its GPU-based P5e and P5en instances, and Microsoft says Maia 200 delivers 30 percent better performance per dollar than the latest hardware previously in its fleet. These are vendor claims, measured on workloads each vendor chose. [6][7]
The main custom accelerators in 2026
Google's TPU line is the most mature. TPU7x, the first Ironwood chip, became generally available on Google Cloud in March 2026 with 192 GB of HBM per chip and pods of up to 9,216 chips. AWS offers two families: Inferentia for inference and Trainium for training and serving, with Trainium3 UltraServers of up to 144 chips available since December 2025. [8][4][9]
Microsoft announced Maia 200 in January 2026 as an inference accelerator built on a 3nm process, with 216 GB of HBM3e and over 10 petaflops of FP4 compute. Microsoft said it runs workloads including OpenAI models and Microsoft Copilot, and launched a preview SDK with PyTorch and Triton support. Meta's MTIA line serves Meta's own ranking, recommendation and generative AI workloads; in March 2026 Meta said MTIA 300 was in production and outlined MTIA 400, 450 and 500 on a cadence of roughly one new chip every six months. [7][10]
Outside the hyperscalers, Intel's Gaudi 3 pairs 128 GB of HBM2E with 24 integrated 200 Gb Ethernet ports and is available on IBM Cloud. Intel has since announced a separate inference GPU, Crescent Island, with customer sampling expected in the second half of 2026, after Gaudi adoption fell short of Intel's own targets. Cerebras builds wafer-scale processors whose latest generation packs about 4 trillion transistors and 900,000 cores on a single wafer, sold as systems or through Cerebras's own cloud. [11][12][13]
Groq, a company focused on AI inference, sells access to its technology through the GroqCloud service. In December 2025 Groq signed a non-exclusive licensing agreement with NVIDIA covering its inference technology, and its founder and president joined NVIDIA. Groq said it would remain independent under a new chief executive and that GroqCloud would keep operating. [15]
Flexibility and software
Software is where GPUs hold their largest advantage. Google recommends GPUs over TPUs for models with custom PyTorch or JAX operations that TPUs do not support, and the TPU7x generation supports only JAX and PyTorch, not TensorFlow. On AWS, custom kernels that go beyond what the Neuron compiler produces must be written in the Neuron Kernel Interface rather than CUDA. [2][8][16]
ASIC vendors are narrowing the gap by targeting the libraries most teams actually use. AWS says vLLM, Hugging Face Transformers and TorchTitan run natively on Trainium; Microsoft's Maia SDK includes PyTorch integration and a Triton compiler; Meta says MTIA builds on PyTorch, vLLM and Triton; and Intel lists PyTorch, ONNX and DeepSpeed support for Gaudi 3. [4][7][10][11]
A research lab testing a new attention variant every week will usually move faster on GPUs, where a custom kernel can be written in CUDA or Triton and run the same day. A company serving one popular open model at high volume faces the opposite trade: the model is stable, supported by vLLM on several accelerators, and a lower cost per token can justify a one-time porting and validation effort.
Performance per watt and system design
Efficiency increasingly comes from system design as much as from the chip. Google's TPU v4 paper describes optical circuit switches that reconfigure the network between 4,096 chips at under 5 percent of system cost, and SparseCores that speed embedding-heavy models by 5 to 7 times using about 5 percent of die area. Google claims its TPU 8t and TPU 8i chips, announced in April 2026, roughly double performance per watt over Ironwood. [17][18]
Different architectures place memory differently. Cerebras compares a CS-3 system, which draws about 23 kW, with an NVIDIA DGX B200, and argues its wafer-scale design yields about 2.2 times better performance per watt. Microsoft's Maia 200 combines 272 MB of on-chip SRAM with HBM3e, and Google's TPU 8i carries 384 MB of on-chip SRAM to keep more of a model's working data close to the compute units during inference. [14][7][18]
Who can actually buy or rent them
Availability is the most practical difference for buyers. TPUs are rentable only on Google Cloud, through Compute Engine, Google Kubernetes Engine and Google's managed AI platform. Trainium and Inferentia are offered only as Amazon EC2 instances. Maia 200 was deployed first in Microsoft's US Central Azure region, and MTIA is used inside Meta's own infrastructure. [2][16][7][10]
Some alternatives are more open. Gaudi 3 is available as virtual servers and OpenShift clusters on IBM Cloud, Cerebras sells systems and a cloud service, and GroqCloud offers hosted inference. GPUs, by contrast, are sold by several vendors and rented by hyperscalers and a large number of specialised GPU clouds, so the same model can move between providers as prices and capacity change. [11][13][15]
How to choose today
We think the ASIC versus GPU question is better framed as platform versus platform. A custom accelerator is worth evaluating when your workload is large, stable and built on well-supported frameworks, and when you are comfortable committing to the one cloud that offers it. GPUs remain the default when your models change often, depend on custom kernels, or need to move between providers. In either case, benchmark your own model and compare cost per token or per training step, not headline petaflops. Kovara lets you compare live GPU prices and providers, which gives you the GPU baseline any ASIC quote should beat.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · CUDA C++ Programming Guide ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- Google Cloud · Introduction to Cloud TPU ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Hopper Architecture In-Depth ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- AWS · AWS Trainium ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- arXiv · In-Datacenter Performance Analysis of a Tensor Processing Unit (Jouppi et al., ISCA 2017) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- AWS · Amazon EC2 Trn2 instances and UltraServers ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Official Microsoft Blog · Maia 200: The AI accelerator built for inference ↗ (opens in a new tab)Company announcement · Checked 29 September 2026
- Google Cloud · TPU7x (Ironwood) ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- About Amazon · Trainium3 UltraServers now available ↗ (opens in a new tab)Company announcement · Checked 29 September 2026
- Meta Newsroom · Expanding Meta's custom silicon to power our AI workloads ↗ (opens in a new tab)Company announcement · Checked 29 September 2026
- IBM · Intel Gaudi 3 AI accelerators on IBM Cloud ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- DatacenterDynamics · Intel unveils Crescent Island data center GPU for inferencing workloads ↗ (opens in a new tab)News report · Checked 29 September 2026
- Cerebras · Wafer-Scale Engine ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- Cerebras Blog · Cerebras CS-3 vs NVIDIA B200: 2024 AI accelerators compared ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- Groq Newsroom · Groq and Nvidia enter non-exclusive inference technology licensing agreement ↗ (opens in a new tab)Company announcement · Checked 29 September 2026
- AWS Neuron Documentation · AWS Neuron SDK ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- arXiv · TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning (Jouppi et al., 2023) ↗ (opens in a new tab)Academic paper · Checked 29 September 2026
- Google Cloud Blog · TPU 8t and TPU 8i technical deep dive ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
