AWS Trainium: the definition
AWS Trainium and AWS Inferentia are machine learning accelerators designed by Amazon Web Services and offered only as Amazon EC2 instances. Inferentia was built for inference, while Trainium targets training and, in its later generations, large-scale inference as well.
The key points
- AWS builds two families of AI chips: Inferentia for inference and Trainium for training and, increasingly, large-model serving.
- Trainium3, a 3nm chip, became available in Trn3 UltraServers of up to 144 chips in December 2025, and Trainium4 is in development.
- All of these chips are programmed through the AWS Neuron SDK, which supports PyTorch, JAX, vLLM and a kernel interface for custom operations.
- They are only available on AWS and compete with GPU instances mainly on price-performance, at the cost of a different software stack.
What Trainium and Inferentia are
Amazon Web Services designs its own machine learning accelerators and rents them through Amazon EC2. There are two product lines. Inferentia is aimed at running trained models, and Trainium is aimed at training them. Both are built around NeuronCores, AWS's compute units, and both are programmed through the same software kit, AWS Neuron. [1][2][8]
The chips are only offered inside AWS. The Neuron documentation lists the instance families that use them: Inf1 and Inf2 for Inferentia, and Trn1, Trn1n, Trn2 and Trn3 for Trainium. Customers therefore meet this silicon as an instance type, a managed service or an UltraServer, not as a card they can install elsewhere. [8]
AWS's motivation is familiar from Google's TPUs. Owning the chip lets a cloud provider tune hardware, networking and pricing together, and reduces dependence on a single outside GPU supplier. For customers the result is another option on the menu, with its own price, performance and porting cost.
Inferentia and Inferentia2
The first Inferentia chip arrived in EC2 Inf1 instances in 2019. Each chip had four NeuronCores and 8 GB of DDR4 memory, and an Inf1 instance could hold up to 16 chips. AWS claimed up to 2.3 times higher throughput and up to 70 percent lower cost per inference than comparable EC2 instances of the time. [1]
Inferentia2 moved to high-bandwidth memory. Each chip has two NeuronCores and 32 GB of HBM, and an Inf2 instance holds up to 12 chips. AWS cites 4 times the total memory and 10 times the memory bandwidth of Inf1, up to 190 teraflops of FP16 per chip, and support for FP32, TF32, BF16, FP16, INT8 and a configurable FP8 format. Inf2 was also the first AWS inference-optimised instance to support distributing one model across chips with a high-speed chip-to-chip link. [1]
Trainium: the first training chip
Trn1 instances carry 16 first-generation Trainium chips, each with two NeuronCores, and up to 512 GB of shared HBM with 9.8 TB per second of bandwidth. AWS rates a Trn1 instance at up to 3 petaflops of FP16 or BF16. The network-optimised Trn1n variant doubles Elastic Fabric Adapter bandwidth from 800 to 1,600 Gbps, which AWS says gives about 20 percent faster training on network-heavy jobs. [3]
AWS marketed Trn1 on cost, claiming up to 50 percent lower cost to train than comparable EC2 instances, and on scale: EC2 UltraClusters of up to 30,000 Trainium chips on a nonblocking petabit-scale network. [3]
Trainium2, UltraServers and Project Rainier
Trn2 instances became generally available on 3 December 2024. Each instance holds 16 Trainium2 chips with 1.5 TB of HBM3 and up to 20.8 petaflops of FP8 compute. AWS introduced the Trn2 UltraServer at the same time, first in preview: four Trn2 servers joined by the NeuronLink interconnect into a 64-chip unit with 6 TB of HBM and up to 83.2 FP8 petaflops. [5][4]
AWS claims Trn2 instances deliver 30 to 40 percent better price-performance than its GPU-based P5e and P5en instances. As with any vendor benchmark, the real gap depends on the model, the batch size and how well the software stack is tuned for each platform. [4]
The largest Trainium2 deployment is Project Rainier, an EC2 UltraCluster of Trn2 UltraServers built for Anthropic. AWS announced in late 2025 that the cluster was active, less than a year after the project was announced, with nearly half a million Trainium2 chips across several US data centers, and said Anthropic expected to be using more than one million Trainium2 chips for training and inference by the end of 2025. [6]
Trainium3 and what comes next
AWS made Trn3 UltraServers available on 2 December 2025. Trainium3 is AWS's first 3nm AI chip. An UltraServer holds up to 144 chips and delivers up to 362 petaflops of FP8, which AWS describes as 4.4 times the compute of the previous generation with 4 times better energy efficiency and almost 4 times the memory bandwidth. [7]
Per chip, AWS lists 144 GB of HBM3e with 4.9 TB per second of bandwidth, and 20.7 TB of HBM3e across a full 144-chip UltraServer. AWS also highlights hardware support for mixture-of-experts routing and dedicated collective-communication cores that sit apart from the compute engines. EC2 UltraClusters 3.0 are designed to connect up to one million Trainium chips. [2][7]
At the Trainium3 launch AWS said Trainium4 was in development, targeting at least 6 times the FP4 processing performance, 3 times the FP8 performance and 4 times the memory bandwidth of Trainium3. No availability date had been given in the sources reviewed as of September 2026. [7]
The Neuron SDK
AWS Neuron is the software layer that makes these chips usable. It includes the NeuronX compiler, which turns models into code for Trainium and Inferentia, the Neuron Runtime that executes them, and monitoring and profiling tools. On the framework side it supports PyTorch through torch-neuronx and JAX through a NeuronX plugin. [8]
Above the compiler sit libraries for large models. NeuronX Distributed provides tensor and pipeline parallelism for training and inference, vLLM on Neuron handles LLM serving with features such as speculative decoding and multi-LoRA serving, and the Neuron Kernel Interface lets engineers write custom kernels when the compiler's output is not fast enough. AWS says vLLM, Hugging Face Transformers and TorchTitan run natively on Trainium. [8][2]
Consider a team serving an open-weight model with vLLM on GPU instances. Moving to Trainium typically means switching to the Neuron build of vLLM, compiling the model ahead of time for fixed sequence lengths and batch sizes, and benchmarking. If the model uses a custom attention kernel written in CUDA, that kernel must be replaced by a Neuron equivalent or rewritten with NKI before results are comparable.
Where they fit versus GPUs
The case for Trainium is price-performance and supply at scale inside AWS. The vendor figures above, and the Project Rainier deployment, show AWS is prepared to put very large clusters behind the chip, and AWS names Anthropic, Databricks, OpenAI, Ricoh and Uber among Trainium users. [2][6][4]
The case for GPUs is portability and ecosystem. A model tuned for NVIDIA hardware can move between hyperscalers, specialised GPU clouds and on-premises clusters, while Trainium and Inferentia capacity exists only in AWS regions that stock it. Trn2, for example, launched in a single region, US East (Ohio). [5][8]
In our view, Trainium makes the most sense for teams that are already committed to AWS, run mainstream model architectures supported by Neuron, and consume enough compute that a double-digit price-performance gain outweighs the engineering time to port and validate. Smaller teams, or those that need to move between providers as prices change, will often find GPUs simpler. Kovara tracks GPU prices across providers so you can see what the GPU alternative costs today; compare those figures against an AWS quote for Trn2 or Trn3 capacity, and measure the result in cost per token or per training step on your own model.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1Can I buy Trainium or Inferentia chips outside AWS?
No. The chips are offered through Amazon EC2 instance families such as Inf1, Inf2, Trn1, Trn2 and Trn3, and through AWS services built on them. They are not sold as standalone hardware.
2Do I need to rewrite my PyTorch model to use Trainium?
Usually not from scratch. The AWS Neuron SDK compiles PyTorch and JAX models for Trainium and Inferentia, and AWS says libraries such as vLLM, Hugging Face Transformers and TorchTitan run on Trainium. Code that depends on hand-written CUDA kernels does need porting, for example to the Neuron Kernel Interface.
3What is a Trainium UltraServer?
An UltraServer links several Trainium servers with AWS's NeuronLink interconnect so they act as one larger unit. A Trn2 UltraServer joins 64 Trainium2 chips across four instances, and a Trn3 UltraServer scales to 144 Trainium3 chips.
4Is Inferentia still relevant now that Trainium also does inference?
Inf2 instances remain listed by AWS for inference, but AWS positions Trainium2 and Trainium3 for both training and serving large models. Which is more cost-effective depends on model size and the specific workload.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- AWS · AWS Inferentia ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AWS · AWS Trainium ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AWS · Amazon EC2 Trn1 instances ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- AWS · Amazon EC2 Trn2 instances and UltraServers ↗ (opens in a new tab)Cloud documentation · Checked 29 September 2026
- Amazon Press Center · AWS Trainium2 instances now generally available ↗ (opens in a new tab)Press release · Checked 29 September 2026
- About Amazon · AWS activates Project Rainier ↗ (opens in a new tab)Company announcement · Checked 29 September 2026
- About Amazon · Trainium3 UltraServers now available ↗ (opens in a new tab)Company announcement · Checked 29 September 2026
- AWS Neuron Documentation · AWS Neuron SDK ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
