AMD Instinct MI300X: the definition
The AMD Instinct MI300X is a data center GPU built from stacked chiplets on AMD's CDNA 3 architecture, with 192 GB of HBM3 memory. It started a product line that continued with the 256 GB MI325X and the CDNA 4 based MI350X and MI355X with 288 GB of HBM3E.
The key points
- The MI300X is built from eight compute chiplets stacked on four I/O dies, with 304 compute units and 192 GB of HBM3.
- The MI325X keeps the same compute but raises memory to 256 GB of HBM3E; the CDNA 4 based MI350X and MI355X reach 288 GB at 8 TB/s and add FP4 and FP6.
- AMD's main competitive argument is memory capacity per GPU, while NVIDIA keeps an advantage in software ecosystem maturity and rack-scale NVLink domains.
- ROCm is AMD's open software stack; framework support has broadened, but porting and tuning effort is still higher than on CUDA for some workloads.
The Instinct MI300 and MI350 families
AMD launched the Instinct MI300X on December 6, 2023. At launch, AMD named Microsoft (with its Azure ND MI300X v5 virtual machines), Oracle Cloud Infrastructure and Meta among its customers, along with server makers Dell, HPE, Lenovo and Supermicro. A sister product, the MI300A, combines GPU and CPU chiplets and powers the El Capitan supercomputer at Lawrence Livermore National Laboratory. [9]
Chiplet design of the MI300X
Where NVIDIA's Hopper is a single large die, the MI300X is assembled from many smaller dies. Eight accelerator complex dies (XCDs) hold the compute units, and these are stacked vertically on four I/O dies that hold the system infrastructure and tie the chiplets together through AMD Infinity Fabric. Eight stacks of HBM3 sit around them on the same package. In total the MI300X has 153 billion transistors, built on a mix of TSMC 5 nm and 6 nm FinFET processes. [6][1]
Each XCD has 40 physical compute units, of which 38 are enabled to improve manufacturing yield, giving 304 compute units and 19,456 stream processors across the package. Each XCD has 4 MB of L2 cache, and a 256 MB last-level cache is shared across the device. The MI300A replaces two of the eight XCDs with three CPU chiplets (CCDs), which is how AMD builds an APU from the same parts. [6][1]
The chiplet approach lets AMD mix process nodes, use smaller dies that yield better, and reuse the same building blocks across products. The trade-off is that data has to cross die boundaries, so cache and memory placement affects performance more visibly than on a monolithic GPU.
192, 256 and 288 GB: memory as the main selling point
Memory capacity has been AMD's clearest differentiator. The MI300X has 192 GB of HBM3 on an 8,192-bit interface with 5.3 TB/s of peak bandwidth. At launch AMD contrasted this with the H100's 80 GB and 3.35 TB/s. The MI325X raises capacity to 256 GB of HBM3E at 6 TB/s. The MI350X and MI355X use eight 12-high HBM3E stacks for 288 GB and 8 TB/s, and an eight-GPU MI355X platform holds about 2.3 TB of HBM3E. [1][9][2][7][3]
A model with 405 billion parameters stored in FP8 needs about 405 GB just for its weights. That fits across three 192 GB MI300X GPUs, or two 288 GB MI355X GPUs, before counting the key-value cache. On 80 GB GPUs the same weights would need at least six devices. Fewer GPUs per model instance means less communication between GPUs and more independent replicas per server.
Compute, precision and power
AMD rates the MI300X at 1.3 petaFLOPS of FP16 or BF16 and 2.6 petaFLOPS of FP8 without sparsity, doubling with sparsity, plus 81.7 teraFLOPS of FP64 vector and 163.4 teraFLOPS of FP64 matrix for scientific computing. Its typical board power is 750 W and it uses the OAM module form factor with a PCIe 5.0 x16 host link. Eight Infinity Fabric links at 128 GB/s each connect it to other GPUs. The MI325X lists the same compute figures at a typical board power of 1,000 W. [1][2]
CDNA 4 reorganises the chiplets. The MI350 series still uses eight XCDs, now built on TSMC N3P, but mounts them on two larger I/O dies on N6 instead of four smaller ones. Each XCD has 36 compute units with 32 enabled, giving 256 compute units and 185 billion transistors. Matrix cores gain the OCP microscaling formats MXFP8, MXFP6 and MXFP4, and AMD doubled the execution resources for 16-bit and smaller data types. Local data share per compute unit grows to 160 KB. [7][4]
The MI355X peaks at 10.1 petaFLOPS for MXFP4 and MXFP6, and at 10.1 petaFLOPS FP8 and 5 petaFLOPS FP16 with sparsity, at a 2,400 MHz peak clock and 1,400 W typical board power. The MI350X runs at up to 2,200 MHz and 1,000 W, is designed for passive air-cooled OAM systems, and reaches 9.2 petaFLOPS of MXFP4. AMD says air-cooled racks hold up to 64 GPUs and direct liquid-cooled racks up to 128. Both use seven Infinity Fabric links at 153 GB/s each. [4][5][10]
ROCm: the software side
ROCm is AMD's open software stack for GPU-accelerated computing. It includes the HIP runtime and programming interface, an LLVM-based compiler, math and deep learning libraries such as rocBLAS, hipBLAS, MIOpen and Composable Kernel, the RCCL collective communication library, and profiling and debugging tools. HIPIFY converts CUDA source code to HIP. The major frameworks, including PyTorch, JAX and vLLM, support ROCm. [8]
AMD has iterated quickly on software. At the MI300X launch it introduced ROCm 6 with support for FlashAttention, HIPGraph and vLLM. With the MI350 series in 2025 it released ROCm 7, which AMD says improves inference performance by up to 4 times and training by up to 3 times compared with ROCm 6.0, and it opened a developer cloud for free GPU access. [9][10]
ServeTheHome, covering the MI350 launch, noted that AMD itself acknowledges it has some way to go on software maturity even as it invests heavily. In practice, popular models served through vLLM or trained through PyTorch usually run, while custom CUDA kernels, less common libraries or very new model architectures can need porting and tuning before they perform well. [11]
Where Instinct competes with NVIDIA
On paper the MI355X is positioned against the NVIDIA B200. AMD's own comparison lists the MI355X ahead of the B200 on FP16, FP8 and especially FP6 peak throughput and on FP64, and it sets the MI355X's 288 GB and 8 TB/s against 180 GB and 7.7 TB/s for the B200. Peak figures from a vendor's own comparison should be treated with care, since delivered performance depends heavily on kernels, libraries and the model being run. [3]
The larger structural gap has been scale-up networking. MI300X and MI350 systems connect eight GPUs per node with Infinity Fabric, while NVIDIA's GB200 NVL72 places 72 GPUs in one NVLink domain. ServeTheHome pointed out that a 128-GPU liquid-cooled MI355X rack is built from sixteen eight-GPU trays, a notably different approach from the 72 GPUs in a single NVLink domain of NVL72. AMD's answer is Helios, a rack-scale system with up to 72 MI400-series GPUs and 260 TB/s of scale-up bandwidth. [11][10][6]
In our view Instinct GPUs are most compelling for memory-bound inference of large models, where more HBM per GPU reduces the number of GPUs needed per replica, and for teams already using framework-level code on PyTorch or vLLM. They are a harder fit for workloads that depend on NVIDIA-specific libraries, or for very large training jobs that benefit from a 72-GPU NVLink domain.
The roadmap as of September 2026
AMD previewed the MI400 series with 432 GB of HBM4, 19.6 TB/s of bandwidth and 40 petaFLOPS of FP4 per GPU, planned for 2026 in the Helios rack. In July 2026 AMD presented Helios and the MI455X at an event in San Francisco, and reported deals with Anthropic and Microsoft. TrendForce reported that AMD expected to ship its first Helios racks later in 2026. As of September 2026, MI300X, MI325X and MI350-series GPUs are the Instinct parts most likely to be available to rent. [10][12][13]
Renting AMD Instinct today
Fewer providers list AMD Instinct GPUs than NVIDIA GPUs, and the software images they offer differ, so it pays to check the ROCm version, driver stack and framework builds before committing. If your model is memory-bound, compare the number of GPUs you would need on an MI300X or MI355X with the number needed on an H100, H200 or B200, and compare the total cost of that configuration rather than the hourly rate per GPU. On Kovara, the MI300X 192GB page shows current listings, the GPU prices view lets you compare rates across providers, and Compare lets you place an Instinct GPU next to its NVIDIA alternatives.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1How much memory does the MI300X have compared with the H100?
The MI300X has 192 GB of HBM3 with 5.3 TB/s of peak bandwidth. The H100 SXM has 80 GB of HBM3 with 3.35 TB/s. The larger capacity lets some models run on fewer GPUs.
2What is the difference between MI300X and MI325X?
They use the same CDNA 3 compute design with 304 compute units and the same peak compute figures. The MI325X moves to 256 GB of HBM3E with 6 TB/s of bandwidth and has a higher typical board power of 1,000 W versus 750 W.
3What is the difference between MI350X and MI355X?
Both are CDNA 4 GPUs with 288 GB of HBM3E and 8 TB/s. The MI355X runs at a higher clock and a typical board power of 1,400 W and is the part used in liquid-cooled platforms, reaching 10.1 petaFLOPS of MXFP4. The MI350X is an air-cooled 1,000 W part rated at 9.2 petaFLOPS.
4Does CUDA code run on AMD Instinct GPUs?
Not directly. AMD's ROCm stack uses HIP, a CUDA-like programming interface, and provides HIPIFY tools to translate CUDA source. Most users run standard frameworks such as PyTorch, JAX or vLLM, which have ROCm builds.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- AMD · AMD Instinct MI300X Accelerators ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD · AMD Instinct MI325X Accelerators ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD · AMD Instinct MI350 Series GPUs ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD · AMD Instinct MI355X GPU specifications ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD · AMD Instinct MI350X GPU specifications ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD ROCm Documentation · AMD Instinct MI300 series microarchitecture ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD ROCm Documentation · AMD Instinct MI350 Series microarchitecture ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD ROCm Documentation · What is ROCm? ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- AMD · AMD Delivers Leadership Portfolio of Data Center AI Solutions with AMD Instinct MI300 Series ↗ (opens in a new tab)Press release · Checked 29 September 2026
- AMD · AMD Instinct MI350 Series and Beyond: Accelerating the Future of AI and HPC ↗ (opens in a new tab)Manufacturer blog · Checked 29 September 2026
- ServeTheHome · AMD MI350 and CDNA 4 Architecture Launched with ROCm 7 ↗ (opens in a new tab)News report · Checked 29 September 2026
- Quartz · AMD launches Helios rack system and MI450 GPUs to rival Nvidia ↗ (opens in a new tab)News report · Checked 29 September 2026
- TrendForce · AMD's first rack-scale AI system Helios challenges NVIDIA with HBM4 memory edge ↗ (opens in a new tab)News report · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
