MIG: the definition
Multi-Instance GPU (MIG) is an NVIDIA technology, introduced with the Ampere A100, that partitions one physical data-center GPU into as many as seven isolated GPU instances. Each instance has its own compute, cache and memory resources, so workloads running side by side cannot slow each other down or crash each other.
The key points
- MIG splits one supported NVIDIA GPU into up to seven hardware-isolated instances, each with its own compute, L2 cache and memory bandwidth.
- Instances are chosen from fixed profiles such as 1g.10gb or 3g.40gb on an 80 GB H100, with equivalent profiles for H200 and B200.
- MIG suits many small inference services, notebooks and multi-tenant clusters where a whole GPU would sit mostly idle.
- Time-slicing and MPS share a GPU without full hardware isolation, while vGPU shares it across virtual machines and can be backed by MIG.
What MIG does
Multi-Instance GPU lets an administrator divide a supported NVIDIA GPU into as many as seven separate GPU instances for CUDA applications. NVIDIA introduced it with the A100 in the Ampere generation and describes each instance as securely partitioned, with its own share of the chip rather than a share of time on the whole chip. [1][5]
The isolation runs through the memory system. According to NVIDIA's MIG user guide, each instance is assigned its own on-chip crossbar ports, L2 cache banks, memory controllers and DRAM address buses. Streaming multiprocessors and other engines, such as copy engines and video decoders, are also divided, which gives each instance a defined quality of service and fault isolation. The practical effect is predictable throughput and latency even when a neighbouring instance is fully loaded. [1]
MIG exists because modern data-center GPUs are larger than many individual workloads need. A small model serving modest traffic can leave most of an H100 idle. Partitioning turns one large accelerator into several right-sized ones without giving up the isolation that separate physical GPUs would provide.
Which GPUs support MIG
NVIDIA supports MIG on GPUs from the Ampere generation onward, meaning compute capability 8.0 or higher, but only on selected models. The supported list includes the A100 in 40 GB and 80 GB versions and the A30 among Ampere parts; the H100, H20, H200 and the H100 inside GH200 among Hopper parts; and the B200 and GB200 among Blackwell data-center parts. The A100, Hopper parts, B200 and GB200 allow up to seven instances, while the A30 allows four. [3]
MIG has also reached professional Blackwell cards. NVIDIA lists the RTX PRO 6000 Blackwell with up to four instances and the RTX PRO 5000 and RTX PRO 4500 with up to two. Consumer GeForce cards do not appear in the supported list. [3]
The Hopper generation brought a second version of MIG. NVIDIA says each H100 instance has roughly 3 times the compute capacity and nearly 2 times the memory bandwidth of an A100 instance, adds per-instance performance monitors, and supports confidential computing at the MIG level for multi-tenant environments. [6]
MIG profiles and how partitions are sized
MIG does not allow arbitrary slices. Each GPU exposes a fixed menu of profiles, named for compute and memory. On an 80 GB A100 or H100 the options are 1g.10gb, up to seven per GPU; 2g.20gb, up to three; 3g.40gb, up to two; 4g.40gb, one; and 7g.80gb, which is the whole GPU. A 1g instance receives one-seventh of the streaming multiprocessors and one-eighth of the memory. [4]
Larger-memory GPUs scale the same pattern. The 141 GB H200 offers 1g.18gb, 2g.35gb, 3g.71gb, 4g.71gb and 7g.141gb, while the 180 GB B200 offers 1g.23gb, 2g.45gb, 3g.90gb, 4g.90gb and 7g.180gb. Both also have a 1g profile with a double memory share, 1g.35gb on H200 and 1g.45gb on B200, of which up to four fit on one GPU. [4]
The user guide describes two levels of partitioning. A GPU instance owns a slice of memory and compute; within it, one or more compute instances can subdivide the streaming multiprocessors while sharing that instance's memory. Changing the layout requires that no processes are using the affected GPU, which is why the Kubernetes tooling stops GPU workloads before reconfiguring a node. [2][7]
An H100 80 GB serving several small language models could be split into seven 1g.10gb instances, one model per instance. If the team later needs to run a larger model that requires about 30 GB of weights and cache, it could instead reconfigure the GPU as two 3g.40gb instances, trading instance count for memory per instance.
Where MIG is useful
NVIDIA positions MIG for workloads that do not fill an entire GPU, letting different jobs run in parallel to raise utilization. For cloud providers it adds a multi-tenancy guarantee: one customer's work cannot affect the performance or scheduling of another's. At the A100 launch NVIDIA framed this as up to seven times more GPU instances at no additional hardware cost. [1][5]
Typical uses are inference for small and mid-sized models, development notebooks, continuous-integration test runners and shared research clusters where many users need some GPU but rarely a whole one. MIG works on bare metal and in containers, with GPU pass-through to Linux virtual machines, and with vGPU on supported hypervisors. [1]
MIG also has limits. The user guide notes that MIG instances do not support graphics APIs and cannot use peer-to-peer transfers or NVLink between instances, so it is not a way to build a multi-GPU training job out of slices. Large training runs still want whole GPUs. [2]
MIG in Kubernetes
In Kubernetes, NVIDIA's GPU Operator manages MIG through a component called MIG Manager. Administrators set a node label, nvidia.com/mig.config, to a named layout such as all-1g.10gb; MIG Manager stops GPU workloads, applies the new geometry and restarts the GPU software stack. [7]
Two strategies decide how instances appear to the scheduler. With the single strategy, every GPU on a node is partitioned the same way. With the mixed strategy, GPUs on the same node can have different layouts, and each profile is exposed as its own resource, for example nvidia.com/mig-1g.10gb, which a pod requests in its resource limits. [7]
MIG vs time-slicing, MPS and vGPU
Time-slicing is the simplest way to share a GPU. The GPU Operator can advertise one GPU as several replicas, and workloads scheduled on them interleave in time. NVIDIA is explicit that there is no memory or fault isolation between replicas, and that requesting more replicas does not guarantee a proportional share of compute. Time-slicing can also be layered on top of MIG instances. [8]
The CUDA Multi-Process Service, or MPS, is a runtime service that lets processes from several applications share a GPU cooperatively, improving utilization and reducing context switching, with options to limit memory and streaming multiprocessors per client. It shares a GPU at the software level rather than partitioning the memory system in hardware the way MIG does. [9][1]
NVIDIA vGPU is a different layer: hypervisor software that shares a physical GPU among virtual machines and requires an NVIDIA software licence. vGPUs can be time-sliced, with best-effort, equal-share or fixed-share scheduling, or MIG-backed, in which case each virtual machine gets a MIG instance with hardware isolation underneath. [10]
MIG and choosing compute today
For buyers, MIG changes the unit of comparison. If your model fits in 10 to 20 GB and needs predictable latency, a fraction of an H100 or H200 may be a better fit than a whole older GPU, provided your platform or provider exposes MIG instances. If you run many such services, renting whole GPUs and partitioning them yourself gives you the same isolation with more control. Use Kovara to compare what full GPUs cost across providers, then divide by the number of MIG instances you would actually run to estimate cost per service.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1Which GPUs support MIG?
MIG is supported from the Ampere generation onward on selected data-center and professional GPUs. NVIDIA's list includes the A30, A100, H100, H20, H200, GH200, B200 and GB200, which support up to seven instances, and several RTX PRO Blackwell cards with two to four instances.
2What does a MIG profile name like 1g.10gb mean?
The first part is the compute slice count, out of seven, and the second part is the memory assigned. On an 80 GB H100, 1g.10gb gives one-seventh of the streaming multiprocessors and about 10 GB of memory, and up to seven of them fit on one GPU.
3Is MIG the same as time-slicing?
No. Time-slicing lets several workloads take turns on the whole GPU, with no memory or fault isolation between them. MIG divides the hardware itself, so each instance has dedicated memory, cache and compute and predictable performance.
4Can I use MIG in Kubernetes?
Yes. NVIDIA's GPU Operator can configure MIG on nodes and expose each instance type as a Kubernetes resource, such as nvidia.com/mig-1g.10gb, that pods request like any other GPU.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · MIG User Guide: Introduction ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · Multi-Instance GPU User Guide ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · MIG User Guide: Supported GPUs ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · MIG User Guide: Supported MIG profiles ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Ampere Architecture In-Depth ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA Technical Blog · NVIDIA Hopper Architecture In-Depth ↗ (opens in a new tab)Technical blog · Checked 29 September 2026
- NVIDIA · GPU Operator with MIG ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · GPU Operator: Time-slicing GPUs in Kubernetes ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · Multi-Process Service ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · Virtual GPU Software User Guide ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
