The key points
- On-demand GPUs are flexible and uninterrupted once running, but availability is best effort and the hourly rate is the highest.
- Spot, preemptible and interruptible GPUs carry the deepest discounts in exchange for the risk of being reclaimed, sometimes with only seconds of warning.
- A discount commitment is not the same as a capacity guarantee: some reservations lock in price only, while capacity blocks and future reservations lock in actual GPUs for set dates.
- Large buyers increasingly sign private multi-year contracts for dedicated clusters, which offer the most certainty but the least flexibility.
The main ways to rent a GPU
Cloud providers sell the same GPU under several commercial models. AWS's documentation, for example, lists on-demand instances, Savings Plans, Reserved Instances, Spot Instances, dedicated hosts and instances, and capacity reservations, including Capacity Blocks for reserving clusters of GPU instances. Google Cloud's guidance for AI clusters similarly distinguishes on-demand, Spot, flex-start, calendar-mode future reservations and longer-term reserved blocks. [1][2]
These options vary along three dimensions: how much you pay per GPU-hour, how likely you are to obtain capacity when you want it, and how likely you are to lose it once running. No model wins on all three. The cheaper the rate, the more risk the buyer usually carries, either of interruption or of paying for capacity that sits idle. Kovara's separate article on spot pricing covers the interruptible end of the spectrum in more depth.
On-demand: pay as you go
On-demand capacity is billed for the time it runs, with no commitment. AWS bills on-demand EC2 instances per second. Once an on-demand machine is running, the provider does not reclaim it; Google describes on-demand VMs as not subject to preemption and running until the user deletes them. [1][2]
The catch for GPUs is obtaining the machine in the first place. Google labels on-demand obtainability as best effort, meaning a request succeeds only if capacity happens to be free at that moment. For popular accelerators in popular regions, that can mean failed launches at busy times, which is why many teams pair on-demand usage with some form of reservation. [2]
Spot, preemptible and interruptible capacity
Spot capacity is spare capacity sold at a discount on the condition that the provider can take it back. Google says its Spot VMs can be discounted by up to 91 percent, including for GPUs, but can be preempted at any time. AWS may interrupt Spot Instances when it needs the capacity, when the Spot price exceeds the buyer's maximum or when request constraints can no longer be met, and gives a two-minute warning. [3][4]
Warning periods differ. Azure provides a best-effort 30-second eviction notice for Spot VMs, offers no SLA for them, and lets users choose whether an evicted VM is deallocated or deleted. Marketplaces can work differently again: Vast.ai's interruptible instances are bid-based, and an instance is stopped if another user bids more for the same machine. [5][12]
Spot suits work that can be checkpointed and resumed: batch inference, hyperparameter sweeps, data processing and training jobs that save state frequently. It is a poor fit for multi-node training runs where losing one node stalls the whole job, or for latency-sensitive production inference without a fallback.
Reserved and committed-use discounts
Reserved or committed pricing trades a term commitment for a lower rate. AWS Savings Plans commit the buyer to a dollar amount of hourly spend for one or three years, with all-upfront, partial-upfront or no-upfront payment. Azure reservations run for one or three years and can be paid upfront or monthly, with self-service refunds capped at 50,000 dollars in a rolling 12-month window. [6][8]
A key detail is whether the commitment reserves hardware or only price. AWS states that Savings Plans do not reserve capacity, and that regional Reserved Instances do not either, whereas zonal Reserved Instances do reserve capacity in a specific availability zone. A buyer who signs a discount commitment expecting guaranteed GPUs may find that the discount applies only to machines they can actually obtain. [6][7]
Specialist GPU clouds often offer shorter reservations. Together AI sells reserved cluster capacity for 1 to 90 days paid upfront and lets customers combine reserved and on-demand nodes in one cluster, while Vast.ai offers discounts of up to 50 percent for reserved rentals that keep on-demand priority. [11][12]
Capacity blocks and future reservations
For training runs of days or weeks, hyperscalers now sell time-boxed GPU clusters. AWS Capacity Blocks for ML let customers reserve up to 64 GPU instances per block, and up to 256 across blocks, up to eight weeks in advance, placed in UltraClusters. Blocks cannot be cancelled and instances are terminated automatically at the end of the reservation. [9]
Capacity Block prices are set by supply and demand at the time of purchase, locked once reserved and charged upfront; Savings Plan and Reserved Instance discounts do not apply to them. Google offers comparable options: calendar-mode future reservations of up to 90 days for up to 80 VMs, flex-start VMs that run for up to seven days when capacity becomes available, and longer reserved blocks arranged through an account team up to a year ahead. [10][2]
Suppose a team needs 32 eight-GPU nodes for a two-week training run next month. On-demand launches might fail on the start date, spot nodes might be reclaimed mid-run, and a one-year commitment would leave most of the term unused. A two-week capacity block or calendar reservation matches the need: the price is known in advance, the cluster is guaranteed for the window, and the team pays nothing after it ends, though it also cannot cancel if plans change.
Private contracts for dedicated capacity
At the largest scale, capacity is sold through negotiated contracts rather than published price lists. Nebius's 2025 agreement with Microsoft, for example, covers dedicated capacity from a specific new data center over multiple years. CoreWeave reports a revenue backlog made up of remaining performance obligations and other amounts expected under committed customer contracts, which stood at 66.8 billion dollars at the end of 2025. [13][14]
Such contracts give the buyer certainty over hardware, location and often networking design, and give the provider revenue it can borrow against. The trade-off is inflexibility: the buyer is committed for years to a hardware generation that newer chips will overtake, and the terms, including penalties, delivery milestones and what happens if capacity arrives late, are negotiated case by case and rarely public.
Choosing a model
Most organisations end up with a mix. A common pattern is a committed baseline sized to steady workloads such as production inference, time-boxed reservations for known training runs, on-demand for bursts and spot for anything that can be interrupted. Google's own guidance frames this choice as a trade-off between obtainability, preemption risk and discount, and AWS documents its purchasing options as a menu of separate choices rather than a single price. [2][1]
The questions to ask any provider are simple but often left unasked: does this commitment reserve capacity or only price, how much warning will I get before an interruption, can I cancel or resell unused time, and what happens if the provider cannot deliver on the start date.
Comparing commercial terms on Kovara
Kovara tracks published GPU prices across roughly 100 providers, so readers can see on the GPU prices page how the same accelerator is priced by different suppliers and, where providers publish them, under different commercial models. Kova can walk through the trade-offs for a particular workload, and buyers who need a time-boxed cluster or a longer commitment can request a quote to confirm real capacity rather than relying on a list price alone.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- AWS · Amazon EC2 billing and purchasing options ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- Google Cloud · Cluster Director: choose a consumption option ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- Google Cloud · Spot VMs ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- AWS · Spot Instance interruption notices ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- Microsoft Learn · About Azure Spot Virtual Machines ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- AWS · What are Savings Plans? ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- AWS · Reserved Instance scope ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- Microsoft Learn · Save compute costs with reservations ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- AWS · Capacity Blocks for ML ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- AWS · Capacity Blocks pricing and billing ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- Together AI Docs · Instant Clusters ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- Vast.ai Documentation · Rental types FAQ ↗ (opens in a new tab)Company documentation · Checked 29 September 2026
- Nebius · Nebius announces multi-billion dollar agreement with Microsoft for AI infrastructure ↗ (opens in a new tab)Company announcement · Checked 29 September 2026
- CoreWeave (SEC filing) · Fourth quarter and fiscal year 2025 results ↗ (opens in a new tab)Company filing · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
