GPU Spot pricing: the definition
GPU Spot pricing is a commercial model in which a provider offers qualifying interruptible capacity under its Spot terms. A lower unit price comes with different allocation and interruption conditions; it is not the same promise as a non-interruptible reservation.
The key points
- Spot terms differ between providers, including interruption handling and billing.
- A published discount and an approved quota do not guarantee capacity is available when a job starts.
- Checkpointing, restart time and repeated work belong in the cost comparison.
- Use measured recovery behavior and a deadline-aware plan rather than choosing solely by hourly price.
What is being discounted?
Amazon EC2 describes Spot Instances as a way to use spare capacity under Spot pricing, while Google Cloud describes Spot VMs as excess capacity that it can reclaim. These provider products share an interruptible character but have distinct rules. The discount is attached to a commercial allocation model, not proof that the underlying GPU performs fewer operations while it is running. [1][3]
A useful comparison keeps hardware and service conditions separate. Two offerings can name the same accelerator while differing in interruption risk, minimum allocation, storage persistence or notice behavior. We recommend documenting those differences before calculating savings. The question is not simply which rate is smaller, but whether the workload can produce its required result under the offered operating conditions.
Quota, price and capacity answer different questions
Google's Spot documentation distinguishes resource quota from the resources actually available for allocation. The service can refuse a request when the relevant capacity is unavailable, even when the account has quota. A previous successful allocation also does not establish that the same configuration will be obtainable immediately after an interruption. [3]
Imagine a hypothetical job needing eight GPUs simultaneously. The account is permitted to request eight, and a public page shows a Spot rate, but only four are allocatable at the desired time. Neither the quota nor the price solves that concurrency constraint. A queueing or fallback plan needs to account for when the full required allocation becomes usable, not only what it would cost while running.
Interruption notices are provider-specific
AWS documents best-effort interruption notices, normally two minutes before a Spot instance is stopped or terminated, with a different behavior for hibernation. That is a specific AWS mechanism. It should not be copied into another provider's description or treated as a guarantee that every shutdown always leaves exactly two minutes for arbitrary cleanup. [2]
Google's current Spot documentation describes a best-effort shutdown period of up to 30 seconds and an optional longer preemption-notice setting in Preview. The exact supported configuration matters. This guide does not recommend designing a recovery strategy around an assumed universal notice duration; inspect the service's current contract and test the intended handling path. [3]
For an illustrative failure test, let a checkpoint upload take four minutes under representative load. Receiving a two-minute notice cannot make that upload finish in time. Saving periodically before an interruption, reducing the amount of state saved or changing the architecture could help. A notice is useful operational input, but it is not a replacement for a recoverable workload.
What must survive an interruption?
A training checkpoint can contain parameters, optimizer state and progress needed to resume the intended run. Saving model weights alone may be sufficient for an inference export but insufficient for the complete training state. PyTorch's checkpoint guidance distinguishes these uses. The recovery procedure also needs the referenced data and software environment to remain accessible. [4]
Suppose a hypothetical run saves every 20 minutes and an interruption occurs seven minutes after the latest valid save. At least those seven minutes of progress may need to be repeated, plus restart and loading work. If the checkpoint was stored only on a resource that disappears with the allocation, the actual loss can be far greater. The persistence boundary belongs in the design, not in an assumption about the word disk.
We recommend exercising restore, not merely observing successful save calls. Verify that a fresh allocation can obtain the checkpoint, recreate the required environment and continue correctly. Also decide how to prevent duplicate outputs when a partly completed batch is retried. A restart that corrupts results is not successful fault tolerance even if the process resumes quickly.
A worked cost comparison
Assume a hypothetical non-interruptible job requires ten GPU-hours at $3 per GPU-hour, giving $30 of compute-only cost. A Spot configuration costs $1.20 but consumes 14 billed GPU-hours because of repeated work and restart overhead. Its compute cost is $16.80. The lower rate still wins in this scenario, but savings are 44 percent of the $30 baseline, not the 60 percent suggested by the unit rates alone.
If the same Spot workflow instead consumes 28 GPU-hours, compute cost rises to $33.60 and exceeds the baseline. With these hypothetical rates, break-even billed usage is 25 GPU-hours because $30 divided by $1.20 equals 25. This arithmetic excludes storage, transfers, minimum charges and the value of meeting a deadline. It is a method for comparing assumptions, not a prediction of interruption frequency.
Wall-clock delay can matter more than billed time
A job can wait for capacity without accumulating the same compute charges as running hardware, depending on the service. Nevertheless, that delay affects the project. We recommend measuring time to an accepted result from submission, not only the sum of successful kernel durations. Record provisioning, queueing, initialization, interruptions and recovery as separate components.
Imagine a hypothetical overnight report that must finish before 08:00. Spot compute costs half as much, but repeated capacity waits delay completion until 10:00. Whether that option is acceptable depends on the report's purpose. A deadline-aware design might use Spot early and switch to another allocation when remaining time becomes insufficient. The fallback must be realistic; another published catalogue entry is not reserved capacity.
Which costs continue when compute stops?
Google documents that retained persistent disks can continue accruing storage charges after a Spot VM stops. Compute lifecycle and storage lifecycle are separate. Other providers have their own rules for stopped resources, reserved storage and data transfer. A cost model should therefore list each billable component and its termination condition instead of assuming that stopping one instance ends every related charge. [3]
As an illustrative bookkeeping exercise, separate GPU runtime, host runtime, persistent storage, checkpoint requests and data movement. Mark which quantities grow during retries and which persist while waiting. Then apply the actual quoted rates to those quantities. This reveals hidden costs without inventing a universal markup percentage for Spot workloads.
Match the workflow to interruption tolerance
Independent batch tasks, restartable experiments and work with durable progress can be candidates for interruptible execution. A tightly coupled job can be harder to recover when losing one participant disrupts the whole allocation. These are architectural considerations, not a promise that every batch job is suitable or that every interactive service is unsuitable. The relevant evidence is the actual recovery behavior and service objective.
For a hypothetical collection of a thousand independent render tasks, a durable queue can track completed outputs and retry unfinished items. For one large synchronized training run, recovery may involve obtaining many resources together and restoring distributed state. Both use GPUs, but their interruption costs have different shapes. Evaluate the unit of recoverable work rather than classify workloads only by hardware.
What to check before choosing Spot
Confirm the provider's interruption terms, allocation unit, billing granularity and persistence rules. Test checkpoint and restart behavior with realistic state sizes. Estimate costs using repeated work, not only uninterrupted runtime, and define a deadline or fallback policy. Keep prices dated and distinguish a published listing from a successfully allocated resource.
The central lesson is that Spot is a risk-and-price trade-off attached to a service contract. It can reduce completed-work cost substantially when the workload handles interruptions well. It can also cost more or finish too late when recovery is expensive. A useful marketplace presents both sides of that decision rather than treating the lowest interruptible rate as equivalent to guaranteed continuous access.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1. Does an approved GPU quota prove Spot capacity is available?
No. Quota permits a request within account limits; the provider must still have suitable allocatable resources.
2. What is missing from comparing only Spot and on-demand hourly rates?
Repeated work, restart and checkpoint costs, storage and transfer charges, allocation delays and the workload's deadline or service objective.
3. Why should recovery be tested before relying on interruption notices?
A notice may be best effort and too short for the required state save. A proven checkpoint and restore path is needed independently of the notice.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- AWS · EC2 Spot Instances ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
- AWS · Spot interruption notices ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
- Google Cloud · Spot VMs and preemption ↗ (opens in a new tab)Provider documentation · Checked 28 September 2026
- PyTorch · Saving and loading models ↗ (opens in a new tab)Framework tutorial · Checked 28 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.