Serving models (inference)
Inference runs continuously once you have users. It is usually limited by memory rather than raw compute, and cost depends on how many requests each GPU can serve.
What matters most
Memory capacity
The model weights and the KV cache must fit in VRAM. Long contexts and many concurrent users grow the cache quickly.
Memory bandwidth
Generating text token by token is often bound by how fast memory can be read, so bandwidth strongly affects speed.
Region and latency
Running close to your users lowers latency, and data residency rules may fix the region for you.
Flexible scaling
Traffic rises and falls. Being able to add or remove GPUs matters more than for a fixed training run.
A typical setup
One to eight GPUs per model replica, chosen so weights plus cache fit comfortably. Larger-memory GPUs can serve bigger models on fewer cards.
How it is usually bought
A mix is typical: reserved or committed capacity for baseline load and on-demand for peaks. Spot suits only work that tolerates interruption.
Common mistakes
- Sizing only for the weights and forgetting the KV cache.
- Choosing the cheapest region without checking latency to users.
- Relying on spot capacity for user-facing traffic.
Terms to know
This is general guidance, not a recommendation for a specific provider or price. Check current offerings and confirm availability and terms with the provider.
Ready to look at real options?
See tracked offerings with their sources and dates, or ask Kova to size it for you.
