Kovara
Enter Platform
AI workloads

KV cache

Memory that stores intermediate results of attention while a language model generates text, so they don’t have to be recomputed.

Why it matters

The KV cache grows with context length and the number of concurrent users, and often decides how many requests a GPU can serve.

Related terms

Want a deeper answer?

Ask Kova to explain KV cache in the context of your workload. Meet Kova →