AI workloads
KV cache
Memory that stores intermediate results of attention while a language model generates text, so they don’t have to be recomputed.
Why it matters
The KV cache grows with context length and the number of concurrent users, and often decides how many requests a GPU can serve.
Want a deeper answer?
Ask Kova to explain KV cache in the context of your workload. Meet Kova →
