AI workloads
Inference
Running a trained model to produce answers or predictions for users.
Why it matters
Inference is usually limited by memory and latency rather than raw compute, and runs continuously once you have users.
Related terms
Want a deeper answer?
Ask Kova to explain Inference in the context of your workload. Meet Kova →
