Inference is the model actually answering, and it is what you pay for every call.
DefinitionWhat it means
Inference is the process of running a trained model forward on a new input to produce an output, as distinct from training, and in production it is the recurring, per-request compute cost that scales directly with usage.
Why it mattersWhy you should care
Inference cost and latency typically dominate an AI product's unit economics once it has real traffic, so choices like model size, batching, hardware, and prompt length are ongoing engineering and business decisions, not one-time training tradeoffs, and teams scrutinize inference-cost reasoning directly.
At a glanceSee it
Inference runs in two gears — a single parallel prefill of the whole prompt, then a token-by-token decode loop over a growing KV cache, which is why long outputs are the slow, expensive part of every call.
The per-call cost has three levers — shrink the weights, pack more requests into each GPU pass, or skip repeated work with prompt caching and speculative decoding.
Where you see itIn the wild
- Cloud provider bills dominated by inference compute rather than training.
- Latency and throughput dashboards for a live model endpoint.
- Planning discussions on estimating inference cost per user.