OverviewWhat it is
LLMOps covers serving models in production: managing latency, cost, scaling, monitoring, and versioning. It is the discipline that keeps a system alive - and affordable - after launch.
At a glanceDeployment, Inference & LLMOps
LLMOps keeps the system alive and affordable after launch, when real traffic and cost curves hit.
CompareThe serving & inference landscape, side by side
What it takes to run a model in production — the engine that serves it, the tricks that speed it up, and the platform practices around it. Filter by kind or search; most real stacks pick an engine, layer optimizations, and wrap it in ops.
Every row has a page — what it does, what it costs you, and how to tell when it is the thing biting you.
MovedThis grew its own section
The forty-two stages a request passes through used to sit on this page. It outgrew being one table on someone else's topic, so it has its own section now — with a full page per stage.
MechanicsHow it works
Requests flow through an API gateway (auth, routing), a cache (cutting repeat cost and latency), to a model server (managed API or self-hosted with quantization), all wrapped in observability - traces, token/cost tracking, and quality monitoring.
Ground levelWhat you actually build
Serving is a pipeline, not a black box. Every lever that cuts the bill trades away a little quality — the lanes tell you how much.
LandscapeTypes & approaches
Click a highlighted type to open its own page — concept, use case, and diagram.
FeasibilityArchitecture & feasibility
Architecture & feasibility
- This is the architecture-feasibility layer: cost-per-token times volume, latency budgets, and scaling decide whether a design survives production.
- Levers: caching, quantization, smaller/routed models, and shorter context - each trades a little quality for large cost/latency wins.
- Managed API vs self-host is the central call: speed and scale vs control, residency, and potentially lower unit cost at high volume.
In practiceWhat it means for building
Cost-per-interaction and latency are product constraints - they shape pricing, UX, and which features are even viable. Model the unit economics early.
You own hosting choice, autoscaling, quantization, caching, and monitoring - these decide whether the system is sustainable at production traffic.
GlossaryKey terms
CheckCheck your understanding
What drives LLM cost in production?
Tokens (in + out), model size, and volume. Levers: smaller/quantized models, caching, shorter context, and routing easy queries to cheaper models.
Managed API vs self-hosting?
API = fast, scalable, less ops, but data leaves and per-token cost grows. Self-host = control, residency, cheaper at scale, but you own the infra.
How do you make an expensive feature feasible?
Cache, quantize, route to cheaper models, trim context, and set token budgets - then monitor unit economics live.
What changedWhat changed here
Updated this page Kimi K3 is available on Amazon Bedrock with a 1M-token context window, making it callable through managed AWS endpoints.
Updated this page A year-long trace study of LLM serving workloads provides evidence on caching and load-balancing behavior.
Three kinds of claim, strongest first. Signal runs every morning.