Home › Operate
⚙️ Operate

Deployment, Inference & LLMOps

Running models reliably and affordably at scale - MLOps for LLMs.

OverviewWhat it is

LLMOps covers serving models in production: managing latency, cost, scaling, monitoring, and versioning. It is the discipline that keeps a system alive - and affordable - after launch.

At a glanceDeployment, Inference & LLMOps

Deployment, Inference & LLMOps diagram

LLMOps keeps the system alive and affordable after launch, when real traffic and cost curves hit.

CompareThe serving & inference landscape, side by side

What it takes to run a model in production — the engine that serves it, the tricks that speed it up, and the platform practices around it. Filter by kind or search; most real stacks pick an engine, layer optimizations, and wrap it in ops.

MovedThis grew its own section

The forty-two stages a request passes through used to sit on this page. It outgrew being one table on someone else's topic, so it has its own section now — with a full page per stage.

MechanicsHow it works

Requests flow through an API gateway (auth, routing), a cache (cutting repeat cost and latency), to a model server (managed API or self-hosted with quantization), all wrapped in observability - traces, token/cost tracking, and quality monitoring.

Ground levelWhat you actually build

Serving is a pipeline, not a black box. Every lever that cuts the bill trades away a little quality — the lanes tell you how much.

Serving is a pipeline, not a black box. Every lever that cuts the bill trades away a little quality — the lanes tell you how much.

LandscapeTypes & approaches

Click a highlighted type to open its own page — concept, use case, and diagram.

FeasibilityArchitecture & feasibility

Architecture & feasibility

  • This is the architecture-feasibility layer: cost-per-token times volume, latency budgets, and scaling decide whether a design survives production.
  • Levers: caching, quantization, smaller/routed models, and shorter context - each trades a little quality for large cost/latency wins.
  • Managed API vs self-host is the central call: speed and scale vs control, residency, and potentially lower unit cost at high volume.

In practiceWhat it means for building

Cost-per-interaction and latency are product constraints - they shape pricing, UX, and which features are even viable. Model the unit economics early.

You own hosting choice, autoscaling, quantization, caching, and monitoring - these decide whether the system is sustainable at production traffic.

GlossaryKey terms

CheckCheck your understanding

What drives LLM cost in production?

Tokens (in + out), model size, and volume. Levers: smaller/quantized models, caching, shorter context, and routing easy queries to cheaper models.

Managed API vs self-hosting?

API = fast, scalable, less ops, but data leaves and per-token cost grows. Self-host = control, residency, cheaper at scale, but you own the infra.

How do you make an expensive feature feasible?

Cache, quantize, route to cheaper models, trim context, and set token budgets - then monitor unit economics live.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Kimi K3 is available on Amazon Bedrock with a 1M-token context window, making it callable through managed AWS endpoints.

    China frontier labs · 18 Sep 2026 · source

  • Updated this page A year-long trace study of LLM serving workloads provides evidence on caching and load-balancing behavior.

    arXiv cs.AI · 16 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning