Home › Guardrails & Responsible AI › Constrained / structured decoding
🛡️ · Operate

Constrained / structured decoding

How a grammar or schema masks each generated token so model output always fits a required structure.

In one line

Constrain the decoder so every token it emits keeps the output valid against a schema or grammar — guaranteeing shape, never meaning.

ConceptWhat it is

Constrained decoding (also called structured or guided decoding) makes malformed output impossible by restricting what the model is allowed to generate at each step. A language model produces a probability distribution over its whole vocabulary at every token; normally it can sample anything. Constrained decoding intercepts that distribution and masks any token that would violate a target grammar or schema, so the only reachable sequences are ones that parse. It operates at the decoding layer, below the prompt — a hard constraint the model cannot talk its way out of, unlike an instruction to "reply in JSON."

It exists because prompt-and-parse is unreliable at scale: even a strong model occasionally wraps JSON in a markdown fence, adds a trailing comment, or invents an enum value, and any downstream code that calls parse() breaks on that one response. Constraining the tokens themselves turns "usually well-formed" into "always well-formed." The essential caveat: it controls shape, not semantics — a schema-valid object can still hold wrong or unsafe values.

How it worksThe mechanics

A schema (typically JSON Schema), regular expression, or context-free grammar is compiled once into a finite state machine that describes every legal continuation. During generation the engine tracks the current parser state; before sampling each token it computes the set of tokens that keep the automaton in a valid state and sets the logits of all other tokens to negative infinity. The model then samples normally from what remains, the chosen token advances the parser state, and the loop repeats — mask, sample, advance — until the structure reaches an accepting (complete) state, at which point decoding stops with a guaranteed-parseable result.

At a glanceSee it

Constrained / structured decoding diagram
Constrained / structured decoding diagram 1

A taxonomy of constraint formalisms — from fixed enums to full recursive grammars — and the automata they compile into, where richer structure costs more compile time and slower decoding.

Constrained / structured decoding diagram 2

Why guaranteed shape can still yield degraded content — when the mask deletes the tokens the model most wanted, renormalizing over the survivors forces it into continuations it rated unlikely.

When to use itWhere it fits

  • Downstream code parses the output and a single malformed response breaks the pipeline — you need first-try-valid every time.
  • Tool and function calling, where arguments must match a named signature and types exactly.
  • Classification or routing into a fixed enum, where the model must not invent a label outside the allowed set.
  • High-volume extraction into a stable JSON shape, where retry loops and repair code are a real cost.

When NOT to use itLimits & anti-patterns

  • Free-form generation — prose, brainstorming, open chat — where imposed structure adds nothing and can hurt fluency.
  • Tasks that need the model to reason freely first; constrain only the final answer block, not the thinking.
  • When you do not control the decoder and the provider exposes no structured mode — you fall back to prompt, parse, and retry.
  • When the reliability you actually need is semantic correctness — constrained decoding will not deliver it; use validation and evals.

Trade-offsAdvantages & costs

Advantages
  • Guarantees well-formed, parseable output, eliminating format-related failures and retry loops.
  • Removes most defensive parsing and JSON-repair code from the application layer.
  • Confines outputs to exact enums and types, cutting hallucinated categories and field names.
  • Makes tool-calling and multi-step agent pipelines far more deterministic to integrate.
Trade-offs & costs
  • Controls shape only — a schema-valid response can still contain wrong, fabricated, or unsafe values.
  • Adds per-token decode overhead to build masks; small with modern engines, but nonzero.
  • Requires engineering to author and maintain schemas or grammars as the interface evolves.
  • Over-constraining can fight the model's natural distribution and degrade answer quality; not every provider exposes the capability.

ExampleIn the real world

An accounts-payable team extracts fields from scanned invoices — invoice number, date, line items, and total — into JSON that a billing service ingests. With prompt-only "reply in JSON," a small fraction of responses arrived wrapped in a markdown fence or with a trailing explanatory sentence, and each one threw at parse time, forcing a retry that doubled latency and cost on those cases. Switching to schema-constrained decoding — the provider's structured-output mode for the hosted model, or Outlines when running an open model in-house — made every response parse on the first pass and locked the currency field to a fixed enum. The team still asserts, in application code, that the total equals the sum of line items: constrained decoding guaranteed the shape of the object, never the arithmetic inside it.

ToolsHow to implement it

  • Outlinesopen-source structured-generation library compiling JSON Schema, regex, and grammars into finite state machines for local models.
  • XGrammara fast grammar-constrained decoding engine integrated into serving stacks like vLLM and SGLang for low-overhead structured output.
  • Provider structured modesOpenAI's Structured Outputs enforces a supplied JSON Schema, its JSON mode guarantees valid JSON, and Anthropic tool use constrains arguments to a tool's input schema.
  • Guidanceand llama.cpp GBNF grammars — grammar-driven generation for open-weight models you run yourself.

Cost & effortWhat it takes

Token cost is unchanged — constraining does not add tokens — and the runtime cost is a small per-step masking overhead that mature engines like XGrammar reduce to near-negligible through automaton precomputation and caching; naive implementations can be noticeably slower on complex grammars. The real spend is engineering: authoring schemas or grammars, keeping them in sync with the consuming interface, and adding semantic validation on top since shape-conformance is not correctness. On hosted APIs it is often a single flag; self-hosting a constrained decoder means adopting a library and controlling the serving path, which is why it is available only where you own or are given access to the decoder.

A living map of modern AI — kept current every morning