A grammar makes invalid output impossible but cannot make the model know the answer, so a constrained model that is wrong fills the shape with confident nonsense.
Why you'd careThe thing you have already noticed
Your service parses the model's JSON and throws a few times a day: a trailing comma, a sentence before the opening brace, a truncated closing bracket. You add retries. Then you switch on a JSON schema and the parse errors drop to exactly zero — not nearly zero, zero — because the server can no longer emit a token that would break the schema. A week later someone notices the hard classification cases have got worse: the model now always returns one of your three labels, including on inputs where the honest answer was none of them. Both effects come from this stage. Format compliance is enforced by erasure, and erasure has no opinion about correctness.
In and outWhat goes in, what comes out
| In | A grammar in GBNF or EBNF, a JSON Schema compiled to one, a regular expression, or an explicit choice list; the automaton's current state; and the current logits vector of length vocab_size. |
|---|---|
| Process | Look up which vocabulary ids are legal from this automaton state — normally precomputed as a bitmask per state when the grammar was compiled — set every illegal logit to −inf, sample from the survivors, then advance the automaton by the characters of the chosen token. Per-step cost is a mask AND, not a grammar walk. |
| Out | A masked logits vector, one sampled legal token id, and the advanced automaton state. Everything downstream, including temperature and top_p, sees the mask as though the model itself had assigned zero probability to the erased tokens. |
Relative ordering and magnitude among the legal tokens survive, so temperature and truncation still work inside the allowed set. What is destroyed is the model's signal that it wanted to say something else. If the model would have put 0.98 on “this document does not say” and your choice set is {yes, no}, the mask deletes that 0.98 and renormalises the remaining 0.02 into an answer that looks confident. Constrained decoding erases the abstention signal, and nothing downstream can recover it unless you log the pre-mask logprobs.
ConceptThe idea underneath
The model defines a probability distribution over token sequences. A grammar defines a formal language L. Constrained decoding samples from the model conditioned on the prefix staying inside the prefixes of L: p'(t | prefix) = p(t | prefix) * 1[prefix + t is a valid prefix of L] / Z. Worth noticing that this is not the same distribution you would get by sampling freely and rejecting invalid outputs. Rejection sampling gives the correctly conditioned distribution; greedy left-to-right masking gives a locally normalised approximation of it. That gap is the source of the cornering failure — the mask can walk you into a state where every continuation the model finds plausible has already been erased.
The engineering problem is the vocabulary. Naively, deciding which of 128,000 token ids are legal from state q means running the automaton over every token's characters, every step. Willard and Louf's contribution was to invert that: index the vocabulary against the automaton once, at compile time, producing a legal-token bitmask per state. Per-step cost collapses to a lookup and a bitwise AND, which is why constrained decoding is essentially free at runtime and expensive only at compile time.
Context-free grammars need a pushdown automaton, so the state is a pair of FSM state and stack. Engines like XGrammar handle this by splitting the vocabulary into a context-independent portion, maskable entirely at compile time, and a small context-dependent remainder checked against the stack at runtime, with a persistent cache over stack tops.
The residual problem is token-boundary misalignment: the grammar is over characters, the model is over tokens, and one token can straddle a grammar boundary — ", is frequently a single token. An engine that aligns badly produces valid strings the model would never have written, which is a distribution shift it has never seen.
At a glanceSee it
The grammar compiles to an automaton whose state selects a legal-token bitmask applied at every step.
The knobsHyperparameters and nuance
- response_format / json_schema(OpenAI-compatible), structured_outputs with
json/regex/choice/grammar(vLLM), grammar in GBNF (llama.cpp) — which constraint form you supply. vLLM renamed this whole surface: the olderguided_json/guided_regex/guided_choice/guided_grammarrequest fields have been folded into thestructured_outputsobject and are gone from current releases. Pin the parameter names to your server version rather than to a blog post. - structured-outputs backend(vLLM:
auto,xgrammar,guidance,outlines,lm-format-enforcer) — now set as thebackendkey inside--structured-outputs-config; the standalone--guided-decoding-backendflag it replaced no longer exists, so an old runbook fails at startup rather than degrading quietly. Backends differ in compile time, CFG support, regex dialect, and which JSON Schema keywords they honour. This is implementation-specific and moves between releases; verify rather than assume. - whitespace_pattern(vLLM
structured_outputs, from Outlines), withdisable_any_whitespaceas the blunt version — controls whether the grammar admits arbitrary whitespace. Permissive settings let the model burn output tokens on indentation; restrictive settings are a common way to corner it. - schema strictness itself
additionalProperties: false,required,enum,maxLength. This is the real knob. Anenumof three values removes the option to abstain; amaxLengthshorter than the answer forces truncation that looks like a model failure. - tool_choicethe same masking mechanism applied to tool calls, but the two providers spell it differently and the values are not interchangeable. OpenAI takes
none,auto,required, or a named function. Anthropic takes atypeofauto,any,tool(with a name), ornone— there is norequiredon Anthropic, andanyis the equivalent. Either way, forcing a call removes the model's option to answer in prose, including when calling a tool is the wrong move. - constraint start positionwhether the grammar applies from token 0 or only after a free-form block. Constraining from token 0 is the single most common cause of accuracy loss under structured output.
EffectHow this stage moves the answer
Switching it on removes an entire class of error: no preamble, no markdown fence around the JSON, no retry loop. What degrades is any reasoning the schema forbids. If your schema is {answer: string} and the mask forces the model into the answer field at token 0, it has nowhere to think, and accuracy on arithmetic and multi-hop questions falls measurably. The structural fix is to put a reasoning field before answer in the schema, or to run the constraint on a second turn after a free-form first one. The other degradation is manufactured confidence: a forced enum converts every “I don't know” into a wrong label at exactly the rate your base uncertainty was. And a grammar tight enough to leave no plausible continuation produces filler — the model emits padding that satisfies the shape because everything meaningful was masked away.
EvalsWhat it does to your measurements
Format validity becomes 100% by construction, so any metric that was secretly measuring parse success jumps and tells you nothing. The number to watch instead is accuracy conditional on a valid parse, compared before and after. Tam et al. measured degradation under strict format restriction, with the largest gaps on reasoning-heavy tasks. Two further traps. First, a harness that scores unparseable output as wrong will report a large gain from constraints that is entirely a scoring artefact — report the pre-constraint parse-failure rate separately so the gain can be attributed. Second, grammar compilation is not free: a large schema can add hundreds of milliseconds to time-to-first-token on its first use, which pollutes latency benchmarks unless the grammar cache is warm. Constrained runs also look more deterministic than they are, because the mask hides variance inside the allowed set rather than removing it.
Failure modesWhen it goes wrong
- The model emits an opening brace and then padding until max_tokensthe schema cornered it; every continuation it considered plausible was masked, so it emits whatever is left.
- An enum classifier never returns the “unknown” caseyou did not put it in the enum, so the mask deleted abstention and renormalised the uncertainty into a confident wrong label.
- Time-to-first-token spikes on the first request with a new schemagrammar compilation on a cold cache; it disappears on the second request and reappears on every deploy.
- A JSON Schema keyword is silently ignoredseveral backends support only a subset (regex
pattern,$ref, and numeric bounds are common gaps) and drop unsupported keywords rather than rejecting the request. Provider- and version-specific. - Non-Latin text is mangled under a regex constraintthe constraint is expressed over characters while the model produces byte-level tokens, and byte-fallback tokens get masked inconsistently.
PapersWhere this comes from
- Efficient Guided Generation for Large Language ModelsWillard & Louf, 2023 (arXiv:2307.09702). Reframed guided generation as indexing the vocabulary against a finite-state machine ahead of time, reducing the per-token cost from a vocabulary-sized scan to a mask lookup. This is why constrained decoding is now effectively free at runtime.
- Grammar-Constrained Decoding for Structured NLP Tasks without FinetuningGeng et al., 2023. Showed that CFG-constrained decoding can match or beat task-specific fine-tuning on several structured prediction tasks, which is the argument for using a grammar instead of training a model to be well-behaved.
- XGrammar: Flexible and Efficient Structured Generation Engine for Large Language ModelsDong et al., 2024. Introduced the context-independent versus context-dependent token split and a persistent pushdown-stack cache, making context-free grammars practical at serving speed. It is the default backend in several inference servers.
- Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language ModelsTam et al., 2024. Measured accuracy under increasingly strict output-format constraints and found consistent degradation, largest on reasoning tasks. It is the evidence behind “let it reason first, constrain second”.