Home › Guardrails & Responsible AI › Schema / output validation
🛡️ · Operate

Schema / output validation

Validating model output against a schema or types, and re-asking when it does not conform.

In one line

Schema validation checks that a model's output matches a required structure and re-asks the model until it does or a retry budget runs out.

ConceptWhat it is

Schema validation (also called output validation or structured-output enforcement) is the guardrail that checks a model's generated output against a declared schema — a JSON Schema, a set of typed fields, or a data model — before any downstream code trusts it. It exists because an LLM emits free-form text: even when you ask for JSON, it can wrap the braces in prose, drop a required field, put a string where a number belongs, or invent an enum value you never defined. Any code that parses that output as if it were a clean API response will break.

The pattern turns a probabilistic generator into something a program can consume safely by making the contract explicit and machine-checkable. When output fails the check, the system does not pass it downstream — it re-asks the model with the validation errors attached, repairs the output, or falls back to a default. This is distinct from constrained decoding, which prevents invalid tokens during generation rather than checking after the fact.

How it worksThe mechanics

You define the target shape once — a Pydantic model, a Zod schema, or raw JSON Schema — and hand it to a validation layer wrapped around the model call. The model generates, the layer parses the raw string and validates each field against its type and constraints, and a policy branch decides what happens next: conforming output is coerced into a typed object and returned to the caller, while non-conforming output triggers a re-ask in which the specific validation errors are appended to a follow-up prompt so the model can correct itself. The loop repeats up to a capped number of retries, after which the system raises an error or returns a safe default rather than looping forever.

At a glanceSee it

Schema / output validation diagram
Schema / output validation diagram 1

Ways to obtain structured output, ranked weakest to strongest — from a polite prompt up to constrained decoding, where an invalid token cannot be emitted at all.

Schema / output validation diagram 2

The failure modes hidden behind a single no branch, and the boundary where schema checks end and semantic validation must take over.

When to use itWhere it fits

  • The output feeds another system — an API call, a database write, a function argument, a UI component — where one malformed field breaks the next step.
  • You need specific typed fields (enums, dates, numbers, nested objects) rather than free-form prose.
  • Building agent or tool-calling pipelines where each step's output is the next step's input and must be machine-parseable.
  • Extraction or classification tasks where the answer must land in a fixed set of categories or fields.

When NOT to use itLimits & anti-patterns

  • The output is free-form prose meant for a human to read, where a rigid structure is not required.
  • The provider already offers native guaranteed-schema mode or constrained decoding, which makes post-hoc re-asks rare — you still validate, but the loop almost never fires.
  • Latency-critical paths where you cannot afford the extra round-trips a re-ask loop can add.
  • The failure you actually fear is semantic (wrong facts), not structural — schema validation confirms shape, never truth.

Trade-offsAdvantages & costs

Advantages
  • Turns unreliable text into a typed object your code consumes directly, removing defensive parsing scattered everywhere.
  • Cheap in the common case: a valid response costs only one extra local validation pass, no additional model call.
  • Errors are caught at the boundary, close to the model, instead of surfacing deep inside downstream logic.
  • The schema doubles as living documentation and as the prompt's implicit spec, keeping intent in one place.
Trade-offs & costs
  • Re-ask loops add tokens, latency, and cost — each retry is another full generation.
  • Validation confirms structure, not correctness: a well-formed object can still hold wrong values.
  • Complex or deeply nested schemas raise the failure rate and the number of re-asks needed to converge.
  • A model that keeps missing the same constraint can burn the entire retry budget and still return nothing usable.

ExampleIn the real world

A team building an invoice-processing pipeline defines a Pydantic model with fields for vendor, invoice number, a list of line items, and a total typed as a decimal. They use Instructor to wrap the model call so the extraction is validated against that model. On a scanned invoice where the model returns the total as the string "1,240.00" with a comma, validation fails on the decimal field; Instructor automatically re-asks with the parse error attached, and the model returns 1240.00. The downstream ledger code only ever receives a validated object, so it never has to guard against a stray comma or a missing field.

ToolsHow to implement it

  • PydanticPython type-annotated data models; the de facto validation layer for parsed LLM output.
  • Instructorpatches LLM clients to return validated Pydantic or Zod objects, with automatic re-asks on failure.
  • Guardrails AIa validator framework with a re-ask loop and a library of field-level checks.
  • Outlinesconstrained decoding that forces token-level conformance to a JSON Schema or regex during generation.

Cost & effortWhat it takes

Cheap by default and expensive on the tail: a conforming response adds only a local validation pass with no extra tokens, but every re-ask is a fresh generation, so cost and latency scale with how often the model misses the schema. Engineering effort is low to start — define a model, wrap the call — and grows with schema complexity, retry-policy tuning, and fallback handling. Native structured-output modes and constrained decoding push the re-ask tail toward zero, at the price of provider lock-in or self-hosting.

A living map of modern AI — kept current every morning