Assertions run plain code checks on a model's output, giving an instant, reproducible pass or fail wherever correctness is programmatically verifiable.
ConceptWhat it is
An assertion or exact-match check is the simplest form of automated evaluation: ordinary code that inspects a model's output and returns true or false. Rather than judging open-ended quality, it verifies a concrete, checkable property — the answer equals an expected string, the response is valid JSON against a schema, a required field is present, or a regular expression matches.
It exists because a large share of LLM work produces structured or verifiable outputs — a classification label, an extracted date, a SQL query, a function-call argument — where there is a genuine right answer. For those cases you do not need an LLM judge or a human reviewer; a deterministic check is faster, cheaper, and fully reproducible, which is exactly what a CI regression gate needs.
How it worksThe mechanics
You define each test case as an input paired with a rule the output must satisfy, run the model to produce a candidate, then apply the assertion — exact equality, substring or contains, regex, JSON-schema validation, numeric tolerance, or a custom function — often after light normalization such as trimming whitespace or lowercasing. Each case resolves to pass or fail with no score in between; results aggregate into a pass rate that gates a merge or deploy, and every failing case is surfaced with its expected-versus-actual diff for fast debugging.
At a glanceSee it
The one abstract ‘match rule’ fans out into a family of concrete checks — exact strings, schema validity, bounded numeric tolerance — each with its own strictness and its own blind spot.
Robust matchers canonicalize both sides — trim, parse, sort keys — before comparing, so the verdict is stable rather than raw-string brittle, though too much normalizing can mask genuine formatting errors.
When to use itWhere it fits
- Outputs with a single correct answer: classification labels, extracted fields, numeric results, or enum values.
- Enforcing structure — valid JSON, a required schema, or a parseable function call — before anything downstream consumes it.
- Fast CI regression gates where every prompt or model change must be checked deterministically at near-zero cost.
- Guarding hard constraints: forbidden words absent, output length within bounds, or a required citation format present.
When NOT to use itLimits & anti-patterns
- Open-ended or subjective outputs — summaries, chat, creative writing — where many wordings are equally correct.
- Anything where valid paraphrases would fail a literal comparison, wrongly penalizing good answers.
- Judging tone, helpfulness, reasoning quality, or faithfulness, which need an LLM judge or human review.
- Tasks with no stable ground truth, where over-tight assertions become brittle and rot as expected outputs shift.
Trade-offsAdvantages & costs
Advantages
- Trivial cost and latency — plain code, no model calls, runs in milliseconds.
- Fully deterministic and reproducible, so a pass or fail means the same thing every run.
- Unambiguous signal that plugs straight into CI as a ship gate.
- Easy to write, read, and debug with a clear expected-versus-actual diff.
Trade-offs & costs
- Only works where correctness is programmatically checkable — a narrow slice of LLM tasks.
- Brittle to formatting: a correct answer phrased differently fails exact match.
- Says nothing about quality, tone, or nuance beyond the literal rule.
- Test sets rot as valid outputs drift, demanding ongoing curation.
ExampleIn the real world
A support-ticket triage feature asks the model to return a JSON object with a category drawn from a fixed enum and an integer priority. The eval suite holds a few hundred labeled tickets; a pytest run sends each through the model and asserts that the output parses as JSON, validates against the schema, and that the category exactly matches the labeled answer, tolerating case. A prompt tweak that raises category accuracy but occasionally wraps the reply in a markdown code fence is caught instantly, because the JSON-parse assertion fails and blocks the merge before it reaches production.
ToolsHow to implement it
- promptfoodeclarative test cases with built-in assertions like equals, contains, is-json, and regex.
- pytestgeneral-purpose harness for writing custom assertion-based eval tests in Python.
- Pydantic / JSON Schemavalidate that structured output conforms to an expected shape and types.
- DeepEvala pytest-style framework for LLM outputs that supports both deterministic assertions and model-graded metrics.
Cost & effortWhat it takes
This is the cheapest evaluation method there is: assertions are ordinary code, add no model or API cost, and execute in milliseconds, so a full suite can run on every commit. The real effort is upfront and ongoing curation — writing correct rules, assembling a labeled golden set, and keeping expected values current as the product changes — plus the judgment to know when a task is genuinely checkable versus when a literal check would wrongly fail valid answers.