Home › Evals & Testing › Assertion / exact-match
✅ · Operate

Assertion / exact-match

Assertion and exact-match: cheap deterministic code checks that pass or fail a verifiable output

In one line

Assertions run plain code checks on a model's output, giving an instant, reproducible pass or fail wherever correctness is programmatically verifiable.

ConceptWhat it is

An assertion or exact-match check is the simplest form of automated evaluation: ordinary code that inspects a model's output and returns true or false. Rather than judging open-ended quality, it verifies a concrete, checkable property — the answer equals an expected string, the response is valid JSON against a schema, a required field is present, or a regular expression matches.

It exists because a large share of LLM work produces structured or verifiable outputs — a classification label, an extracted date, a SQL query, a function-call argument — where there is a genuine right answer. For those cases you do not need an LLM judge or a human reviewer; a deterministic check is faster, cheaper, and fully reproducible, which is exactly what a CI regression gate needs.

How it worksThe mechanics

You define each test case as an input paired with a rule the output must satisfy, run the model to produce a candidate, then apply the assertion — exact equality, substring or contains, regex, JSON-schema validation, numeric tolerance, or a custom function — often after light normalization such as trimming whitespace or lowercasing. Each case resolves to pass or fail with no score in between; results aggregate into a pass rate that gates a merge or deploy, and every failing case is surfaced with its expected-versus-actual diff for fast debugging.

At a glanceSee it

Assertion / exact-match diagram
Assertion / exact-match diagram 1

The one abstract ‘match rule’ fans out into a family of concrete checks — exact strings, schema validity, bounded numeric tolerance — each with its own strictness and its own blind spot.

Assertion / exact-match diagram 2

Robust matchers canonicalize both sides — trim, parse, sort keys — before comparing, so the verdict is stable rather than raw-string brittle, though too much normalizing can mask genuine formatting errors.

When to use itWhere it fits

  • Outputs with a single correct answer: classification labels, extracted fields, numeric results, or enum values.
  • Enforcing structure — valid JSON, a required schema, or a parseable function call — before anything downstream consumes it.
  • Fast CI regression gates where every prompt or model change must be checked deterministically at near-zero cost.
  • Guarding hard constraints: forbidden words absent, output length within bounds, or a required citation format present.

When NOT to use itLimits & anti-patterns

  • Open-ended or subjective outputs — summaries, chat, creative writing — where many wordings are equally correct.
  • Anything where valid paraphrases would fail a literal comparison, wrongly penalizing good answers.
  • Judging tone, helpfulness, reasoning quality, or faithfulness, which need an LLM judge or human review.
  • Tasks with no stable ground truth, where over-tight assertions become brittle and rot as expected outputs shift.

Trade-offsAdvantages & costs

Advantages
  • Trivial cost and latency — plain code, no model calls, runs in milliseconds.
  • Fully deterministic and reproducible, so a pass or fail means the same thing every run.
  • Unambiguous signal that plugs straight into CI as a ship gate.
  • Easy to write, read, and debug with a clear expected-versus-actual diff.
Trade-offs & costs
  • Only works where correctness is programmatically checkable — a narrow slice of LLM tasks.
  • Brittle to formatting: a correct answer phrased differently fails exact match.
  • Says nothing about quality, tone, or nuance beyond the literal rule.
  • Test sets rot as valid outputs drift, demanding ongoing curation.

ExampleIn the real world

A support-ticket triage feature asks the model to return a JSON object with a category drawn from a fixed enum and an integer priority. The eval suite holds a few hundred labeled tickets; a pytest run sends each through the model and asserts that the output parses as JSON, validates against the schema, and that the category exactly matches the labeled answer, tolerating case. A prompt tweak that raises category accuracy but occasionally wraps the reply in a markdown code fence is caught instantly, because the JSON-parse assertion fails and blocks the merge before it reaches production.

ToolsHow to implement it

  • promptfoodeclarative test cases with built-in assertions like equals, contains, is-json, and regex.
  • pytestgeneral-purpose harness for writing custom assertion-based eval tests in Python.
  • Pydantic / JSON Schemavalidate that structured output conforms to an expected shape and types.
  • DeepEvala pytest-style framework for LLM outputs that supports both deterministic assertions and model-graded metrics.

Cost & effortWhat it takes

This is the cheapest evaluation method there is: assertions are ordinary code, add no model or API cost, and execute in milliseconds, so a full suite can run on every commit. The real effort is upfront and ongoing curation — writing correct rules, assembling a labeled golden set, and keeping expected values current as the product changes — plus the judgment to know when a task is genuinely checkable versus when a literal check would wrongly fail valid answers.

A living map of modern AI — kept current every morning