Home › Evals & Testing › Regression suite
✅ · Operate

Regression suite

A curated golden set of cases re-run on every change to block regressions in CI

In one line

Freeze a set of known-good cases and re-run them automatically on every change so any drop in quality fails the build before users see it.

ConceptWhat it is

A regression suite is a curated collection of representative cases — a golden set of inputs paired with known-good expected outputs or assertions — that you re-run automatically every time something changes: a prompt edit, a model version bump, a retrieval or chunking tweak, or ordinary application code. Its single job is to catch regressions: silent quality drops where a change that helps one case quietly breaks ten others. Because LLM behavior is non-deterministic and globally coupled, a fix you eyeball as safe often is not, and manual spot-checking cannot cover the surface area.

The pattern is borrowed directly from software testing, adapted for graded rather than exact outputs. Rather than demanding byte-identical answers, each case carries a check — exact match, a regex or JSON-schema assertion, a similarity threshold, or an LLM-as-judge rubric — and the suite passes only when the aggregate score clears a stored baseline. Wired into CI, it becomes a ship-gate: no change merges until the golden set holds.

How it worksThe mechanics

You start by assembling a golden set from real traffic, past incidents, and known edge cases, and you record the expected behavior for each as a machine-checkable assertion. On every pull request or model change, CI runs each case through the current system, scores the output against its check, and compares the aggregate pass rate against the stored baseline. If the suite holds, the change is green and can merge; if a case that used to pass now fails, the build goes red and the diff points at exactly which inputs broke. From there you either fix the regression or, when the new behavior is genuinely correct, deliberately re-bless the baseline so it becomes the new expectation.

At a glanceSee it

Regression suite diagram
Regression suite diagram 1

Check types form a spectrum — from cheap brittle exact matches to costly tolerant graded judges, and most real suites mix both.

Regression suite diagram 2

A failing case is not automatically a bug — triage separates flakiness, a true regression, and an intended change that should re-bless the baseline.

When to use itWhere it fits

  • You have a maturing system with recurring, well-understood tasks where correct behavior is stable enough to freeze as expected outputs
  • You change prompts, models, or retrieval frequently and need confidence that each change does not silently break past wins
  • You want an objective, automated merge gate rather than relying on manual spot-checks before every deploy
  • Past bugs and incidents keep recurring — each becomes a permanent case so the same regression cannot ship twice

When NOT to use itLimits & anti-patterns

  • The product is still in rapid discovery where the target behavior shifts weekly — you would spend all your time rewriting baselines
  • Outputs are open-ended or creative with no stable notion of correct, so any fixed expectation is arbitrary
  • You need to measure absolute quality or discover unknown failure modes — that is exploratory evaluation, not regression guarding
  • The team will not commit to curating the set — a stale, unmaintained suite gives false confidence and blocks merges for the wrong reasons

Trade-offsAdvantages & costs

Advantages
  • Catches silent regressions before users do, turning a vague sense that things seem fine into a measured pass or fail
  • Fully automated and repeatable, so it scales to every commit without consuming human reviewer time
  • Encodes institutional memory — every past incident becomes a permanent guard against its own recurrence
  • Makes changes auditable: a red build names the exact cases that broke, shrinking debugging effort
Trade-offs & costs
  • Rots without active curation — cases drift from reality, expected outputs go stale, and the gate loses meaning
  • Only guards the behaviors you thought to include; it is blind to failure modes absent from the set
  • Graded checks add cost and flakiness — LLM-judge scoring is itself non-deterministic and can wobble around the threshold
  • A brittle suite that fails on cosmetic differences trains the team to rubber-stamp or disable it

ExampleIn the real world

A support-automation team runs a customer-facing assistant over a knowledge base. They assemble roughly two hundred golden cases — real questions each with a graded rubric: refund answers must cite the policy, out-of-scope questions must defer to a human, and none may leak internal notes. The suite runs in CI on every prompt change and model upgrade. When they trial a newer model, the aggregate looks better, but the suite goes red on eleven cases: the new model has started answering billing questions it should escalate. Because the regression is caught in the pull request rather than in production, they tighten the escalation instruction, watch the eleven cases go green again, and only then merge. A case from a prior incident — where the assistant once invented a discount code — stays in the set permanently, guaranteeing that specific failure can never ship again.

ToolsHow to implement it

  • promptfoo — declarative YAML test cases with assertions, similarity, and LLM-judge checks, built to run in CI
  • DeepEval — pytest-style LLM evaluation with metrics and regression-friendly assertions
  • LangSmith and Braintrust — hosted datasets, scoring, and baseline comparison across runs and model versions
  • Standard CI plus a snapshot library such as pytest or Jest for the exact-match and schema portions of the gate

Cost & effortWhat it takes

The recurring cost is CI compute on every change, which for LLM-judge cases means real API spend that scales with suite size times commit frequency — keep the set focused and reserve expensive graded checks for the cases that need them. The larger and easily underestimated cost is human curation: someone must add cases from new incidents, retire stale ones, and re-bless baselines when behavior legitimately changes. A regression suite is cheap to run and moderately expensive to keep honest; its failure mode is not compute but neglect, where an un-maintained set rots into either noise everyone ignores or a gate that blocks the wrong things.

A living map of modern AI — kept current every morning