Home › Operate
✅ Operate

Evals & Testing

How you know the system is actually good - and staying good.

OverviewWhat it is

Evals are the systematic measurement of quality: accuracy, relevance, safety, and cost. Without them you are shipping on vibes. Robust evals are the single biggest differentiator between a flashy demo and a dependable production system.

At a glanceEvals & Testing

Evals & Testing diagram

No evals = shipping on vibes. A regression set lets you change a prompt or model with confidence.

CompareThe evaluation landscape, side by side

How teams actually measure LLM quality — cheap automated checks, human judgment, and production signals. Filter by kind or search; most real systems layer several.

Every row has a page — what it does, what it costs you, and how to tell when it is the thing biting you.

MechanicsHow it works

Build a golden dataset, run the system, and score outputs - via automatic metrics, an LLM-as-judge, or human review. Wire it into CI so every prompt or model change is tested for regressions, and A/B test in production for real signal.

Ground levelWhat you actually build

A set, a scorer and a gate. Each lane is a place the suite can keep running long after it stopped meaning anything.

A set, a scorer and a gate. Each lane is a place the suite can keep running long after it stopped meaning anything.

LandscapeTypes & approaches

Click a highlighted type to open its own page — concept, use case, and diagram.

FeasibilityArchitecture & feasibility

Architecture & feasibility

  • Evals are infrastructure: a golden dataset, an automated scorer, and a CI gate. This is what makes changing models or prompts feasible without regressions.
  • For RAG, measure retrieval and generation separately (context recall, faithfulness) so you know which half to fix.
  • LLM-as-judge scales grading cheaply but must be calibrated against human labels to be trustworthy.

In practiceWhat it means for building

Define what 'good' means for the user and tie it to metrics. Evals are how you make the ship / no-ship call defensibly instead of arguing anecdotes.

Build eval sets and CI gates so prompt/model swaps are tested, not guessed. Regression suites protect quality as dependencies drift.

Choose a methodWhat your evals dashboard looks like

Evaluation is not one thing. If you grade with a model, with code, by hand, by rubric, by panel or by routing between them, your inputs, outputs, the way you validate the method itself, and the dashboard you need are all different. Pick a method to see the shape it requires.

Pick one

We ran thisThis is the method behind our own kit, so the numbers on this site for it are measured, not illustrative. The shape below is what it needs to contain; the real figures live on the kit’s Evals pages. The judge itself was then validated against an adjudication of all 100 answers — see Judge validation, which scores the graders rather than the system.

Pass rate per criterion, against a baseline

200
cases graded
5
criteria
84%
overall pass
Factual accuracy142 of 200
71%
Hallucination138 of 200
69%
Tone match196 of 200
98%
Completeness176 of 200
88%
SEO readiness184 of 200
92%

How to read it. Each criterion separately, because an average hides the one that is failing. The null-grader baseline sits beside it, or every number reads as good.

Input it needs
  • the question
  • a reference answer
  • the candidate answer
  • a rubric in the prompt, and the model and settings that will judge
Output it gives
  • a verdict per row, with the judge's stated reason
  • an aggregate pass rate
Validate the method byADJUDICATION against human ground truth, then a confusion matrix. Sample the rows where it disagrees with a second grader AND a sample of the rows where they agree — two graders agreeing does not make them right.
Cost measured intokens — a second bill on the same traffic, and every row leaves your network
It cannot produceits own accuracy. A judge cannot score itself.
Use it whenthe criterion needs interpretation — tone, faithfulness, whether a paraphrase counts
Do not use itas the first thing you reach for. It costs money and sends every row to a vendor.

GlossaryKey terms

CheckCheck your understanding

How do you evaluate a GenAI feature?

Build a representative test set, define metrics, automate scoring (metrics + LLM-judge, human-checked), gate releases in CI, and A/B test live.

What's special about evaluating LLMs?

Outputs are open-ended and non-deterministic, so you need rubric-based judging, distributions over many cases, and human calibration.

How do evals make model swaps feasible?

A regression eval set tells you instantly whether a new model or prompt kept quality - so change is safe, not scary.

What changedWhat changed here

MeasuredCounted in the daily brief, with every item listed

As of 2026-09-25 — evaluation and leaderboard coverage in the daily brief: 2 items in the last 7 days

Written inYou approved this and it changed the page
  • Updated this page A controlled clinical benchmark checks whether model confidence tracks evidence quality and uncertainty, providing a way to test calibration before deployment.

    arXiv cs.AI · 17 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

With thanksFollows the approach taught in AI Evals For Engineers & PMs on Maven, by Shreya Shankar and Hamel Husain. Thanks & what I learned →
A living map of modern AI — kept current every morning