OverviewWhat it is
Evals are the systematic measurement of quality: accuracy, relevance, safety, and cost. Without them you are shipping on vibes. Robust evals are the single biggest differentiator between a flashy demo and a dependable production system.
At a glanceEvals & Testing
No evals = shipping on vibes. A regression set lets you change a prompt or model with confidence.
CompareThe evaluation landscape, side by side
How teams actually measure LLM quality — cheap automated checks, human judgment, and production signals. Filter by kind or search; most real systems layer several.
Every row has a page — what it does, what it costs you, and how to tell when it is the thing biting you.
MechanicsHow it works
Build a golden dataset, run the system, and score outputs - via automatic metrics, an LLM-as-judge, or human review. Wire it into CI so every prompt or model change is tested for regressions, and A/B test in production for real signal.
Ground levelWhat you actually build
A set, a scorer and a gate. Each lane is a place the suite can keep running long after it stopped meaning anything.
LandscapeTypes & approaches
Click a highlighted type to open its own page — concept, use case, and diagram.
FeasibilityArchitecture & feasibility
Architecture & feasibility
- Evals are infrastructure: a golden dataset, an automated scorer, and a CI gate. This is what makes changing models or prompts feasible without regressions.
- For RAG, measure retrieval and generation separately (context recall, faithfulness) so you know which half to fix.
- LLM-as-judge scales grading cheaply but must be calibrated against human labels to be trustworthy.
In practiceWhat it means for building
Define what 'good' means for the user and tie it to metrics. Evals are how you make the ship / no-ship call defensibly instead of arguing anecdotes.
Build eval sets and CI gates so prompt/model swaps are tested, not guessed. Regression suites protect quality as dependencies drift.
Choose a methodWhat your evals dashboard looks like
Evaluation is not one thing. If you grade with a model, with code, by hand, by rubric, by panel or by routing between them, your inputs, outputs, the way you validate the method itself, and the dashboard you need are all different. Pick a method to see the shape it requires.
We ran thisThis is the method behind our own kit, so the numbers on this site for it are measured, not illustrative. The shape below is what it needs to contain; the real figures live on the kit’s Evals pages. The judge itself was then validated against an adjudication of all 100 answers — see Judge validation, which scores the graders rather than the system.
Pass rate per criterion, against a baseline
How to read it. Each criterion separately, because an average hides the one that is failing. The null-grader baseline sits beside it, or every number reads as good.
| Input it needs |
|
| Output it gives |
|
| Validate the method by | ADJUDICATION against human ground truth, then a confusion matrix. Sample the rows where it disagrees with a second grader AND a sample of the rows where they agree — two graders agreeing does not make them right. |
| Cost measured in | tokens — a second bill on the same traffic, and every row leaves your network |
| It cannot produce | its own accuracy. A judge cannot score itself. |
| Use it when | the criterion needs interpretation — tone, faithfulness, whether a paraphrase counts |
| Do not use it | as the first thing you reach for. It costs money and sends every row to a vendor. |
TEMPLATE — we did not run this. Every figure below is illustrative. Our own testing used an LLM as judge; this panel exists so you can see what a deterministic code grader dashboard would need to contain before you build one.
Score distribution, and where the cut was made
How to read it. The threshold is only meaningful beside the sweep it came from. A single tuned number shown alone is a test score you tuned on.
| Input it needs |
|
| Output it gives |
|
| Validate the method by | A THRESHOLD SWEEP, published — the score distribution and where the cut was made. A threshold shown without its sweep is a tuned number presented as a finding. Check both tails by hand: what it passes and what it fails. |
| Cost measured in | free — runs in-process, sends nothing, repeats exactly |
| It cannot produce | any judgement about meaning. It cannot recognise a paraphrase. |
| Use it when | the data cannot leave your network, or you need the same verdict twice |
| Do not use it | as a correctness verdict on free text. It measures wording, not meaning. |
TEMPLATE — we did not run this. Every figure below is illustrative. Our own testing used an LLM as judge; this panel exists so you can see what a human review dashboard would need to contain before you build one.
Do your reviewers agree with each other?
| Reviewer | Rows read | Minutes/row | Said acceptable |
|---|---|---|---|
| A | 60 | 4.0 | 82% |
| B | 60 | 3.2 | 88% |
| C | 40 | 5.1 | 75% |
How to read it. One reviewer is a sample of one opinion, so the second reader IS the validation. Raw agreement flatters — report it corrected for chance. Cost is in MINUTES.
| Input it needs |
|
| Output it gives |
|
| Validate the method by | INTER-RATER AGREEMENT. Have a second reviewer grade an overlapping sample and report how often they agree, corrected for chance. A single reviewer cannot be validated at all — the second reader IS the validation. |
| Cost measured in | reviewer-minutes — minutes per row times the number of reviewers. It does not get cheaper at scale. |
| It cannot produce | a rate on its own. One reviewer is a sample of one opinion. |
| Use it when | you need ground truth, or the criterion is genuinely a matter of judgement |
| Do not use it | on every row at volume. Use it on a sample and let a cheaper method triage. |
TEMPLATE — we did not run this. Every figure below is illustrative. Our own testing used an LLM as judge; this panel exists so you can see what a weighted rubric review dashboard would need to contain before you build one.
Which criterion actually moved the total
| Criterion | Weight | Score /5 | Contribution |
|---|---|---|---|
| Accuracy | 0.40 | 4.2 | 1.68 |
| Grounding | 0.30 | 3.6 | 1.08 |
| Tone | 0.20 | 4.5 | 0.90 |
| Brevity | 0.10 | 3.1 | 0.31 |
How to read it. The weights must be written down BEFORE scoring, and the page must say whether the ranking survives re-weighting. A ranking that flips on a 0.1 change is a property of the weights.
| Input it needs |
|
| Output it gives |
|
| Validate the method by | SENSITIVITY ANALYSIS. Re-run the totals under different weightings and report whether the ranking survives. A ranking that flips when a weight moves by a tenth is a property of the weights, not of the thing being scored. |
| Cost measured in | reviewer-minutes — scored per criterion, so cost multiplies by the number of criteria |
| It cannot produce | a defensible ranking if the weights were chosen after seeing the results. |
| Use it when | several things matter at once and they are not equally important |
| Do not use it | when one criterion actually decides everything — say so and grade on that. |
TEMPLATE — we did not run this. Every figure below is illustrative. Our own testing used an LLM as judge; this panel exists so you can see what a human error analysis dashboard would need to contain before you build one.
Why it failed, counted by cause
| Cause | Count | Starts at which stage | Zero means |
|---|---|---|---|
| Retrieval missed the doc | 14 | retrieval | — |
| Answered from prior knowledge | 9 | generation | — |
| Dropped in chunking | 8 | chunking | — |
| Never extracted | 0 | extraction | CANNOT FIRE by construction |
How to read it. A cause counted zero is either genuinely absent or unable to fire, and those look identical on a chart. Every zero must say which. Each row needs a real example under it, not just a number.
| Input it needs |
|
| Output it gives |
|
| Validate the method by | THE ZERO ROW IS THE TELL. A cause with a count of zero is either genuinely absent or CANNOT FIRE by construction, and those look identical on a chart. Every zero must say which it is, or the taxonomy is decoration. |
| Cost measured in | reviewer-minutes — the most expensive method per row and usually the highest-value |
| It cannot produce | a pass rate. It explains failures; it does not count successes. |
| Use it when | always, at least once. It is what turns a score into a fix. |
| Do not use it | as the thing you report as a number. It is a diagnosis, not a metric. |
TEMPLATE — we did not run this. Every figure below is illustrative. Our own testing used an LLM as judge; this panel exists so you can see what a hybrid, with escalation dashboard would need to contain before you build one.
Which leg decided, and how often it escalated
| Leg | Rows decided | Its own validated accuracy | Cost unit |
|---|---|---|---|
| Code grader | 880 | 85% (threshold sweep) | free |
| LLM judge | 120 | adjudicated, see above | tokens |
How to read it. The routing rule silently decides most of the volume, so it needs its own audit: sample what the cheap leg settled and check the expensive one would have agreed. Each leg keeps its own accuracy; they do not average.
| Input it needs |
|
| Output it gives |
|
| Validate the method by | VALIDATE EACH LEG SEPARATELY, then check the ROUTING: sample the rows the cheap leg settled and confirm the expensive one would have agreed. An unchecked routing rule silently decides most of your volume. |
| Cost measured in | tokens — the point of the design — most rows never reach the expensive leg |
| It cannot produce | a single accuracy figure. Each leg has its own, and they do not average. |
| Use it when | volume is high and the expensive method is only needed on the hard rows |
| Do not use it | before the individual legs are validated. Routing between two unknowns is one unknown. |
GlossaryKey terms
CheckCheck your understanding
How do you evaluate a GenAI feature?
Build a representative test set, define metrics, automate scoring (metrics + LLM-judge, human-checked), gate releases in CI, and A/B test live.
What's special about evaluating LLMs?
Outputs are open-ended and non-deterministic, so you need rubric-based judging, distributions over many cases, and human calibration.
How do evals make model swaps feasible?
A regression eval set tells you instantly whether a new model or prompt kept quality - so change is safe, not scary.
What changedWhat changed here
As of 2026-09-25 — evaluation and leaderboard coverage in the daily brief: 2 items in the last 7 days
Updated this page A controlled clinical benchmark checks whether model confidence tracks evidence quality and uncertainty, providing a way to test calibration before deployment.
Three kinds of claim, strongest first. Signal runs every morning.