Home › Reference › Evals › Dashboard
✅ Evals dashboard

Evals & Error Analysis

A worked eval run with LLM-as-judge scoring, failing-case logs, and the error-analysis loop that turns red into shipped quality.

The ideaEvals in action — an error-analysis walkthrough

This is a live-style eval run for an AI product-description generator (a fashion e-commerce build). A weaker model generates and a stronger model judges — each output scored against five criteria. The dashboard shows why the first prompt failed, and how error analysis — reading the failing logs — turned a two-line prompt change into a large quality jump.

ResultsPass rate by criterion

v1 hallucinates finishes and textures the metadata never mentioned. Toggle to v2 to see the grounded prompt's effect.

Error analysisRead the failing cases

The single most important step in applied AI: don't just see the red — open the logs. Click a case to see the output, the metadata it should have stuck to, the judge's reasoning, and the fix.

DiagnosisFrom symptom to fix

The loopHow error analysis works

error analysis loop

Run the eval, read every failure, diagnose the root cause, patch the prompt (or data), add the failure to the test set, and rerun.

SetupThe LLM-as-judge rig

Why weaker-generates / stronger-judges

  • Generation model:a smaller, cheaper, faster model (e.g. GPT-4.1) at low temperature — better for high-volume production.
  • Judge model:a stronger model (e.g. GPT-5 mini) scoring each criterion, kept consistent across every prompt variant.
  • Graders:pass/fail labellers for factuality, hallucination and completeness; numeric 1-5 scorers (threshold 3) for tone and SEO.
  • Discipline:effort goes into the system + user prompt, not into paying for a bigger generation model.

Add every failure to the test set

  • When you find a failure, don't just fix the prompt — add the case to your eval set so the fix can't regress.
  • Run evals like unit tests for intelligence: in CI, or with a data-science team, but the PM owns the criteria.
With thanksThe error-analysis walkthrough follows AI Evals For Engineers & PMs on Maven, by Shreya Shankar and Hamel Husain. Thanks & what I learned →
A living map of modern AI — kept current every morning