The ideaEvals in action — an error-analysis walkthrough
This is a live-style eval run for an AI product-description generator (a fashion e-commerce build). A weaker model generates and a stronger model judges — each output scored against five criteria. The dashboard shows why the first prompt failed, and how error analysis — reading the failing logs — turned a two-line prompt change into a large quality jump.
ResultsPass rate by criterion
v1 hallucinates finishes and textures the metadata never mentioned. Toggle to v2 to see the grounded prompt's effect.
Error analysisRead the failing cases
The single most important step in applied AI: don't just see the red — open the logs. Click a case to see the output, the metadata it should have stuck to, the judge's reasoning, and the fix.
DiagnosisFrom symptom to fix
The loopHow error analysis works
Run the eval, read every failure, diagnose the root cause, patch the prompt (or data), add the failure to the test set, and rerun.
SetupThe LLM-as-judge rig
Why weaker-generates / stronger-judges
- Generation model:a smaller, cheaper, faster model (e.g. GPT-4.1) at low temperature — better for high-volume production.
- Judge model:a stronger model (e.g. GPT-5 mini) scoring each criterion, kept consistent across every prompt variant.
- Graders:pass/fail labellers for factuality, hallucination and completeness; numeric 1-5 scorers (threshold 3) for tone and SEO.
- Discipline:effort goes into the system + user prompt, not into paying for a bigger generation model.
Add every failure to the test set
- When you find a failure, don't just fix the prompt — add the case to your eval set so the fix can't regress.
- Run evals like unit tests for intelligence: in CI, or with a data-science team, but the PM owns the criteria.