LLM-as-judge uses a smarter model as a scalable stand-in for human graders.
ConceptWhat it is
LLM-as-judge uses a separate, typically more capable, language model to score or compare outputs against a rubric, in place of a human grader, for tasks where correctness is subjective or open-ended.
It exists because human evaluation does not scale to thousands of outputs, while simple string-matching metrics cannot capture qualities like helpfulness, tone, or reasoning quality that matter in real applications.
How it worksThe mechanics
The judge model is given the original prompt, the candidate output, and a scoring rubric or reference answer, then asked to assign a score or pick the better of two candidates; results are aggregated across many examples to produce a quality metric that tracks changes across model or prompt versions.
At a glanceSee it
The reliability layer: a judge carries known biases — position, verbosity, and self-preference — and each has a fix. Calibrating against human labels is what earns the metric trust.
Two judging modes: absolute scoring against a rubric, or pairwise A-versus-B aggregated into an Elo. Pairwise stays steadier when quality is hard to pin to a single number.
When to use itWhere it fits
- Scoring open-ended tasks like summarization or chat quality at scale.
- Running pairwise comparisons between two model or prompt versions.
- Continuously monitoring output quality where labeled ground truth does not exist.
- Supplementing human review to cover far more examples for the same budget.
When NOT to use itLimits & anti-patterns
- High-stakes decisions needing certified human judgment, such as medical or legal review.
- Tasks with an objective, checkable answer where a simple offline metric is cheaper and more reliable.
- Situations where judge-model bias or blind spots could systematically favor certain answer styles.
Trade-offsAdvantages & costs
Advantages
- Scales to thousands of outputs far cheaper than human review.
- Captures nuanced quality dimensions simple metrics miss.
- Enables fast iteration on prompts and model choices.
- Can be customized with rubrics for domain-specific criteria.
Trade-offs & costs
- Judge models carry their own biases and blind spots.
- Judge scores can be gamed by outputs that look good superficially.
- Adds cost and latency from an extra model call per evaluation.
- Needs periodic validation against human judgment to stay trustworthy.
ExampleIn the real world
An AI writing assistant team uses GPT-4o as a judge to score thousands of generated marketing blurbs for tone and persuasiveness daily, reserving human review for a small audited sample.
ToolsHow to implement it
- Ragasincludes LLM-judge metrics for RAG-specific quality dimensions.
- LangSmithsupports configuring LLM-as-judge evaluators on traced runs.
- OpenAI Evalsframework for defining model-graded evaluation tasks.
- Claude or GPT-4ocommonly used as the judge model for its reasoning strength.
Cost & effortWhat it takes
Costs an additional model call per evaluated example, typically using a stronger and pricier model than the one being tested. Moderate engineering effort to design and validate rubrics.
What changedWhat changed here
Updated this page Research shows that carefully induced judging rubrics reduce over-crediting when an LLM judges agent performance.
Three kinds of claim, strongest first. Signal runs every morning.