Home › Evals & Testing › LLM-as-judge
✅ · Operate

LLM-as-judge

Using a strong LLM to score another model's outputs when there is no single correct answer.

In one line

LLM-as-judge uses a smarter model as a scalable stand-in for human graders.

ConceptWhat it is

LLM-as-judge uses a separate, typically more capable, language model to score or compare outputs against a rubric, in place of a human grader, for tasks where correctness is subjective or open-ended.

It exists because human evaluation does not scale to thousands of outputs, while simple string-matching metrics cannot capture qualities like helpfulness, tone, or reasoning quality that matter in real applications.

How it worksThe mechanics

The judge model is given the original prompt, the candidate output, and a scoring rubric or reference answer, then asked to assign a score or pick the better of two candidates; results are aggregated across many examples to produce a quality metric that tracks changes across model or prompt versions.

At a glanceSee it

LLM-as-judge diagram
LLM-as-judge diagram 1

The reliability layer: a judge carries known biases — position, verbosity, and self-preference — and each has a fix. Calibrating against human labels is what earns the metric trust.

LLM-as-judge diagram 2

Two judging modes: absolute scoring against a rubric, or pairwise A-versus-B aggregated into an Elo. Pairwise stays steadier when quality is hard to pin to a single number.

When to use itWhere it fits

  • Scoring open-ended tasks like summarization or chat quality at scale.
  • Running pairwise comparisons between two model or prompt versions.
  • Continuously monitoring output quality where labeled ground truth does not exist.
  • Supplementing human review to cover far more examples for the same budget.

When NOT to use itLimits & anti-patterns

  • High-stakes decisions needing certified human judgment, such as medical or legal review.
  • Tasks with an objective, checkable answer where a simple offline metric is cheaper and more reliable.
  • Situations where judge-model bias or blind spots could systematically favor certain answer styles.

Trade-offsAdvantages & costs

Advantages
  • Scales to thousands of outputs far cheaper than human review.
  • Captures nuanced quality dimensions simple metrics miss.
  • Enables fast iteration on prompts and model choices.
  • Can be customized with rubrics for domain-specific criteria.
Trade-offs & costs
  • Judge models carry their own biases and blind spots.
  • Judge scores can be gamed by outputs that look good superficially.
  • Adds cost and latency from an extra model call per evaluation.
  • Needs periodic validation against human judgment to stay trustworthy.

ExampleIn the real world

An AI writing assistant team uses GPT-4o as a judge to score thousands of generated marketing blurbs for tone and persuasiveness daily, reserving human review for a small audited sample.

ToolsHow to implement it

  • Ragasincludes LLM-judge metrics for RAG-specific quality dimensions.
  • LangSmithsupports configuring LLM-as-judge evaluators on traced runs.
  • OpenAI Evalsframework for defining model-graded evaluation tasks.
  • Claude or GPT-4ocommonly used as the judge model for its reasoning strength.

Cost & effortWhat it takes

Costs an additional model call per evaluated example, typically using a stronger and pricier model than the one being tested. Moderate engineering effort to design and validate rubrics.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Research shows that carefully induced judging rubrics reduce over-crediting when an LLM judges agent performance.

    arXiv cs.AI · 16 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning