Home › Evals & Testing › Key term › LLM-as-judge
Key term · Operate

LLM-as-judge

Using one model to score another model's answers at scale.

In one line

LLM-as-judge replaces slow human grading with a calibrated model grader.

DefinitionWhat it means

LLM-as-judge is the practice of prompting a capable model with a rubric to score or compare another model's outputs, standing in for expensive human review so that thousands of responses can be graded quickly and consistently, provided the judge is calibrated against human ratings.

Why it mattersWhy you should care

LLM-as-judge is what makes continuous evaluation affordable at production scale, but a miscalibrated or biased judge silently masks real quality regressions, so teams shipping AI products routinely audit judge scores against a smaller human-labeled sample to confirm the judge still agrees with people.

At a glanceSee it

LLM-as-judge diagram
LLM-as-judge diagram 1

A judge earns trust through a calibration loop—scored against a human gold set and tuned until agreement holds, then rechecked as the model drifts.

LLM-as-judge diagram 2

The judge's systematic biases—position, verbosity, and self-preference—each map to a standard mitigation.

Where you see itIn the wild

  • Automated eval pipelines scoring chatbot or RAG answers before release.
  • Vendor benchmark leaderboards that disclose their judge model and rubric.
  • Review discussions on how to detect and correct judge bias.
A living map of modern AI — kept current every morning