Home › Evals & Testing › Pairwise / Elo
✅ · Operate

Pairwise / Elo

Rank models and prompts by having judges pick A over B, aggregated into Elo ratings

In one line

Instead of scoring outputs absolutely, you show a judge two answers to the same prompt, record which wins, and convert thousands of these A-versus-B verdicts into a relative Elo leaderboard.

ConceptWhat it is

Pairwise comparison evaluates candidates by asking a judge to choose between two outputs for the same prompt rather than assigning each a number. It exists because absolute scoring is unreliable: people and LLM judges struggle to consistently call an answer a 7 out of 10, but they reliably agree on which of two answers is better. Collecting many such forced choices sidesteps the calibration problem and yields cleaner signal.

The verdicts are aggregated into a rating with a system borrowed from competitive games: Elo, which nudges a score up or down after each match, or a Bradley-Terry model, which fits all match results at once and produces confidence intervals. The output is a relative ranking — a leaderboard, not an absolute quality grade. This is the method behind public model arenas, where crowdsourced blind votes rank frontier models against each other.

How it worksThe mechanics

Pick the candidates to compare and a representative prompt set. For each battle, sample an output from two candidates, present them side by side with the order randomized to blunt position bias, and collect a preference of A, B, or tie from a human or an LLM judge. Log every outcome, then either update Elo scores incrementally after each vote or fit a Bradley-Terry model over the full battle log for a maximum-likelihood ranking. Keep running comparisons until the ratings and their confidence intervals stabilize, then read off the leaderboard.

At a glanceSee it

Pairwise / Elo diagram
Pairwise / Elo diagram 1

Inside the update step — each verdict nudges both ratings by K times the gap between predicted and actual, so upsets move the numbers far more than expected wins.

Pairwise / Elo diagram 2

The judge itself is the error source — three systematic biases distort raw verdicts, and even a debiased protocol can still yield a leaderboard where preferences loop back on themselves.

When to use itWhere it fits

  • You need to rank several models, prompts, or fine-tunes against each other and only relative order matters.
  • Your quality dimension is subjective or holistic (helpfulness, tone, writing quality) where absolute rubrics are hard to pin down.
  • Graders give noisy, poorly calibrated absolute scores but agree well on head-to-head preference.
  • You want a defensible, statistically grounded leaderboard with confidence intervals to justify a launch decision.

When NOT to use itLimits & anti-patterns

  • You need an absolute pass or fail, or a score you can track against a fixed threshold over time.
  • You are evaluating a single system with no comparison baseline to play it against.
  • You have many candidates and a tight budget; all-pairs comparison scales poorly and needs many votes per pair.
  • The task has checkable ground truth (exact match, unit tests, retrieval hit) where cheaper deterministic metrics suffice.

Trade-offsAdvantages & costs

Advantages
  • Forced binary choices are easier and more consistent for judges than absolute scoring, cutting rater noise.
  • Randomized blind pairs reduce anchoring and calibration drift across raters and over time.
  • Bradley-Terry and Elo yield principled rankings with uncertainty estimates, not just bare point scores.
  • Scales to crowdsourced or LLM judges, and new candidates join by playing into the existing pool.
Trade-offs & costs
  • Produces only relative standing; it cannot tell you whether the top-ranked model is actually good enough.
  • Comparison count grows with candidates and with votes needed for significance, so cost climbs fast.
  • Susceptible to position, verbosity, and self-preference bias, especially when an LLM is the judge.
  • Ratings drift as the prompt distribution or the judge changes, so a leaderboard needs periodic refreshing.

ExampleIn the real world

A team building a customer-support assistant must choose between two prompt templates running on two base models. They assemble roughly 300 representative support questions, generate a reply from each variant, and run pairwise battles where support agents pick the better answer with the order randomized. After a few thousand comparisons they fit a Bradley-Terry model, producing an Elo-style leaderboard with confidence intervals. The stronger model paired with the second prompt leads by a margin whose interval clears the runner-up, so they ship that combination. They never learn an absolute quality score, but they have a clear, defensible relative winner.

ToolsHow to implement it

  • LMArena (Chatbot Arena) — the crowdsourced pairwise battle platform that popularized Elo rankings for LLMs.
  • Bradley-Terry fitting via libraries such as choix or statsmodels for maximum-likelihood rankings with confidence intervals.
  • promptfoo and LangSmith — both ship built-in pairwise and model-comparison evaluators for your own bake-offs.
  • Arena-Hard-Auto — an automated pairwise benchmark that uses a strong LLM as the judge.

Cost & effortWhat it takes

Cheap to prototype, expensive to make rigorous. The dominant cost is the number of judgments: tight confidence intervals need many votes per pair, and comparing every pair of candidates scales roughly with the square of the candidate count, so a large bake-off balloons quickly. Human judges give the most trustworthy verdicts but are slow and costly; an LLM judge cuts per-comparison cost to cents while adding bias you must audit. Budget also for upkeep, since leaderboards go stale as models, prompts, or judges change and must be re-run.

A living map of modern AI — kept current every morning