🎯 · Models

RLHF

Train a reward model on human preferences, then optimize the policy against it.

In one line

Humans rank outputs, a reward model learns their taste, and reinforcement learning bakes it in.

ConceptWhat it is

RLHF, reinforcement learning from human feedback, aligns a model to human preferences by first training a separate reward model on human-ranked comparisons of model outputs, then using reinforcement learning, typically PPO, to optimize the original model against that reward signal. It exists because instruction tuning alone teaches format and topic-following but does not directly optimize for the subtler qualities humans actually prefer, like helpfulness, tone, and safety trade-offs.

RLHF was the core technique behind making early ChatGPT-style assistants feel notably more aligned than their raw instruction-tuned predecessors.

How it worksThe mechanics

Human annotators rank multiple model outputs for the same prompt from best to worst; a reward model is trained to predict these preference rankings; the base policy model is then fine-tuned using reinforcement learning, most commonly PPO, to maximize expected reward while a KL penalty keeps it from drifting too far from the original model.

At a glanceSee it

RLHF diagram
RLHF diagram 1

The reward model is trained on comparison pairs — a pairwise loss forces the chosen answer's score above the rejected one, turning subjective human rankings into a single scalar reward.

RLHF diagram 2

Inside the single PPO box is a real loop — the policy samples, is scored while a frozen reference plus KL penalty holds it in place, and is nudged each round until over-optimisation starts gaming the reward.

When to use itWhere it fits

  • Consumer-facing assistants where subtle tone, helpfulness, and safety alignment materially affect user trust.
  • Products with enough scale to justify collecting substantial human preference data.
  • Cases where instruction tuning alone leaves the model technically correct but not well-calibrated in style or judgment.
  • Organizations with the infrastructure to run the full reward model plus RL training pipeline.

When NOT to use itLimits & anti-patterns

  • Small teams without the infrastructure or budget for the reward model plus RL pipeline, where DPO gets similar results more simply.
  • Narrow, well-defined tasks where instruction tuning alone is already sufficient.
  • Situations needing fast iteration, since RLHF training is notoriously unstable and slow to tune.

Trade-offsAdvantages & costs

Advantages
  • Directly optimizes for what humans actually prefer, not just format-following.
  • Historically produced the largest jump in perceived assistant quality after instruction tuning.
  • Reward model can be reused across multiple RL fine-tuning iterations.
  • Flexible framework that can incorporate safety, helpfulness, and other preference dimensions.
Trade-offs & costs
  • Complex, multi-stage pipeline that is notoriously unstable and hard to tune.
  • Requires substantial human preference labeling, which is expensive and slow to collect.
  • Risk of reward hacking, where the policy exploits reward model quirks rather than genuinely improving.
  • Heavier infrastructure and expertise requirement than simpler alignment methods like DPO.

ExampleIn the real world

OpenAI's original InstructGPT and early ChatGPT used RLHF with a trained reward model and PPO to align GPT-3-based models to human preference rankings, well beyond what instruction tuning alone achieved.

ToolsHow to implement it

  • Hugging Face TRLprovides PPOTrainer and reward modeling utilities for RLHF pipelines.
  • OpenAI PPO implementationthe reinforcement learning algorithm most commonly paired with reward models.
  • Anthropic and OpenAI preference datasetsreference examples of human ranking data formats used to train reward models.
  • Weights and Biasestracking reward model accuracy and PPO training stability.

Cost & effortWhat it takes

High cost and engineering effort: requires human preference labeling at scale, a separate reward model, and unstable RL training; latency to a finished aligned model is measured in weeks for most teams.

A living map of modern AI — kept current every morning