Home › The Road to LLMs › RLHF / alignment
🛤️ · Foundations

RLHF / alignment

Training models to prefer helpful, honest responses using human feedback on their outputs.

In one line

RLHF teaches a model which of its own answers humans actually prefer, turning a raw text predictor into a helpful assistant.

ConceptWhat it is

RLHF, reinforcement learning from human feedback, fine-tunes a pre-trained language model using human preference judgments rather than more text prediction. It was the key technique, popularized by OpenAI's InstructGPT and then ChatGPT, that turned raw next-token predictors into helpful, instruction-following assistants.

It exists because a model trained only to predict the next word has no notion of helpfulness, honesty, or safety; RLHF, and newer variants like DPO, explicitly optimizes for what humans actually want, which is called alignment.

How it worksThe mechanics

Human raters rank several model outputs for the same prompt from best to worst, a separate reward model is trained to predict those rankings, and the base language model is then updated with reinforcement learning to produce outputs the reward model scores highly.

At a glanceSee it

RLHF / alignment diagram
RLHF / alignment diagram 1

Inside the RL step, PPO maximizes the reward-model score minus a KL penalty that anchors the policy to a frozen reference model — drop that anchor and the model reward-hacks.

RLHF / alignment diagram 2

RLHF is one point in a design space: feedback can come from humans or from an AI constitution, and preferences can flow through a reward model plus PPO or be optimized directly with DPO.

When to use itWhere it fits

  • Turning a raw pre-trained model into an instruction-following assistant.
  • Reducing harmful, biased, or unhelpful outputs before public release.
  • Aligning a model's tone and behavior to a specific product's needs.
  • Improving a model's ability to follow nuanced, multi-part instructions.

When NOT to use itLimits & anti-patterns

  • Early-stage research models still being evaluated on raw capability, not behavior.
  • Narrow, deterministic tasks where simple supervised fine-tuning is sufficient and cheaper.
  • Situations lacking budget for the human-labeling pipeline RLHF requires.

Trade-offsAdvantages & costs

Advantages
  • Dramatically improves helpfulness and instruction-following.
  • Reduces harmful, toxic, or off-policy outputs.
  • Lets a small amount of human judgment steer a huge model.
  • Newer methods like DPO cut cost versus classic RLHF.
Trade-offs & costs
  • Expensive and slow, requiring large-scale human labeling.
  • Can over-optimize for rater preferences, causing sycophancy.
  • Reward models can be gamed, called reward hacking.
  • Alignment quality is only as good as the raters and their guidelines.

ExampleIn the real world

OpenAI's shift from GPT-3 to InstructGPT and ChatGPT used RLHF to transform a model that just completed text into one that reliably follows instructions and refuses harmful requests.

ToolsHow to implement it

  • TRL by Hugging Faceopen-source library for RLHF and DPO training.
  • InstructGPT recipethe reference methodology most labs adapted for alignment.
  • Anthropic's Constitutional AIalignment approach reducing reliance on raw human labels.
  • Scale AIcommercial human-feedback labeling provider.

Cost & effortWhat it takes

High cost driven by human labeling at scale, often hundreds of thousands of ranked comparisons; moderate additional compute versus pre-training; DPO variants lower cost by skipping the separate reward model.

A living map of modern AI — kept current every morning