Home › Fine-Tuning & Alignment › Key term › RLHF
Key term · Models

RLHF

Using human preferences to reward better model behavior.

In one line

RLHF trains a model to prefer responses that humans rated as better.

DefinitionWhat it means

Reinforcement Learning from Human Feedback, RLHF, first trains a reward model on human comparisons between candidate responses, then uses reinforcement learning, typically PPO, to update the language model so it generates responses the reward model scores highly. It is applied after SFT to further align outputs with nuanced human preferences like helpfulness, tone, and harmlessness.

Why it mattersWhy you should care

RLHF is largely responsible for turning capable but blunt base models into the helpful, calibrated assistants shipped by major labs; it is how preferences that are hard to write as explicit rules, such as being appropriately cautious without being unhelpfully evasive, get baked into model behavior. It is also expensive and complex to run well, which is a big part of why simpler alternatives like DPO gained traction.

At a glanceSee it

RLHF diagram
RLHF diagram 1

The PPO inner loop that a single Policy Update box hides — generate, score, then subtract a KL penalty against the frozen SFT reference so the policy cannot drift into reward-hacked gibberish.

RLHF diagram 2

The two real design forks behind one RLHF arrow — whether humans or another model label the pairs, and whether you keep an explicit reward model plus PPO or collapse it into DPO.

Where you see itIn the wild

  • Alignment pipelines behind major consumer chat assistants.
  • Human preference labeling platforms collecting response comparisons.
  • Research papers reporting reward model accuracy and PPO training curves.
A living map of modern AI — kept current every morning