RLHF trains a model to prefer responses that humans rated as better.
DefinitionWhat it means
Reinforcement Learning from Human Feedback, RLHF, first trains a reward model on human comparisons between candidate responses, then uses reinforcement learning, typically PPO, to update the language model so it generates responses the reward model scores highly. It is applied after SFT to further align outputs with nuanced human preferences like helpfulness, tone, and harmlessness.
Why it mattersWhy you should care
RLHF is largely responsible for turning capable but blunt base models into the helpful, calibrated assistants shipped by major labs; it is how preferences that are hard to write as explicit rules, such as being appropriately cautious without being unhelpfully evasive, get baked into model behavior. It is also expensive and complex to run well, which is a big part of why simpler alternatives like DPO gained traction.
At a glanceSee it
The PPO inner loop that a single Policy Update box hides — generate, score, then subtract a KL penalty against the frozen SFT reference so the policy cannot drift into reward-hacked gibberish.
The two real design forks behind one RLHF arrow — whether humans or another model label the pairs, and whether you keep an explicit reward model plus PPO or collapse it into DPO.
Where you see itIn the wild
- Alignment pipelines behind major consumer chat assistants.
- Human preference labeling platforms collecting response comparisons.
- Research papers reporting reward model accuracy and PPO training curves.