RLHF trains a model to prefer answers that humans actually rate as good.
DefinitionWhat it means
RLHF, Reinforcement Learning from Human Feedback, is a post-training step where human raters compare pairs of model outputs, that preference data trains a reward model, and the base model is then optimized against that reward model using reinforcement learning. The result is a model that is not just capable but aligned with what people find helpful, honest, and safe, which is what separates a raw pre-trained model from a usable assistant like ChatGPT or Claude.
Why it mattersWhy you should care
RLHF is the step that converts a next-token predictor into a product people can trust with real tasks, and it is why the same base model can be tuned toward different personalities or safety postures for different customers. For AI product teams, understanding RLHF explains both the strengths of instruction-tuned models and their failure modes, such as reward hacking or overly cautious refusals.
At a glanceSee it
The single RL box unpacked — the policy chases reward while a KL leash to the frozen reference keeps it from drifting into gibberish that games the score.
RLHF sits in a family of methods — the two forks are who labels the pairs and whether you train an explicit reward model or skip it with DPO.
Where you see itIn the wild
- Model cards describing an alignment or safety fine-tuning stage
- Discussions of reward hacking or sycophancy in assistant behavior
- Comparisons between a base model and its chat-tuned counterpart