Home › Fine-Tuning & Alignment › Key term › DPO
Key term · Models

DPO

Aligning a model on preferences without a separate reward model.

In one line

DPO directly trains a model on preference pairs, skipping the reward model and RL loop.

DefinitionWhat it means

Direct Preference Optimization, DPO, reframes preference alignment as a single supervised loss computed directly on pairs of preferred and rejected responses, mathematically derived to have the same optimum as RLHF without needing a separate reward model or an online reinforcement learning loop. Training looks much closer to ordinary supervised fine-tuning.

Why it mattersWhy you should care

DPO delivers RLHF-like alignment quality with far less infrastructure and instability, no reward model to train, no PPO hyperparameters to tune, which is why it became the default preference-tuning method for many open and commercial models by the mid-2020s. It lowered the cost of alignment enough that smaller teams could meaningfully customize model behavior on their own preference data.

At a glanceSee it

DPO diagram
DPO diagram 1

Opens the DPO-loss black box — a frozen reference and a trainable policy define an implicit reward whose chosen-minus-rejected margin the logistic loss maximizes each step, updating only the policy.

DPO diagram 2

Because DPO optimizes only the margin, the loss can be satisfied by driving the rejected response down faster than the chosen one — leaking probability into unseen text unless a larger beta keeps the policy anchored.

Where you see itIn the wild

  • Open-weight model release notes citing DPO as the alignment step.
  • Fine-tuning libraries offering a DPO trainer alongside SFT trainers.
  • Preference datasets structured as chosen and rejected response pairs.
A living map of modern AI — kept current every morning