Era 6 · Generative models and systems · 2023

51 DPO

Direct Preference Optimization: Your Language Model is Secretly a Reward Model · Rafailov, Sharma, Mitchell, Manning, Ermon & Finn · Stanford · NeurIPS
🟥 read in full~2 horiginal ↗
The gist in 20 seconds. RLHF without the RL: the preference objective is solved directly as a plain classification loss on preference pairs — no separate reward model, no PPO loop. The model is its own reward model. Alignment got simpler and more stable → the default for open models.

Context

RLHF (#44) is powerful but a COMPLICATED pipeline: a separate reward model plus unstable RL (PPO) with a pile of hyperparameters. Rafailov et al. ask: can you get the result of RLHF WITHOUT the RL?

The idea and the mechanism

The mathematical insight: the optimal RLHF policy has a closed form in terms of the reward, which means the reward also has a closed form in terms of the policy. Substitute that into the preference model and the RLHF objective turns into a plain classification loss directly on (preferred, rejected) pairs. No separate reward model and no PPO loop needed.

optimization · probability The derivation: how an RL problem collapses into classification

Step 1. The RLHF objective maxπ E[r] − β KL(π ‖ πref) has a well-known closed-form solution:

π*(y|x) = 1Z(x) πref(y|x) · exp(r(x,y)β)

Step 2. Express the reward in terms of the policy (just rearranging):

r(x,y) = β log π*(y|x)πref(y|x) + β log Z(x)

Step 3. Substitute into the Bradley-Terry preference model P(yw≻yl) = σ(rw−rl) — the normalizer Z(x) cancels (same x!), and what is left is a loss directly on the policy:

L = −E log σ(β log πθ(yw)πref(yw) − β log πθ(yl)πref(yl))

No reward model, no PPO — ordinary classification on pairs. The language model itself implicitly is the reward model (hence the paper's subtitle). This three-line derivation is the real value of the work.

PyTorch The DPO loss
import torch.nn.functional as F
# pi_*, ref_* — log-probabilities of the chosen/rejected answer (policy / reference)
def dpo_loss(pi_w, pi_l, ref_w, ref_l, beta=0.1):
    logits = beta * ((pi_w - ref_w) - (pi_l - ref_l))
    return -F.logsigmoid(logits).mean()     # just classification on pairs
RLHF (#44): reward PPO (RL) DPO (#51): one classification loss on pairs simpler, more stable,no separate model, no RL
RLHF: a reward model plus unstable RL. DPO folds the same thing into a single direct loss on preference pairs.
Analogy. To teach someone taste you could hire a separate taster (the reward model) and tune the cook through them by trial and error (RL) — slow and shaky. Or you could tell the cook directly: "this dish is better than that one" — and let them nudge their own recipes in the right direction. DPO is "teach on pairs, directly", with no taster in the middle.

Why it matters

It simplified alignment sharply; it became the default recipe for open models and the basis for a whole family of variants (IPO, KTO, ORPO). Understanding its derivation means understanding how modern alignment works without RL.

Connections

The same result (alignment to preferences), but without a separate reward model and PPO. DPO took exactly the same RLHF setup and folded it algebraically into a single step — it beat RLHF on simplicity, not on the goal.

↔ supersedes33. PPO

PPO was the RL engine of alignment; DPO shows that for preference learning you do not need RL at all. Where RLHF spins an unstable RL loop, DPO does an ordinary supervised step — more stable and cheaper.

→ complements45. Constitutional AI

Two independent simplifications of RLHF: CAI takes the human out of the labeling (AI feedback), DPO takes the RL out of the training. They combine: collect the preferences with AI feedback (CAI) and train on them with the DPO loss.

Questions worth asking

If DPO is simpler and more stable, why would anyone still use PPO/RLHF?

RL keeps some advantages: it works with online data (generating and scoring on the fly), whereas DPO learns from a fixed set of pairs and can end up memorizing their distribution. For harder signals (multi-step rewards, verifiable tasks) RL is more flexible. In practice many frontier labs combine the two: DPO as a cheap base, RL fine-tuning where it genuinely earns its place.

"The model is its own reward model" — is that a metaphor or literally true?

Literally true, in the precise sense of the derivation: the quantity β log(πθ/πref) is the implicit reward corresponding to that policy. No separate scoring network is needed — the reward is baked into the ratio of the policy's probabilities to the reference's. This is not an analogy but a consequence of the closed form of the optimal RLHF policy.

Where does DPO break down in practice?

The known problems: sensitivity to the quality and coverage of the pair set, a tendency to push down the probability of both answers (the good one along with the bad) when badly tuned, and a dependence on a good reference. Hence the stream of variants (IPO fixes preference overfitting, KTO works without pairs, ORPO drops the reference). DPO is a strong base, but not a silver bullet, and tuning it is an art of its own.

What to read in the original

Read it in full — the real value is the DERIVATION (a few lines that turn an RL problem into classification); understanding it means understanding modern alignment without RL.