Data, Pre-Training & Post-Training · 7 min

Aligning to Preferences: RLHF and the Simpler DPO

How preference optimization turns human A/B judgments into model behavior: the reward-model-plus-PPO pipeline, and DPO's collapse of it into one classification loss.

Supervised fine-tuning teaches a model to imitate good answers. What it cannot teach is that answer A is better than answer B when both look plausible. A demonstration only ever shows one target, and the cross-entropy loss treats every other string as equally wrong. Preference optimization fills that gap. You collect judgments of the form "given this prompt, output A beats output B," and you push the model's probability mass toward the winners. That is the whole game. The two dominant methods, classic RLHF and DPO, differ only in how they turn those pairwise judgments into a gradient.

The mental model: a preference is a coin flip

The mathematical spine of both methods is the Bradley-Terry model, a 1952 framework for paired comparisons (it was built for things like taste tests, not, as often assumed, for ranking chess players; that lineage runs through Elo instead). Bradley-Terry says each response y to a prompt x carries a latent scalar score r(x, y), and the probability a human prefers y_w (winner) over y_l (loser) is the logistic of the score gap:

P(y_w > y_l | x) = sigmoid( r(x, y_w) - r(x, y_l) )

If the winner scores two nats higher, the human "should" prefer it about 88% of the time. This is a model of noisy raters. It never claims the winner gets chosen 100% of the time, which matches reality: on InstructGPT's data, labelers agreed with each other only about 73% of the time, and researcher-labeler agreement sat near 77%. Any pipeline that assumes clean labels is lying to itself.

RLHF, the three-stage pipeline

Ouyang et al. (2022), the InstructGPT paper, is the canonical recipe. Three stages, run in order:

  [1] SFT              [2] Reward Model         [3] PPO
  demos -> pi_SFT      A/B prefs -> r_phi       maximize r_phi
  (imitation)          (Bradley-Terry fit)      minus KL to pi_SFT

Stage 1, SFT. Fine-tune on human-written demonstrations (about 13k prompts for InstructGPT). This gives you pi_SFT, a reference policy that already speaks the right dialect.

Stage 2, the reward model. Take a prompt, sample K completions (InstructGPT used K = 4 to 9), and have a human rank them. A ranking of K items yields C(K,2) pairwise comparisons. Rank four items and you get six pairs out of one annotation session, which is why ranking beats isolated A/B votes. You then train a separate network r_phi by minimizing the Bradley-Terry negative log-likelihood over all pairs from a prompt in one batch:

L(phi) = -E[ log sigmoid( r_phi(x, y_w) - r_phi(x, y_l) ) ]

The reward model is just the base transformer with its token head swapped for a single scalar head reading off the last position. It outputs "how good," not "what word comes next." InstructGPT used a 6B reward model against a 175B policy, kept deliberately smaller for training stability.

Stage 3, PPO. Now optimize the policy pi_theta to produce high-reward completions using Proximal Policy Optimization, a reinforcement-learning algorithm. Raw reward maximization is a trap. The policy will hunt down adversarial gibberish that r_phi happens to score highly (reward hacking), because r_phi is only accurate near the distribution it was trained on. The fix is a per-token KL penalty pulling pi_theta back toward pi_SFT:

objective = E[ r_phi(x, y) ] - beta * KL( pi_theta || pi_SFT )

InstructGPT added a third term (PPO-ptx) that mixes in the pretraining gradient to stop the model forgetting general skills. The headline result still lands hard: a 1.3B InstructGPT was preferred by humans over the 175B GPT-3. Alignment bought more than a 100x parameter increase did.

Builder note. Stage 3 is where teams bleed time. PPO here means holding four models in memory at once (policy, reference, reward, and a value/critic head) plus generating fresh samples every step. It is memory-hungry, sensitive to the KL coefficient and learning rate, and prone to a silent collapse where reward climbs while output quality tanks. Budget for it.

DPO: skip the reward model entirely

Rafailov et al. (2023) asked a sharp question. If the reward is only a means to an end, do we need to build it at all? Their paper, Your Language Model Is Secretly a Reward Model, shows you don't.

The move is algebra, not a new algorithm. The KL-constrained objective from Stage 3 has a known closed-form optimum:

pi*(y|x)  proportional to  pi_ref(y|x) * exp( r(x, y) / beta )

Invert it to solve for the reward that a given policy is optimal for:

r(x, y) = beta * log( pi_theta(y|x) / pi_ref(y|x) )  +  beta * log Z(x)

Z(x) is a nasty partition function summing over every possible output, normally intractable. But drop this reward expression into the Bradley-Terry loss and the reward shows up only as a difference, r(x,y_w) - r(x,y_l). Since Z(x) depends on the prompt alone and not the response, it cancels. What survives is a loss you compute directly from the policy's own log-probabilities:

L_DPO = -log sigmoid( beta * [ log pi_theta(y_w|x)/pi_ref(y_w|x)
                              - log pi_theta(y_l|x)/pi_ref(y_l|x) ] )

Read it plainly. Raise the policy's log-prob on the winner relative to the frozen reference, lower it on the loser, with beta setting how hard you are allowed to pull away from that reference. The reward model didn't vanish; it went implicit, living inside the log-ratio. No separate reward network, no sampling during training, no RL loop. Just a forward pass over your preference pairs and a logistic loss. It is supervised learning wearing an RL trench coat.

A worked micro-step. Say beta = 0.1, and for one pair the policy and reference agree exactly, so both log-ratios are 0. The argument to sigmoid is 0, sigmoid returns 0.5, and the loss is -log 0.5 ≈ 0.69: maximal uncertainty, strong gradient. Training nudges the winner's log-ratio up. Once the winner sits five nats above the loser (0.1 * 5 = 0.5 inside the sigmoid, giving 0.62), the loss softens and the gradient fades. Here beta behaves like a temperature. Too low and the model wanders arbitrarily far from pi_ref and starts producing degenerate text; too high and it barely moves.

What actually goes wrong

Reward hacking (RLHF). The classic failure: the policy games r_phi. Length bias is the everyday version, since reward models often confuse "longer" with "better," so PPO learns to pad.

Distribution shift (DPO). DPO's subtle weakness is that it only ever sees a fixed, offline set of pairs. PPO explores by sampling new completions and scoring them live; DPO cannot judge anything outside its dataset. If your preference pairs came from a different model than the one you are tuning, you are optimizing on off-policy data and results degrade. Practitioners often generate the pairs from the SFT model itself, or run DPO in iterative rounds, to keep the data on-distribution.

Both share one root risk: the labels. Bradley-Terry assumes preferences are transitive and consistent. Humans are neither. Biased or careless comparisons produce a confidently wrong model.

Defender note. Preference data is the attack surface. Whoever writes or filters the A/B pairs writes the model's values. A poisoned subset that consistently prefers some subtle behavior (a backdoor phrase, a slanted refusal policy) gets optimized in as faithfully as any legitimate preference. Under DPO there is no reward model to audit separately; the intent bakes straight into the weights. Treat annotation provenance as a supply-chain control.

Researcher note. DPO's simplicity spawned a family. IPO adds a regularizer against the overfitting DPO shows when a pair is labeled deterministically; KTO drops pairs altogether, learning from single thumbs-up/down signals through a prospect-theory value function. The real design axis is what supervision signal you can actually collect: pairs, rankings, or lone binary labels.

The practical bottom line: DPO delivers most of RLHF's benefit at a fraction of the engineering cost, which is why it dominates open-weight fine-tuning. PPO still earns its keep when you need online exploration or a reusable reward model, for instance to score outputs at inference time, which leads straight into the next module's treatment of reward models as standalone safety classifiers.

Sources

  • Ouyang et al., Training Language Models to Follow Instructions with Human Feedback (InstructGPT), 2022 — arXiv:2203.02155
  • Rafailov et al., Direct Preference Optimization: Your Language Model Is Secretly a Reward Model, 2023 — arXiv:2305.18290
  • Christiano et al., Deep Reinforcement Learning from Human Preferences, 2017 — arXiv:1706.03741 (origin of learning a reward model from comparisons)
  • Schulman et al., Proximal Policy Optimization Algorithms, 2017 — arXiv:1707.06347
  • Bradley & Terry, Rank Analysis of Incomplete Block Designs, Biometrika, 1952 (the paired-comparison model) — https://academic.oup.com/biomet/article-abstract/39/3-4/324/326091
  • Azar et al., A General Theoretical Paradigm to Understand Learning from Human Preferences (IPO), 2023 — arXiv:2310.12036
  • Ethayarajh et al., KTO: Model Alignment as Prospect Theoretic Optimization, 2024 — arXiv:2402.01306