Era 5 · The LLM era · 2022

44 InstructGPT (RLHF)

Training language models to follow instructions with human feedback · Ouyang, Wu, Jiang et al. · OpenAI · NeurIPS
🟥 read in full~2 horiginal ↗
The gist in 20 seconds. Turned RLHF into an operational recipe (SFT → reward model → PPO) for aligning GPT-3 with what the user actually wants. People prefer the answers of the 1.3B InstructGPT to those of the 175B GPT-3. The recipe that led straight to ChatGPT.

Context

GPT-3 is powerful but raw — it continues text rather than following instructions, and it can be useless, untruthful or toxic. Ouyang et al. align it with human intent.

The idea and the mechanism

RLHF in three stages. SFT: humans write reference answers, the model is fine-tuned on them. Reward Model: for a prompt, several answers are collected and humans RANK them; a reward model is trained to predict the preference. RL (PPO): the LLM policy is optimized to maximize the RM's reward, with a KL penalty for drifting too far from the SFT model.

reinforcement learning The three stages formally

1. SFT — ordinary fine-tuning on demonstrations: maximize log P(y | x).

2. Reward Model. From the rankings we take pairs "chosen yw ≻ rejected yl" and train r under the Bradley-Terry model:

P(yw ≻ yl) = σ(r(yw) − r(yl)),   L = − log σ(r(yw) − r(yl))

3. RL (PPO). Maximize the reward with a KL penalty towards the SFT model:

maxθ Ey∼πθ[r(y)] − β · KL(πθ ‖ πSFT) + γ · Ex∼Dpre[log πθ(x)]

The KL penalty is critical: without it the policy "hacks" the reward model (reward hacking) and degenerates into unreadable text with a high reward. The penalty keeps it near a sane SFT distribution.

PPO-ptx — the third term and InstructGPT's signature detail: it mixes in the gradient of the original pre-training (log-likelihood on the pre-training data Dpre). Without it, alignment charges a tax (the alignment tax) — a regression on standard NLP benchmarks; the ptx term damps that down and keeps the model's general capabilities.

PyTorch The reward-model loss + the RL objective
import torch.nn.functional as F
# reward model: Bradley-Terry on (chosen, rejected) pairs
def rm_loss(r_chosen, r_rejected):
    return -F.logsigmoid(r_chosen - r_rejected).mean()

# PPO RL objective: maximize the reward minus the KL penalty towards SFT
# objective = E[r(y)] − beta·KL(pi_theta||pi_sft) + gamma·E_pre[log pi_theta]  (PPO-ptx)
1. SFT 2. Reward Modelrankings 3. RL (PPO)+ KL penalty demonstrations → preferences → policy optimization
Three stages: fine-tune on examples, learn a reward from human preferences, optimize the model against that reward while tethered to SFT.
Analogy. Training an intern. First you show them samples of good work (SFT). Then, instead of writing out every rule, you simply judge their drafts — "this one is better, that one is worse" — and they build an internal model of your preferences (the reward model). Finally they produce work on their own, guided by that model, but you do not let them wander off into nonsense (the KL penalty towards a sane baseline).

Why it matters

This is the RECIPE ChatGPT was built from (Claude uses a related approach) → the thing that turned LLMs from "autocomplete" into an assistant for hundreds of millions of people. The result: the 1.3B InstructGPT is preferred to the 175B GPT-3 — for usefulness, alignment matters more than size; truthfulness rose and toxicity fell alongside.

Connections

← the engine33. PPO

The RL stage of alignment is PPO, where the policy is a language model and the reward is a reward model. A modest 2017 RL algorithm turned out to be a load-bearing part of what aligned ChatGPT.

← aligns38. GPT-3

GPT-3 supplied a powerful but unfocused base. RLHF adds no knowledge — it focuses the model on following instructions and being useful. The capabilities come from GPT-3, the obedience from RLHF.

→ simplified in51. DPO

RLHF on PPO is complicated (a separate reward model plus unstable RL). DPO will show that the same result is reachable without RL — with a single classification loss. This work set the bar that DPO cleared on simplicity.

Questions worth asking

What is reward hacking, and why does the KL penalty hold it back?

The reward model is an imperfect proxy for human preferences. A policy optimizing it to the limit finds loopholes: answers with a high reward that are bad for people (verbosity, sycophancy, odd patterns). The KL penalty punishes drift from a sane SFT distribution, limiting how far the policy can go in gaming the reward model. It is a trade-off between "optimize the reward" and "do not break the language".

Where does the sycophancy of post-RLHF models come from?

Human annotators tend to score answers higher when those answers are pleasant and agree with them — and the reward model soaks that up. Optimizing it, the policy learns to nod along even when it is wrong. This is a direct artefact of optimizing human preferences rather than truth — a well-known and hard alignment problem.

Why does a 1.3B aligned model beat a raw 175B — is that not a paradox?

No: size gives capability, alignment gives usefulness. The raw 175B knows more but does not grasp what is being asked of it — it continues the text instead of doing the task. The small aligned model does exactly what was requested. On the metric "how useful is this answer to the user", obedience beats erudition — hence the result.

What to read in the original

Read it in full — this is the canonical alignment pipeline; understand the role of each stage and why the KL penalty is there.