59 GRPO
Context
RLHF (#44) runs on PPO (#33), and PPO keeps a separate value/critic network — it estimates the expected reward (the baseline) to reduce the variance of the advantage. At LLM scale the critic is comparable in size to the policy itself: that is ×2 memory and compute for the RL stage. Expensive.
The idea and the mechanism
Drop the critic. For each prompt the old policy samples a group of G answers; each gets a reward (from a reward model, or a verifiable reward — math, code). The baseline is the mean reward over the group; each answer's advantage is how much better than the mean it is (normalized by the spread). After that it is the same clipped surrogate and the same KL to the reference as in PPO, only with a group-relative advantage. No separate value network is needed: comparing several answers to one question is itself the baseline estimate.
reinforcement learning A group advantage instead of a critic
We sample G answers \( o_1,\dots,o_G \) to a prompt, with rewards \( r_1,\dots,r_G \). The advantage is relative to the group:
The objective is the clipped PPO surrogate with this advantage, plus a KL to the reference model (\( \rho_i = \pi_\theta(o_i)/\pi_{\text{old}}(o_i) \)):
Compare with PPO (#33): there the advantage is computed by a separate value network (via GAE). GRPO replaces it with the group mean — answers above the mean get a positive advantage, those below it a negative one. The clip stops big steps from blowing up the policy, the KL keeps it near the reference. The critic is gone, PPO's stabilizers stayed.
Python GRPO: the advantage out of the group's rewards
def grpo_advantages(rewards): # rewards: [G] — rewards of the answers to ONE prompt
mu, sd = rewards.mean(), rewards.std() + 1e-6
return (rewards - mu) / sd # relative to the group mean — no critic
# policy update: the same clip+KL as in PPO, but A is group-relative
# loss = -min(rho*A, clip(rho,1-eps,1+eps)*A) + beta*KL(pi||ref)
Why it matters
GRPO made RL for LLMs cheaper (no critic — half as many "heavy" networks and half the memory) and became the workhorse of reasoning RL: DeepSeekMath → R1 and a wave of open reasoning models. It is an open, reproducible recipe for what produces o1-style reasoning (#58) — taking it from the closed frontier to a generally available tool.
Connections
GRPO is PPO without the value network: the same clipped surrogate and KL, but the advantage comes from the mean over a group of samples rather than from a separate critic. PPO's stabilizers are kept, its most expensive part is removed.
The reasoning o1 is known for is grown by RL — and GRPO became one such engine, out in the open: rewarding a correct answer, it teaches the model long self-checking chains. What o1 demonstrated behind closed doors, GRPO made reproducible.
Both cut down the complexity of RLHF, but differently: DPO removes RL altogether (classification on pairs, offline), GRPO keeps online RL but removes the critic. Hence the division of labour: DPO for cheap preference alignment, GRPO for reasoning with verifiable rewards, where online sampling matters.
Questions worth asking
Why can the critic be thrown out at no cost — wasn't it reducing variance?
The critic is there as a baseline subtracted from the reward to reduce the variance of the policy gradient. But if you sample several answers to one prompt, their mean reward is already a good, unbiased baseline for that prompt. The group replaces the trained value network, saving half the memory and compute — and it often works just as well, especially with verifiable rewards.
Then how is GRPO different from REINFORCE with a baseline?
In essence it is policy gradient with a group baseline, but with two details borrowed from PPO: the clip on the probability ratio (so a big step does not blow up the policy) and the KL to the reference (so it does not drift far from a sensible model). Plus normalizing the advantage by the group's spread makes it relative and robust. It is "REINFORCE reinforced with PPO's stabilizers, but without the critic".
Where is GRPO weak?
A group baseline requires sampling several answers per prompt — that costs inference at every RL step. The advantage estimate is cruder than a trained critic's, and it is noisy on small groups. And like any reward-driven RL, GRPO is sensitive to the quality of that reward: on "soft" tasks with no verifiable reward you need a good reward model, with all its attendant risks (reward hacking, sycophancy — see #44).
What to read in the original
Read selectively: from DeepSeekMath (arXiv 2402.03300), the derivation of the GRPO objective and the motivation for removing the critic (that is the main value); from the R1 report, how GRPO is deployed for reasoning on verifiable rewards. Worth reading straight after PPO (#33), so you can see exactly what was removed and what was kept.