Era 6 · Generative models and systems · 2024

59 GRPO

DeepSeekMath: Pushing the Limits of Mathematical Reasoning… · Shao et al. · DeepSeek-AI · 2024 (the GRPO algorithm; the RL engine inside DeepSeek-R1)
🟧 read selectively~1.5 horiginal ↗
The gist in 20 seconds. PPO without a critic. For a prompt we sample a GROUP of G answers and reward each one; the advantage is computed relative to the group mean (and normalized). A separate value network — half the cost of PPO — is not needed: the group is its own baseline. That halves the number of "heavy" networks in RL for LLMs; it is the core of DeepSeek's reasoning training (#53, R1) and the open recipe for what sits behind o1-style reasoning (#58).

Context

RLHF (#44) runs on PPO (#33), and PPO keeps a separate value/critic network — it estimates the expected reward (the baseline) to reduce the variance of the advantage. At LLM scale the critic is comparable in size to the policy itself: that is ×2 memory and compute for the RL stage. Expensive.

The idea and the mechanism

Drop the critic. For each prompt the old policy samples a group of G answers; each gets a reward (from a reward model, or a verifiable reward — math, code). The baseline is the mean reward over the group; each answer's advantage is how much better than the mean it is (normalized by the spread). After that it is the same clipped surrogate and the same KL to the reference as in PPO, only with a group-relative advantage. No separate value network is needed: comparing several answers to one question is itself the baseline estimate.

reinforcement learning A group advantage instead of a critic

We sample G answers \( o_1,\dots,o_G \) to a prompt, with rewards \( r_1,\dots,r_G \). The advantage is relative to the group:

\[ A_i = \frac{r_i - \mathrm{mean}(r_1,\dots,r_G)}{\mathrm{std}(r_1,\dots,r_G)} \]

The objective is the clipped PPO surrogate with this advantage, plus a KL to the reference model (\( \rho_i = \pi_\theta(o_i)/\pi_{\text{old}}(o_i) \)):

\[ \mathcal{J}(\theta) = \mathbb{E}\Big[\min\big(\rho_i A_i,\ \mathrm{clip}(\rho_i,\,1-\epsilon,\,1+\epsilon)\,A_i\big)\Big] - \beta\,\mathrm{KL}\!\left(\pi_\theta \,\Vert\, \pi_{\text{ref}}\right) \]

Compare with PPO (#33): there the advantage is computed by a separate value network (via GAE). GRPO replaces it with the group mean — answers above the mean get a positive advantage, those below it a negative one. The clip stops big steps from blowing up the policy, the KL keeps it near the reference. The critic is gone, PPO's stabilizers stayed.

Python GRPO: the advantage out of the group's rewards
def grpo_advantages(rewards):                  # rewards: [G] — rewards of the answers to ONE prompt
    mu, sd = rewards.mean(), rewards.std() + 1e-6
    return (rewards - mu) / sd                  # relative to the group mean — no critic

# policy update: the same clip+KL as in PPO, but A is group-relative
# loss = -min(rho*A, clip(rho,1-eps,1+eps)*A) + beta*KL(pi||ref)
prompt G answers r=0.9r=0.2r=0.7r=0.4 baseline =group mean A = r − baselineupdate (no critic) above the mean → reinforce · below → weaken
Several answers to one prompt are rewarded; the group mean serves as the baseline; each answer's advantage is its deviation from that mean. PPO's separate value network is no longer needed.
Analogy. Instead of keeping a separate judge (the critic) who tells you in advance what an answer is worth, just ask the student to solve the problem several ways and compare those against each other: the ones that came out better than the average of their own batch get reinforced, the ones that came out worse get weakened. A group of answers to one question is its own point of reference — no judge required.

Why it matters

GRPO made RL for LLMs cheaper (no critic — half as many "heavy" networks and half the memory) and became the workhorse of reasoning RL: DeepSeekMath → R1 and a wave of open reasoning models. It is an open, reproducible recipe for what produces o1-style reasoning (#58) — taking it from the closed frontier to a generally available tool.

Connections

← simplifies33. PPO

GRPO is PPO without the value network: the same clipped surrogate and KL, but the advantage comes from the mean over a group of samples rather than from a separate critic. PPO's stabilizers are kept, its most expensive part is removed.

← the engine for58. o1 / test-time compute

The reasoning o1 is known for is grown by RL — and GRPO became one such engine, out in the open: rewarding a correct answer, it teaches the model long self-checking chains. What o1 demonstrated behind closed doors, GRPO made reproducible.

↔ simplifies RLHF differently51. DPO

Both cut down the complexity of RLHF, but differently: DPO removes RL altogether (classification on pairs, offline), GRPO keeps online RL but removes the critic. Hence the division of labour: DPO for cheap preference alignment, GRPO for reasoning with verifiable rewards, where online sampling matters.

Questions worth asking

Why can the critic be thrown out at no cost — wasn't it reducing variance?

The critic is there as a baseline subtracted from the reward to reduce the variance of the policy gradient. But if you sample several answers to one prompt, their mean reward is already a good, unbiased baseline for that prompt. The group replaces the trained value network, saving half the memory and compute — and it often works just as well, especially with verifiable rewards.

Then how is GRPO different from REINFORCE with a baseline?

In essence it is policy gradient with a group baseline, but with two details borrowed from PPO: the clip on the probability ratio (so a big step does not blow up the policy) and the KL to the reference (so it does not drift far from a sensible model). Plus normalizing the advantage by the group's spread makes it relative and robust. It is "REINFORCE reinforced with PPO's stabilizers, but without the critic".

Where is GRPO weak?

A group baseline requires sampling several answers per prompt — that costs inference at every RL step. The advantage estimate is cruder than a trained critic's, and it is noisy on small groups. And like any reward-driven RL, GRPO is sensitive to the quality of that reward: on "soft" tasks with no verifiable reward you need a good reward model, with all its attendant risks (reward hacking, sycophancy — see #44).

What to read in the original

Read selectively: from DeepSeekMath (arXiv 2402.03300), the derivation of the GRPO objective and the motivation for removing the critic (that is the main value); from the R1 report, how GRPO is deployed for reasoning on verifiable rewards. Worth reading straight after PPO (#33), so you can see exactly what was removed and what was kept.