Group-relative advantages (GRPO)

A group-relative advantage compares each rollout's reward with the rewards of other rollouts sampled for the same task, so the group itself serves as the baseline. For agent RL it trades the learned value model for extra generation per task, and generation is the part of the system that already scales out. GRPO (Group Relative Policy Optimization), introduced in DeepSeekMath, is the best-known method built on this baseline. The name also covers a bundle of loss choices that later methods keep or replace independently.

Comparing a hard repository fix with an easy arithmetic question mixes task difficulty into the signal. Comparing several attempts at the same repository fix measures which sampled behavior did better from the same starting conditions.

Computing group-relative advantages

For a task xx, sample a group of nn rollouts from the current policy and score them, giving rewards r1,…,rnr_1,\ldots,r_n. The group mean is the baseline:

rˉ=1n∑j=1nrj,Ai=ri−rˉ.\bar r = \frac{1}{n}\sum_{j=1}^{n} r_j, \qquad A_i = r_i - \bar r .

The original GRPO also divides by the group's standard deviation:

Ai=ri−rˉstd⁡(r1,…,rn),A_i = \frac{r_i - \bar r}{\operatorname{std}(r_1,\ldots,r_n)},

with a small constant added to the denominator so that uniform groups do not divide by zero.

With an outcome reward, AiA_i is copied to every sampled token of rollout ii and fed to the policy-gradient loss. In the rewards 1,0,0,11, 0, 0, 1 group worked through in advantages and baselines, mean-centering gives ±0.5\pm 0.5, and the population standard deviation is 0.50.5, so std division gives ±1\pm 1.

Worked example

A harder task where only one of four rollouts passes the tests has rewards 1,0,0,01, 0, 0, 0. The mean is 0.250.25 and the standard deviation is 0.25⋅0.75≈0.433\sqrt{0.25 \cdot 0.75}\approx 0.433. Three normalizations give:

Rolloutrir_iri−rˉr_i - \bar r÷ std\div\ \text{std}÷ rˉ\div\ \bar r (MaxRL)
11+0.75+0.75+1.73+1.73+3.0+3.0
2–40−0.25-0.25−0.58-0.58−1.0-1.0

All three put the lone success above the baseline and give it a larger advantage than either success in the 1,0,0,11,0,0,1 group, since a rare success is more informative. They disagree on how much larger, and that choice sets how much weight each task receives as a function of its pass rate.

Normalizing by the standard deviation

For binary rewards with group pass rate pp, the std-normalized advantage of a success is

1−pp(1−p)=1−pp.\frac{1-p}{\sqrt{p(1-p)}} = \sqrt{\frac{1-p}{p}} .

At p=1/2p = 1/2 it is 11. At p=1/16p = 1/16 it is 15≈3.9\sqrt{15}\approx 3.9, and at p=15/16p = 15/16 each failure gets −3.9-3.9. Compared with plain mean-centering, the division multiplies a group's advantages by 1/p(1−p)1/\sqrt{p(1-p)}, which upweights groups whose rewards are nearly uniform: tasks that are almost always solved or almost always failed.

Liu et al. call this a question-level difficulty bias and propose Dr. GRPO, which drops the std division (and GRPO's per-response length normalization, covered under length bias). Without either normalization the update is the plain policy gradient of expected reward, under which a binary group's total advantage mass is proportional to p^(1−p^)\hat p(1-\hat p) and peaks at p^=1/2\hat p = 1/2 (advantage mass).

The case for keeping std normalization is scale invariance. It makes advantages comparable across tasks whose rewards live on different scales, such as a 0–1 pass rate in one environment and a 0–10 rubric score in another. The near-uniform case needs a guard. With only the tiny constant in the denominator, a group like (1,1,1,0.99)(1, 1, 1, 0.99) produces advantages of order one (+0.58+0.58 and −1.73-1.73) from a difference that may be judge noise.

Normalizing by the mean (MaxRL)

MaxRL (Tajwar et al.) starts from the observation that a binary verifier defines a likelihood: the policy's probability pθ(x)p_\theta(x) of solving task xx. Expected-reward RL follows ∇θpθ(x)\nabla_\theta p_\theta(x), while maximum likelihood follows ∇θlog⁡pθ(x)=∇θpθ(x)/pθ(x)\nabla_\theta \log p_\theta(x) = \nabla_\theta p_\theta(x)/p_\theta(x), which weights hard tasks by 1/p1/p. The paper shows that standard RL is a first-order approximation of the likelihood objective and that dividing the centered reward by the group mean makes the gradient unbiased for a truncation of the likelihood whose order grows with nn. The first-order truncation is expected reward, so more rollouts per task move the objective toward maximum likelihood. Groups with no success get zero advantage. The authors report that MaxRL Pareto-dominates the methods they compared on every model and task tested, with up to 20× better test-time scaling efficiency than GRPO.

The three normalizations weight a binary group with pass rate pp as follows:

NormalizationAdvantage of one successTotal mass ∑i∣Ai∣\sum_i \lvert A_i\rvert
None (mean-centering)1−p1-p2np(1−p)2np(1-p)
÷ std\div\ \text{std} (GRPO)(1−p)/p\sqrt{(1-p)/p}2np(1−p)2n\sqrt{p(1-p)}
÷ rˉ\div\ \bar r (MaxRL)(1−p)/p(1-p)/p2n(1−p)2n(1-p)

At p=1/16p = 1/16 the total mass is 0.12n0.12n, 0.48n0.48n and 1.88n1.88n respectively. None of the choices changes which rollouts sit above the baseline, only how the update divides its weight among easy and hard tasks.

Uniform groups

When every rollout earns the same reward, all advantages are zero under every normalization above. An all-success group contains only successful attempts, and an all-failure group contains no observed success, though that does not prove the policy cannot solve the task. For a binary task with success probability pp, a group of nn is uniform with probability

P(uniform)=pn+(1−p)n.P(\text{uniform}) = p^n + (1-p)^n .
ppn=4n=4n=8n=8n=16n=16
0.50.130.008≈0
0.10.660.430.19
0.020.920.850.72

Larger groups help most near the middle and slowly at the extremes. At p=0.02p = 0.02, even sixteen rollouts leave 72% of groups uniformly failed, a task distribution problem that difficulty filtering addresses.

Choosing the group size

Larger groups estimate each task's reward distribution more reliably and make mixed outcomes more likely, at the cost of spending more rollouts on one task instead of covering more tasks. DeepSeekMath sampled 64 outputs per question, and many later recipes use 8 or 16. The right size depends on reward sparsity, rollout cost, and how varied the policy's strategies are.

The comparison is only meaningful when the rollouts share their start conditions: repository state, tool budget, hidden tests and sampling settings. A stochastic judge adds within-group variance of its own, so some of the advantage then comes from the scorer rather than the agent.

What GRPO bundles

GRPO as published combines the group baseline with PPO-style clipping, a KL penalty to a reference model, and per-response length normalization. The group baseline is the part most descendants keep. DAPO, Dr. GRPO, GSPO and MaxRL change different pieces, and RL algorithms for LLMs decomposes them. How the scalar advantage is spread over turns and tokens is a separate question, covered in credit assignment.