Group-relative advantages (GRPO)
A group-relative advantage compares each rollout's reward with the rewards of other rollouts sampled for the same task, so the group itself serves as the baseline. For agent RL it trades the learned value model for extra generation per task, and generation is the part of the system that already scales out. GRPO (Group Relative Policy Optimization), introduced in DeepSeekMath, is the best-known method built on this baseline. The name also covers a bundle of loss choices that later methods keep or replace independently.
Comparing a hard repository fix with an easy arithmetic question mixes task difficulty into the signal. Comparing several attempts at the same repository fix measures which sampled behavior did better from the same starting conditions.
Computing group-relative advantages
For a task , sample a group of rollouts from the current policy and score them, giving rewards . The group mean is the baseline:
The original GRPO also divides by the group's standard deviation:
with a small constant added to the denominator so that uniform groups do not divide by zero.
With an outcome reward, is copied to every sampled token of rollout and fed to the policy-gradient loss. In the rewards group worked through in advantages and baselines, mean-centering gives , and the population standard deviation is , so std division gives .
Worked example
A harder task where only one of four rollouts passes the tests has rewards . The mean is and the standard deviation is . Three normalizations give:
| Rollout | (MaxRL) | |||
|---|---|---|---|---|
| 1 | 1 | |||
| 2–4 | 0 |
All three put the lone success above the baseline and give it a larger advantage than either success in the group, since a rare success is more informative. They disagree on how much larger, and that choice sets how much weight each task receives as a function of its pass rate.
Normalizing by the standard deviation
For binary rewards with group pass rate , the std-normalized advantage of a success is
At it is . At it is , and at each failure gets . Compared with plain mean-centering, the division multiplies a group's advantages by , which upweights groups whose rewards are nearly uniform: tasks that are almost always solved or almost always failed.
Liu et al. call this a question-level difficulty bias and propose Dr. GRPO, which drops the std division (and GRPO's per-response length normalization, covered under length bias). Without either normalization the update is the plain policy gradient of expected reward, under which a binary group's total advantage mass is proportional to and peaks at (advantage mass).
The case for keeping std normalization is scale invariance. It makes advantages comparable across tasks whose rewards live on different scales, such as a 0–1 pass rate in one environment and a 0–10 rubric score in another. The near-uniform case needs a guard. With only the tiny constant in the denominator, a group like produces advantages of order one ( and ) from a difference that may be judge noise.
Normalizing by the mean (MaxRL)
MaxRL (Tajwar et al.) starts from the observation that a binary verifier defines a likelihood: the policy's probability of solving task . Expected-reward RL follows , while maximum likelihood follows , which weights hard tasks by . The paper shows that standard RL is a first-order approximation of the likelihood objective and that dividing the centered reward by the group mean makes the gradient unbiased for a truncation of the likelihood whose order grows with . The first-order truncation is expected reward, so more rollouts per task move the objective toward maximum likelihood. Groups with no success get zero advantage. The authors report that MaxRL Pareto-dominates the methods they compared on every model and task tested, with up to 20× better test-time scaling efficiency than GRPO.
The three normalizations weight a binary group with pass rate as follows:
| Normalization | Advantage of one success | Total mass |
|---|---|---|
| None (mean-centering) | ||
| (GRPO) | ||
| (MaxRL) |
At the total mass is , and respectively. None of the choices changes which rollouts sit above the baseline, only how the update divides its weight among easy and hard tasks.
Uniform groups
When every rollout earns the same reward, all advantages are zero under every normalization above. An all-success group contains only successful attempts, and an all-failure group contains no observed success, though that does not prove the policy cannot solve the task. For a binary task with success probability , a group of is uniform with probability
| 0.5 | 0.13 | 0.008 | ≈0 |
| 0.1 | 0.66 | 0.43 | 0.19 |
| 0.02 | 0.92 | 0.85 | 0.72 |
Larger groups help most near the middle and slowly at the extremes. At , even sixteen rollouts leave 72% of groups uniformly failed, a task distribution problem that difficulty filtering addresses.
Choosing the group size
Larger groups estimate each task's reward distribution more reliably and make mixed outcomes more likely, at the cost of spending more rollouts on one task instead of covering more tasks. DeepSeekMath sampled 64 outputs per question, and many later recipes use 8 or 16. The right size depends on reward sparsity, rollout cost, and how varied the policy's strategies are.
The comparison is only meaningful when the rollouts share their start conditions: repository state, tool budget, hidden tests and sampling settings. A stochastic judge adds within-group variance of its own, so some of the advantage then comes from the scorer rather than the agent.
What GRPO bundles
GRPO as published combines the group baseline with PPO-style clipping, a KL penalty to a reference model, and per-response length normalization. The group baseline is the part most descendants keep. DAPO, Dr. GRPO, GSPO and MaxRL change different pieces, and RL algorithms for LLMs decomposes them. How the scalar advantage is spread over turns and tokens is a separate question, covered in credit assignment.