RL algorithms for LLMs: PPO, GRPO, and variants

RL algorithms for LLMs, such as PPO, GRPO and their variants, are policy-gradient methods that weight the log-probability gradient of each sampled token by an advantage, and each named algorithm is one setting of a few independent design choices. For an agent run the settings matter more than the name, because each choice governs a specific failure: length bias, groups that carry no signal, or instability once the data is off-policy.

Design axes

Five choices separate the methods in use.

  1. Advantage estimator. How a rollout's reward becomes a per-token weight: a learned value model with GAE, the group mean, a leave-one-out mean of the other rollouts, or MaxRL's mean-normalized group credit. GRPO covers the group computation.
  2. Normalization. Whether advantages are divided by the group's standard deviation, and how per-token losses are combined into one batch loss: per response, per token in the batch, or divided by a constant. The token choice creates or removes length bias, which the policy-gradient loss works through.
  3. Loss and trust region. How far one update may move each token: a symmetric ratio clip, clip-higher, a clip on a sequence-level ratio, a truncated importance weight, or a mask in ratio or probability space. PPO clipping and importance sampling cover the mechanics.
  4. Off-policy handling. How data from an older or different sampler enters the loss: several minibatch steps per batch, asynchronous generation with bounded staleness, a token or sequence importance ratio against the recorded behavior probability, sequence masks, and MoE routing replay.
  5. KL. Whether a penalty to a reference model is used, and whether it sits in the reward or the loss (KL penalties).

Worked example

One group scored under three published recipes shows how the axes combine. Four rollouts of a task earn rewards (1,0,0,0)(1, 0, 0, 0), with advantages computed as on the GRPO page. The success took 200 tokens, each failure took 800, and the generation limit is 1,000 tokens. Each token's weight in the update is its advantage times the recipe's normalizer:

RecipeAdvantage, success / failureNormalizerWeight per token, success / failure
GRPO+1.73+1.73 / −0.58-0.581/(4Li)1/(4L_i)+2.2×10−3+2.2\times10^{-3} / −1.8×10−4-1.8\times10^{-4}
DAPO+1.73+1.73 / −0.58-0.581/2,6001/2{,}600+6.7×10−4+6.7\times10^{-4} / −2.2×10−4-2.2\times10^{-4}
Dr. GRPO+0.75+0.75 / −0.25-0.251/(4⋅1,000)1/(4\cdot 1{,}000)+1.9×10−4+1.9\times10^{-4} / −6.3×10−5-6.3\times10^{-5}

GRPO pushes each success token 12 times as hard as each failure token, a factor of 3 from the advantages and 4 from dividing each rollout by its own length. DAPO and Dr. GRPO give every token in the batch the same normalizer, so their ratio is the advantage ratio of 3. Dr. GRPO's total update for this group is also about a fifth of GRPO's, which acts like a lower learning rate. MaxRL would change only the advantage column, to +3+3 / −1-1.

Comparison table

AlgorithmAdvantageNormalizationRatioTrust regionReference KL
PPO (RLHF form)Value model + GAEMean over tokens (varies)TokenSymmetric clip, ϵ=0.2\epsilon = 0.2In the reward
RLOOLeave-one-out meanOne term per responseNoneNoneIn the reward
GRPO(r−mean)/std(r - \text{mean})/\text{std}1/L1/L per response, then group meanTokenSymmetric clipIn the loss, k3k_3, β=0.04\beta = 0.04
Dr. GRPOr−meanr - \text{mean}Token sum over a constantTokenSymmetric clipNone
DAPO(r−mean)/std(r - \text{mean})/\text{std}All tokens in the batchTokenClip-higher, [0.8,1.28][0.8, 1.28]None
GSPO(r−mean)/std(r - \text{mean})/\text{std}One term per responseSequence, geometric meanSequence clip, 3×10−43\times10^{-4} / 4×10−44\times10^{-4}Not stated
CISPOGroup-relativeAll tokens in the groupTokenTruncated, stop-gradient weightNone
GMPO(r−mean)/std(r - \text{mean})/\text{std}Geometric mean over a response's tokensTokenClip (e−0.4,e0.4)(e^{-0.4}, e^{0.4}) in log spaceNone
SAPO(r−mean)/std(r - \text{mean})/\text{std}1/L1/L per responseTokenSigmoid soft gateNone
MaxRL(r−mean)/mean(r - \text{mean})/\text{mean}From the base recipeFrom the base recipeFrom the base recipeFrom the base recipe

What each row changes from its neighbors:

  • PPO trains a value model, often as large as the policy, and reuses each batch for several epochs, which is what makes the clip necessary. The value model is the main cost and is hard to fit when the only reward arrives at the end of a long rollout.
  • RLOO takes one on-policy step per batch, so it needs no ratio and no clip.
  • GRPO replaces the value model with group statistics and keeps PPO's token clip.
  • Dr. GRPO removes the std division and the per-response 1/L1/L, and drops the KL. The GRPO page covers the std debate.
  • DAPO adds clip-higher, dynamic sampling of groups with mixed rewards, a token-level loss and a graded penalty for overlong responses.
  • GSPO moves ratio and clip to the sequence level, which is why its clip range is so narrow.
  • CISPO clips the importance weight instead of the objective, so every token keeps a gradient.
  • GMPO replaces the arithmetic mean over a response's token terms with a geometric mean, which damps outlier ratios and permits a wider clip.
  • SAPO replaces the hard clip with a smooth sigmoid gate. Each token's gradient weight peaks at ρt=1\rho_t = 1 and decays as the ratio moves away, faster for negative advantages.
  • MaxRL changes only the advantage. Its objective moves from expected reward toward maximum likelihood of success as the group grows.

Choosing among them

In ScaleRL's ablations, removing any one of these choices from a tuned recipe mostly changed how fast the run improved, while the loss type also changed where it ended up. For agents with verifiable rewards, a common combination is a group baseline with no value model, no reference KL, token-level normalization over the batch, and a trust region that does not silence rare tokens. The GSPO authors report that sequence-level ratios train MoE models stably without routing replay. Zheng et al., also at Qwen, later trained a 30B MoE stably with token-level importance correction: on its own for on-policy training, and with clipping plus routing replay once each batch took several gradient steps. Sequence ratios are one route to MoE stability, and routing replay is another. PPO with a value model remains the natural fit when rewards arrive mid-rollout and the value function can be learned (credit assignment).