KL penalties and reference models
A KL penalty in LLM reinforcement learning is a term that discourages the policy from drifting away from a fixed reference model , usually the checkpoint the run started from. It adds to the loss, or subtracts the equivalent from the reward. The coefficient sets how much reward the policy gives up to stay close, and so how far an agent can move from its starting behavior to learn a new skill.
Two different anchors
Two mechanisms in policy optimization both measure how far a policy has moved, and they are often confused.
| Reference-model KL penalty | Behavior-policy trust region | |
|---|---|---|
| Compared against | A fixed model, often the initial checkpoint | The policy that sampled this batch |
| Moves during the run | No | Yes, every step |
| Question it answers | How far from the starting model is the policy allowed to go overall? | How far can one update trust this batch? |
| Typical form | PPO clipping, a ratio or probability mask, or a KL to |
The trust region is a statistical device for off-policy data: the closer the batch is to the current policy, the further one update can safely move. The reference KL is a budget on total distribution shift, and it holds even when every rollout is perfectly on-policy. A penalty on toward the behavior policy is still a trust region, even though it looks like a KL term.
Why a reference penalty exists
The reference penalty came from RLHF against learned reward models. A reward model is only accurate near the distribution it was trained on. Far from it, the policy can find outputs the reward model scores highly but humans would not, which is reward hacking. Keeping the policy near keeps it where the reward is trustworthy, and it helps preserve fluency and general capabilities that the reward does not measure. Ziegler et al. and InstructGPT both train against a reward with subtracted.
With verifiable rewards, the argument weakens. A unit test does not become less accurate as the policy moves, and the goal is often for the policy to move a long way. Dr. GRPO sets on the grounds that rule-based verifiers remove the distribution-shift concern, and DAPO and CISPO also drop the term. Dropping it also removes the cost of running a reference model forward on every batch. On-policy RL also stays near its starting model in KL even with (forgetting), so the explicit anchor is partly redundant.
Estimating KL from samples: k1, k2, k3
The exact KL sums over the whole vocabulary at every position. Most implementations instead estimate it from the sampled tokens alone, using only two log-probabilities per token. John Schulman's note on approximating KL names three estimators. For with tokens sampled from , let
| Estimator | Formula | Bias | Notes |
|---|---|---|---|
| unbiased | High variance, negative for many samples | ||
| biased, usually slightly | Always nonnegative, low variance | ||
| unbiased | Always nonnegative, low variance |
adds , which has expectation zero under , to . That keeps it unbiased while making every sample nonnegative. The zero expectation needs to put no probability on tokens can never sample; two full softmaxes satisfy this.
A worked example with two sampled tokens:
| Token | ||||||
|---|---|---|---|---|---|---|
| a | 0.5 | 0.2 | 0.4 | 0.916 | 0.420 | 0.316 |
| b | 0.3 | 0.6 | 2.0 | −0.693 | 0.240 | 0.307 |
Token b is one the policy has made less likely than the reference did. scores it as negative divergence, which is correct on average across samples but noisy for any one token. and report both tokens as drift.
In-reward versus in-loss placement
In the reward. Classic RLHF with PPO subtracts , which is , from the per-token reward. The penalty then flows through returns and advantages like any other reward, and the policy-gradient machinery turns it into a correct gradient of the KL-regularized objective. The KL value itself is treated as a constant, never differentiated.
In the loss. GRPO adds per token directly to the loss, with in DeepSeekMath, and backpropagates through it. The paper's reason is to keep the KL out of the advantage calculation.
These are not interchangeable, because the gradient of an unbiased estimate is not necessarily an unbiased estimate of the gradient. For a single sampled token, ignoring later positions:
- The gradient of in the loss has zero expectation, so it does nothing on average.
- The gradient of in the loss is , whose expectation under is the gradient of the forward KL , not the reverse KL the objective states.
- The gradient of in the loss matches the gradient of the reverse KL.
Tang and Munos analyze these pitfalls, including a second one: in-loss penalties at each token ignore how that token changes the contexts of later tokens, which yields only a partial gradient. In-reward placement avoids both, at the cost of mixing the KL into the reward signal.
When to keep it
A reference KL is worth its extra forward pass when the reward is a learned model or a judge that can be exploited, when retaining broad behavior is an explicit goal, or when a run is unstable in ways that point to drift. It is usually unnecessary when rewards are verifiable and the task is the only objective.
Long runs that keep the penalty meet a second problem. As the policy improves, the KL term grows until it dominates the loss and updates shrink. ProRL handles this with reference resets: when validation performance stalls or degrades, it hard-resets to a recent snapshot of the policy and reinitializes the optimizer state. The penalty then bounds drift within each training stage instead of across the whole run.