Importance sampling
Importance sampling estimates an expectation under one distribution from samples drawn under another, by weighting each sample with the ratio of its probabilities under the two. In agent RL it lets the trainer compute a policy-gradient update for the current policy from tokens sampled by a behavior policy , whether an older checkpoint, an inference engine with different numerics, or both. It corrects for where the data came from; how far one update may move the policy is a separate trust-region question, handled by PPO clipping and the masks below.
The importance ratio
For a sampled token at context , the per-token importance ratio is
It equals 1 when the two policies agree on that token. Above 1, the current policy likes the token more than the sampler did, so the token is underrepresented in the batch and gets upweighted. Below 1, it gets downweighted. The identity behind it is
which holds whenever puts nonzero probability everywhere does. In a policy-gradient loss, the importance-weighted term for one token is , with gradient .
The denominator must come from the distribution that sampled the token. Reloading an old checkpoint to recompute it mixes policy lag with trainer–inference mismatch, because kernels, precision, temperature and top-p all change the probabilities. The standard practice is to record the inference engine's log-probability for every sampled token and ship it with the rollout.
What a token ratio corrects
A token ratio corrects one decision: how likely is to choose at a context that produced. It does not account for reaching different prefixes on the way to the reward. Token-level losses in the PPO family are therefore surrogate objectives, whose gradient approaches the true policy gradient only as approaches .
Zheng et al. at Qwen make this precise. The token-level objective is a first-order approximation of the sequence-level objective, and it holds only when both the training–inference discrepancy and policy staleness are small. The token ratio is part of that approximation. In their experiments on a 30B MoE model with FP8 inference and BF16 training, removing the correction for training–inference mismatch caused rapid training collapse and a sharp drop in entropy. Clipping, masks and MoE routing replay all work by keeping close enough to for the approximation to hold.
Token-level and sequence-level ratios
The token-level ratio treats each token as its own action. PPO, GRPO, DAPO and CISPO all work this way. The sequence-level ratio treats the whole response of length as one action, and its exact value is the product of the token ratios:
Because reward is usually assigned per response, this product is the exact change of measure for responses, when the environment's replies depend only on the actions taken. Its weakness is compounding: a 1,000-token response with every has .
GSPO replaces the product with its length-normalized geometric mean . Three tokens with ratios 1.1, 0.9 and 1.2 have an exact sequence ratio of 1.188 and a GSPO ratio of , and every token in the response shares that one weight. The length root keeps the weight on one scale for any response length, at the cost of no longer being an exact correction, and one badly mismatched token can no longer be singled out.
Controlling extreme ratios
Exact importance weights are unbiased but can have very high variance. A token that was rare under and common under receives a huge weight and can dominate a batch. A large ratio also signals a coverage problem: if almost never sampled a region that now prefers, the batch contains little evidence about it, and reweighting cannot create that evidence. Practical losses give up some unbiasedness to bound the weights in one of three ways: truncating the weight, masking tokens, or masking whole sequences.
Truncated importance sampling
Truncated importance sampling (TIS) caps the weight from above and treats it as a constant:
Tokens with still contribute, with weight , and tokens with small ratios keep their small weight. Yao et al. applied TIS to the ratio between the training framework's and the rollout engine's probabilities for the same weights. They showed that this mismatch silently turns nominally on-policy RL into off-policy RL, and that the correction stabilizes training, including with quantized rollouts.
CISPO uses the same shape inside a REINFORCE-style objective: a clipped, stop-gradient weight multiplies . The MiniMax-M1 authors tuned only the upper bound. Because the weight is clipped and the update is not, every token keeps a gradient.
Token masks
A token mask drops a token's policy-gradient contribution entirely when its probabilities disagree too much and keeps the plain importance-weighted term for everything else. Tokens outside the trust region contribute nothing, and tokens inside keep their uncapped weight.
A ratio mask keeps a token when lies inside a fixed band . IcePop, from the Ring-1T report, applies this idea to trainer–inference mismatch on a trillion-parameter MoE model. At the checkpoint that sampled the data, it computes a calibration ratio between the training engine's and the inference engine's probability for the token. A token with keeps as a fixed weight on an ordinary PPO clipped surrogate, and any other token gets weight zero.
A probability-space mask compares the probabilities themselves and keeps a token when . DPPO argues that a single token's ratio is a noisy measure of how far the policy moved. Ratio bounds over-penalize rare tokens, where a small shift in probability is a large ratio, and under-constrain common ones, where a large shift is a modest ratio. DPPO treats the sampled token's absolute difference as a binary approximation of the total variation distance.
The two masks can reach opposite verdicts:
| Token | Ratio band | Probability mask, | |||
|---|---|---|---|---|---|
| A | 0.90 | 0.50 | 0.56 | kept | masked () |
| B | 0.01 | 0.20 | 20 | masked | kept () |
| C | 0.40 | 0.45 | 1.13 | kept | kept |
Token A is a confident prediction the policy abandoned, which the probability mask flags and the ratio band tolerates. Token B is a rare token that became plausible. The ratio band drops it, and the probability mask keeps it with weight 20, because a probability mask bounds the ratio only at .
Masks can be symmetric or sign-aware. PPO clipping zeroes a token's gradient only when the ratio has left the interval in the direction its advantage favors, and DPPO's mask follows the same rule. IcePop's calibration mask is symmetric and drops an out-of-band token whatever the sign of its advantage.
Sequence-level masks
A sequence-level mask drops a whole rollout when its aggregate divergence from is too large. The first-order argument above motivates it: a token can have an unremarkable ratio while sitting in a rollout whose prefix would rarely produce, and there the token-level surrogate no longer tracks the sequence-level objective.
DeepSeek-V3.2 calls its version off-policy sequence masking. For each rollout it computes the mean per-token log-ratio
a sample estimate of the per-token KL from to along that rollout, and masks the rollout when the value exceeds a threshold and its advantage is negative. Positive rollouts stay whatever their divergence. The authors' reasoning is that a model learns most from its own mistakes, while negative samples far from the current policy can mislead the update. Because is the inference engine's recorded probability, the one statistic covers both staleness and trainer–inference mismatch, which DeepSeek reports improves stability in runs that were otherwise unstable.