Importance sampling

Importance sampling estimates an expectation under one distribution from samples drawn under another, by weighting each sample with the ratio of its probabilities under the two. In agent RL it lets the trainer compute a policy-gradient update for the current policy πθ\pi_\theta from tokens sampled by a behavior policy μ\mu, whether an older checkpoint, an inference engine with different numerics, or both. It corrects for where the data came from; how far one update may move the policy is a separate trust-region question, handled by PPO clipping and the masks below.

The importance ratio

For a sampled token yty_t at context (x,y<t)(x, y_{<t}), the per-token importance ratio is

ρt=πθ(yt∣x,y<t)μ(yt∣x,y<t).\rho_t = \frac{\pi_\theta(y_t \mid x, y_{<t})}{\mu(y_t \mid x, y_{<t})}.

It equals 1 when the two policies agree on that token. Above 1, the current policy likes the token more than the sampler did, so the token is underrepresented in the batch and gets upweighted. Below 1, it gets downweighted. The identity behind it is

Ey∼πθ[f(y)]=Ey∼μ ⁣[πθ(y)μ(y)f(y)],\mathbb{E}_{y \sim \pi_\theta}[f(y)] = \mathbb{E}_{y \sim \mu}\!\left[\frac{\pi_\theta(y)}{\mu(y)} f(y)\right],

which holds whenever μ\mu puts nonzero probability everywhere πθ\pi_\theta does. In a policy-gradient loss, the importance-weighted term for one token is −ρtAt-\rho_t A_t, with gradient −ρtAt∇θlog⁡πθ(yt∣x,y<t)-\rho_t A_t \nabla_\theta \log \pi_\theta(y_t \mid x, y_{<t}).

The denominator must come from the distribution that sampled the token. Reloading an old checkpoint to recompute it mixes policy lag with trainer–inference mismatch, because kernels, precision, temperature and top-p all change the probabilities. The standard practice is to record the inference engine's log-probability for every sampled token and ship it with the rollout.

What a token ratio corrects

A token ratio corrects one decision: how likely πθ\pi_\theta is to choose yty_t at a context that μ\mu produced. It does not account for πθ\pi_\theta reaching different prefixes on the way to the reward. Token-level losses in the PPO family are therefore surrogate objectives, whose gradient approaches the true policy gradient only as πθ\pi_\theta approaches μ\mu.

Zheng et al. at Qwen make this precise. The token-level objective is a first-order approximation of the sequence-level objective, and it holds only when both the training–inference discrepancy and policy staleness are small. The token ratio is part of that approximation. In their experiments on a 30B MoE model with FP8 inference and BF16 training, removing the correction for training–inference mismatch caused rapid training collapse and a sharp drop in entropy. Clipping, masks and MoE routing replay all work by keeping πθ\pi_\theta close enough to μ\mu for the approximation to hold.

Token-level and sequence-level ratios

The token-level ratio treats each token as its own action. PPO, GRPO, DAPO and CISPO all work this way. The sequence-level ratio treats the whole response yy of length LL as one action, and its exact value is the product of the token ratios:

ρ(y)=∏t=1Lρt.\rho(y) = \prod_{t=1}^{L} \rho_t.

Because reward is usually assigned per response, this product is the exact change of measure for responses, when the environment's replies depend only on the actions taken. Its weakness is compounding: a 1,000-token response with every ρt=1.01\rho_t = 1.01 has ρ(y)=1.011000≈21,000\rho(y) = 1.01^{1000} \approx 21{,}000.

GSPO replaces the product with its length-normalized geometric mean ρ(y)1/L\rho(y)^{1/L}. Three tokens with ratios 1.1, 0.9 and 1.2 have an exact sequence ratio of 1.188 and a GSPO ratio of 1.1881/3≈1.0591.188^{1/3} \approx 1.059, and every token in the response shares that one weight. The length root keeps the weight on one scale for any response length, at the cost of no longer being an exact correction, and one badly mismatched token can no longer be singled out.

Controlling extreme ratios

Exact importance weights are unbiased but can have very high variance. A token that was rare under μ\mu and common under πθ\pi_\theta receives a huge weight and can dominate a batch. A large ratio also signals a coverage problem: if μ\mu almost never sampled a region that πθ\pi_\theta now prefers, the batch contains little evidence about it, and reweighting cannot create that evidence. Practical losses give up some unbiasedness to bound the weights in one of three ways: truncating the weight, masking tokens, or masking whole sequences.

Truncated importance sampling

Truncated importance sampling (TIS) caps the weight from above and treats it as a constant:

∇θJ≈Ey∼μ ⁣[min⁡(ρt,C) At ∇θlog⁡πθ(yt∣x,y<t)].\nabla_\theta J \approx \mathbb{E}_{y \sim \mu}\!\left[\min(\rho_t, C)\, A_t\, \nabla_\theta \log \pi_\theta(y_t \mid x, y_{<t})\right].

Tokens with ρt>C\rho_t > C still contribute, with weight CC, and tokens with small ratios keep their small weight. Yao et al. applied TIS to the ratio between the training framework's and the rollout engine's probabilities for the same weights. They showed that this mismatch silently turns nominally on-policy RL into off-policy RL, and that the correction stabilizes training, including with quantized rollouts.

CISPO uses the same shape inside a REINFORCE-style objective: a clipped, stop-gradient weight sg⁡(clip⁡(ρt,1−ϵlow,1+ϵhigh))\operatorname{sg}(\operatorname{clip}(\rho_t, 1-\epsilon_{\text{low}}, 1+\epsilon_{\text{high}})) multiplies Atlog⁡πθ(yt∣x,y<t)A_t \log \pi_\theta(y_t \mid x, y_{<t}). The MiniMax-M1 authors tuned only the upper bound. Because the weight is clipped and the update is not, every token keeps a gradient.

Token masks

A token mask drops a token's policy-gradient contribution entirely when its probabilities disagree too much and keeps the plain importance-weighted term for everything else. Tokens outside the trust region contribute nothing, and tokens inside keep their uncapped weight.

A ratio mask keeps a token when ρt\rho_t lies inside a fixed band [rlow,rhigh][r_{\text{low}}, r_{\text{high}}]. IcePop, from the Ring-1T report, applies this idea to trainer–inference mismatch on a trillion-parameter MoE model. At the checkpoint that sampled the data, it computes a calibration ratio ktk_t between the training engine's and the inference engine's probability for the token. A token with 0.5≤kt≤50.5 \le k_t \le 5 keeps ktk_t as a fixed weight on an ordinary PPO clipped surrogate, and any other token gets weight zero.

A probability-space mask compares the probabilities themselves and keeps a token when ∣πθ(yt)−μ(yt)∣≤ϵ\lvert\pi_\theta(y_t) - \mu(y_t)\rvert \le \epsilon. DPPO argues that a single token's ratio is a noisy measure of how far the policy moved. Ratio bounds over-penalize rare tokens, where a small shift in probability is a large ratio, and under-constrain common ones, where a large shift is a modest ratio. DPPO treats the sampled token's absolute difference as a binary approximation of the total variation distance.

The two masks can reach opposite verdicts:

Tokenμ(yt)\mu(y_t)πθ(yt)\pi_\theta(y_t)ρt\rho_tRatio band [0.2,5][0.2, 5]Probability mask, ϵ=0.3\epsilon = 0.3
A0.900.500.56keptmasked (∣Δ∣=0.40\lvert\Delta\rvert = 0.40)
B0.010.2020maskedkept (∣Δ∣=0.19\lvert\Delta\rvert = 0.19)
C0.400.451.13keptkept

Token A is a confident prediction the policy abandoned, which the probability mask flags and the ratio band tolerates. Token B is a rare token that became plausible. The ratio band drops it, and the probability mask keeps it with weight 20, because a probability mask bounds the ratio only at (μ+ϵ)/μ(\mu + \epsilon)/\mu.

Masks can be symmetric or sign-aware. PPO clipping zeroes a token's gradient only when the ratio has left the interval in the direction its advantage favors, and DPPO's mask follows the same rule. IcePop's calibration mask is symmetric and drops an out-of-band token whatever the sign of its advantage.

Sequence-level masks

A sequence-level mask drops a whole rollout when its aggregate divergence from μ\mu is too large. The first-order argument above motivates it: a token can have an unremarkable ratio while sitting in a rollout whose prefix πθ\pi_\theta would rarely produce, and there the token-level surrogate no longer tracks the sequence-level objective.

DeepSeek-V3.2 calls its version off-policy sequence masking. For each rollout it computes the mean per-token log-ratio

1L∑t=1Llog⁡μ(yt∣x,y<t)πθ(yt∣x,y<t),\frac{1}{L}\sum_{t=1}^{L} \log \frac{\mu(y_t \mid x, y_{<t})}{\pi_\theta(y_t \mid x, y_{<t})},

a sample estimate of the per-token KL from μ\mu to πθ\pi_\theta along that rollout, and masks the rollout when the value exceeds a threshold δ\delta and its advantage is negative. Positive rollouts stay whatever their divergence. The authors' reasoning is that a model learns most from its own mistakes, while negative samples far from the current policy can mislead the update. Because μ\mu is the inference engine's recorded probability, the one statistic covers both staleness and trainer–inference mismatch, which DeepSeek reports improves stability in runs that were otherwise unstable.