On-policy and off-policy RL
Training data is on-policy for an update when it was sampled from the policy being updated, and off-policy when some other distribution sampled it: an older checkpoint, a different model, or the same weights under different decoding settings. For agent RL, the size of that gap determines how many updates an expensive rollout can feed and how much correction the loss needs for its gradient to improve the target policy.
The label describes a pair: one piece of data and one update, or in the terms of policies, whether the behavior policy that sampled the tokens matches the target policy being updated. A rollout is on-policy for the checkpoint that sampled it and becomes off-policy once the trainer takes a step. A teacher's demonstration is off-policy from the moment it is written, even if it was generated a second ago.
Three regimes
Distance from the target is continuous, but three regimes cover the systems in use:
- On-policy. for every token in the batch. A strictly on-policy loop generates a batch, takes one optimizer step and discards the batch, so rollouts are single-use and generation cannot overlap training.
- A little off-policy. is a recent version of , a bounded number of optimizer steps old, and the rollout recorded for every sampled token. The importance ratio between the two stays near 1, and a token-level importance-sampling correction, bounded by a clip or mask, closes the gap.
- Off-policy. is far from or unknown: another model's outputs, a human-written dataset, or old replay. Ratios are extreme or cannot be computed, so methods drop them. Supervised fine-tuning imitates the data directly, and offline preference methods such as DPO learn from comparisons.
Asynchronous RL belongs to the middle regime. It samples continuously from policy versions a few steps behind the trainer, bounds that staleness, and reweights each token by its recorded behavior probability. It is sometimes filed under offline RL, which learns from a fixed dataset whose behavior policy may be unknown and arbitrarily far away. An async run has the recorded for every token and keeps it within a few updates of .
Worked example
A few tokens show why each regime uses different machinery. In the first two rows an agent's policy samples the command grep with . The third row is a token a teacher wrote.
| Case | |||
|---|---|---|---|
| Trained in the same step | 0.40 | 0.40 | 1.00 |
| Trained two steps later (async) | 0.40 | 0.46 | 1.15 |
Teacher wrote find, which the student rates 0.002 | 0.90 | 0.002 | 0.0022 |
The first row needs no correction. The second is a little off-policy: a weight of 1.15 is well inside any clip or mask, and the update proceeds as a slightly reweighted policy gradient. The third row is a teacher token. Its ratio is tiny, and a 200-token demonstration whose token ratios have a geometric mean of 0.5 has a sequence ratio of . No importance correction survives that, which is why methods that train on teacher tokens drop the ratio and imitate.
Why the distinction matters for the gradient
The policy gradient is an expectation over rollouts drawn from the target policy:
A batch sampled from reflects 's frequencies. Tokens that liked more than are overrepresented, and the naive estimate is biased toward them. Reweighting by removes the bias for each next-token choice at the contexts produced, which is a good approximation while the two policies stay close. The ratio needs from the moment of sampling, so a rollout that did not record it cannot be corrected later.
Online versus on-policy
Online means the system keeps collecting experience while it learns. Offline means it learns from a fixed dataset. The two axes usually travel together, but they separate. An online system can train on data that is a few steps stale by the time it arrives, because the trainer moved while the rollout was generating. A saved batch is on-policy for the checkpoint that produced it, even after the run has stopped sampling.
Sources of mismatch
Production systems relax the strict loop in several ways, each moving the data further from on-policy:
| Source of mismatch | What differs between and |
|---|---|
| Several optimizer steps per batch (PPO epochs) | Weights move between minibatches of the same data |
| Asynchronous RL | Rollouts were sampled by a checkpoint one or more steps old |
| Weight updates during a rollout | Different tokens of one rollout came from different checkpoints |
| Different inference and training kernels | Same weights, slightly different probabilities |
| Replay of old data or data from another model | Behavior policy may be many steps or a whole model away |
Each relaxation buys data efficiency or hardware utilization with estimator bias. Two properties characterize a run: how far may drift from before data is dropped, and whether the per-token behavior probability was recorded so the loss can account for the drift.
Student and teacher tokens
A teacher can supply training signal in two ways, and they fall on opposite sides of the on-policy line. When the teacher writes the rollouts and the student trains on them with cross-entropy, as in supervised fine-tuning or hard distillation, every token is off-policy: the student learns what to do at contexts the teacher reached, which may be contexts the student never produces. At test time the student conditions on its own earlier tokens, and a mistake the teacher never made leads to a context with no training data.
When the student writes the rollouts and the teacher scores each of the student's tokens with its log-probabilities, the data is on-policy. That is on-policy distillation: the student is corrected at the contexts it visits, with a dense per-token signal in place of a single reward.