Asynchronous RL
Asynchronous RL runs rollout generation and gradient updates at the same time: inference keeps sampling with the latest weights it has while the trainer computes the next update from finished rollouts. For agent RL, where rollouts take minutes and vary widely in length, it removes most of the idle time of a synchronous loop at the price of policy lag: each batch was generated by a policy a few versions older than the one being updated.
That makes asynchronous training a little bit off-policy. The behavior policy is a recent checkpoint of the same model, the lag is bounded by a fixed number of versions, and the gap is corrected token by token with the importance ratio against probabilities recorded at sampling time. This is a small, measured departure from on-policy data, not offline RL, where the behavior policy may be distant or unknown.
The synchronous baseline
A synchronous loop alternates two phases. Inference generates a full batch with policy and waits, then the trainer computes from that batch and waits while inference loads the new weights. Every rollout is on-policy for the update that consumes it.
The cost is idle hardware. If generation takes 6 minutes and the update 2, the trainer idles 75% of the time and inference idles 25%. Agent rollouts make this worse, because each batch waits for its slowest rollout and rollout durations have a heavy tail: one coding task that takes 20 minutes holds up hundreds that took 2.
Overlapping generation and training
In an asynchronous loop, while the trainer computes step , inference is already generating the rollouts step will use, with the weights it already has. Completed rollouts stream into a buffer, and the trainer takes a batch when enough have arrived. Flow control usually lets inference run at most a fixed number of steps ahead of the trainer. When the two phases take similar time, a one-step overlap hides most of the idle time.
The idea goes back to distributed actor-learner systems such as IMPALA, which introduced V-trace to correct for the lag. For LLMs, AReaL and PipelineRL are fully asynchronous designs built around long generations.
Rollouts that outlive a weight update
A long agent rollout can span several optimizer steps, so new weights arrive while it is still generating. There are two ways to handle this.
Pin the rollout to its starting version. Keep the old weights available until every rollout that began on them finishes. Each rollout then has a single behavior policy, at the price of serving two versions at once or draining in-flight work before switching.
Update in flight. Pause generation, swap the weights, and resume the same requests on the new weights (weight broadcast). Each rollout's tokens then come from a sequence of versions, and nothing is served twice. PipelineRL calls these in-flight weight updates. The INTELLECT-3 report (Prime Intellect, 2025) measured about 1,500 seconds per step at 65,536-token sequence length with in-flight updates, and more than twice that without them.
Rollouts cross a version boundary
A rollout that spans versions has no single behavior checkpoint, but each of its tokens still has one. The sampler records the probability it used for every token under whatever weights were live at that moment, so a token-level correction against those probabilities stays well defined, while a correction that assumed one checkpoint per rollout would not. Tokens sampled after the swap also attend to keys and values computed under the old weights, a hybrid covered in KV cache and prefix reuse, and the recorded probabilities already include its effect. A rollout's age becomes a span of versions, and the conservative age is measured from the oldest one.
Worked example
A staleness bound allows a rollout whose oldest version is to train in a batch that updates version only if . With , two rollouts start while inference serves .
- A math rollout finishes in 40 seconds under and trains in the batch that updates . Its lag is .
- A coding rollout runs for 25 minutes. Its first 3,000 tokens are sampled under , the next 4,500 under , and the rest under , so it records the span . It finishes while the trainer is collecting the batch that updates , a lag of .
The coding rollout is dropped, even though its last tokens are one version old. With it would train, each token reweighted against the probability recorded when it was sampled. Time spent waiting in the buffer counts toward lag in the same way as time spent generating.
Bounding lag
Rollouts that would exceed the bound are cancelled while still generating, to save compute, or dropped from the buffer before a batch ships. Long, difficult tasks are the most likely to exceed it, as the coding rollout above does, so a tight bound removes the hardest successful rollouts from training. Batch assembly has a similar bias: taking whichever rollouts finish first fills batches with short, easy tasks. MiniMax's Forge uses a windowed FIFO, in which only rollouts within a window of the generation queue (for example 30% of a batch) may be taken out of order, so slow tasks inside the window hold the batch.
Integer version age is only a proxy for how off-policy the data is. After a large update two adjacent versions can differ a lot, while several small updates may barely move token probabilities. Importance-ratio statistics logged alongside age, and drop rates by environment and rollout length, show whether lag is costing signal. The same ratio also absorbs trainer–inference mismatch, which exists even at zero lag.
How much staleness is tolerable
Published runs tolerate a handful of versions with standard clipped token losses, and far more with losses built for stale data.
- INTELLECT-2 (Prime Intellect, 2025) ablated asynchrony on a 1.5B model: with rollouts up to four steps old, the reward curve matched the synchronous baseline. The 32B run itself used two-step asynchrony.
- ScaleRL (Khatri et al., 2025), a study of RL recipes that used more than 400,000 GPU-hours, found that PipelineRL reached about the same asymptotic reward as conventional off-policy PPO with better compute efficiency. Of the bounds it compared, 8 steps did better than 4, and the final recipe uses 8.
- INTELLECT-3, a 106B-parameter MoE trained on 512 H200s, discarded rollouts generated more than 8 policy versions ago.
- M2PO (Zheng et al., 2025) constrains the second moment of the importance weights instead of clipping each ratio, and trained stably on data at least 256 updates stale while matching on-policy accuracy.
The first three support a bound of about 4 to 8 versions as a default. M2PO locates the failure in the variance of the importance weights, which grows with staleness, so ratio statistics are the signal to watch when raising the bound.