Rollout generation
Rollout generation is the stage of RL training that runs the current policy on tasks and collects the finished rollouts: an inference engine serves the model while agent workers drive it through environments until each rollout ends. It is usually the most expensive stage of agent RL, because group baselines need several rollouts per task, each rollout makes many model calls, and each call may be followed by a wait on a tool or sandbox.
Rollouts finish at different times
The unit of scheduling
An inference engine sees independent requests. The training system needs whole rollouts, each carried from the first prompt to termination and returned as one record. A worker receives a task, sends the first model request, turns the reply into an action through the harness, passes it to the environment, appends the observation, and repeats until the rollout stops.
The scheduler therefore admits rollouts, and each admitted rollout then issues an unknown number of requests. A math problem may take one request and a second of compute. A coding task may take forty requests over ten minutes, most of it spent waiting on builds and tests.
Groups and group identity
GRPO and other group-relative methods sample several rollouts of the same task and compare their rewards. The members of a group share the task, its initial state, and the sampling settings, which are part of the behavior policy (policies), so that reward differences measure the policy's variation and not the problem's.
Members of a group do not need to run together. They can be dispatched one at a time, land on different workers, and finish in any order. What has to survive is the label saying which rollouts belong together. A group identifier separate from the task identifier lets the same task be sampled again while an earlier group for it is still in flight. A request that fails for infrastructure reasons still counts toward its group's size, so the group closes without waiting for a rollout that will never arrive, and the failure stays out of the reward comparison (rollouts and traces).
Serving features for RL traffic
Three inference-engine features matter more for RL than for chat serving:
- Continuous batching replaces finished sequences immediately, so short rollouts do not wait for the longest one in a batch.
- Prefix caching lets a group's members share the prefill of their common prompt, and lets each turn reuse the previous turn's context when the new prompt extends the exact tokens already cached (KV cache).
- Session affinity routes every turn of one rollout to the same replica, where its cached prefix lives.
Concurrency and admission control
More concurrent rollouts hide more environment latency, because the engine can decode for one rollout while another waits on a tool. The limit is KV-cache memory. Take an 8B model with 36 layers, 8 key-value heads and head dimension 128, served in bf16. Each token of context holds bytes, about 147 KB, of keys and values. An 80 GB GPU running at 90% memory utilization with 16 GB of weights leaves about 56 GB for the cache, or about 380,000 tokens. That is room for about 47 rollouts at 8,000 tokens of context each, or about 12 at 32,000.
Agent contexts grow as they run, so a pool that fits at dispatch can outgrow the cache ten turns later. The engine then preempts sequences or evicts prefixes it will need again, recomputes them, and throughput falls sharply. The right concurrency depends on the model, the hardware and rollout length, and it moves during training as the policy's behavior changes, so a fixed cap is wrong for part of the run. Admission control instead watches the engine's cache usage and queue depth: it admits more rollouts while the cache has headroom and stops admitting, or cancels the youngest rollouts, before it fills.
Heavy-tailed durations
Rollout durations have a heavy tail. Suppose 2% of rollouts run ten times longer than the median. A synchronous batch of 128 rollouts contains at least one of them with probability , so nearly every step waits about ten median durations for its last rollout while most inference capacity sits idle.
Four responses are common.
- Overlap. Generation for the next batch starts while the tail of the current one finishes, at the cost of training on slightly stale data (async RL).
- Over-provisioning. More rollouts are dispatched than the batch needs and the first to finish are used. Unless the late rollouts are kept for a later batch, a batch filled this way over-represents short rollouts, so the update is computed on a distribution skewed toward shorter behavior.
- Partial rollouts. Each step generates at most a fixed token budget per rollout. Kimi k1.5 saves the unfinished part of a long generation in a replay buffer and continues it in the next iteration, so a long rollout spans several policy versions instead of stalling one step.
- Length-aware scheduling. Synchronous systems reshape the rollout phase around the tail. RollPacker moves prompts expected to produce long responses into dedicated rollout rounds and reports 2.03–2.56× shorter end-to-end training time than veRL on up to 128 H800 GPUs. Seer uses the similar output lengths of rollouts that share a prompt to split and schedule groups, and reports up to 2.04× higher rollout throughput and 72–94% lower tail latency than synchronous baselines.
Timeouts and turn limits cut the tail as well, but a truncated rollout is a different outcome from a completed or failed one, and the environment has to say which (rollouts and traces).
Measuring generation
Useful metrics follow the path of experience:
- time in model inference versus time in the environment
- rollout length, truncation rate and error rate by environment
- reward variance within groups
- the fraction of generated tokens that reach the trainer
Raw tokens per second overstates progress when many of those tokens belong to failed, stale or zero-signal rollouts. Counting only tokens that reach an update gives the measure used in inference-heavy workloads.