KV cache and prefix reuse

The KV cache is the per-layer store of attention keys and values that a transformer inference engine keeps for the tokens it has already processed, so that generating the next token does not recompute attention inputs for the whole context. Prefix caching extends this across requests: when a new request begins with a token prefix the engine has already processed, it reuses the stored keys and values and only prefills the new suffix. In RL for agents, where many requests share long prefixes, prefix reuse is one of the largest levers on generation cost.

Where RL rollouts share prefixes

Two structures in RL training repeat the same context many times.

Groups. The rollouts in a group share one prompt. With 16 rollouts of a 4,000-token task prompt, prefilling each separately processes 64,000 prompt tokens. With prefix caching the prompt is prefilled once and the other 15 rollouts start decoding immediately, so 4,000 tokens are processed.

Turns. Each turn's prompt is the previous turn's prompt plus the model's reply and the observation that followed it. Suppose a rollout has ten turns and each adds 2,000 tokens of context. Without reuse, turn kk prefills all 2000k2000k tokens, for a total of 2000×(1+2+⋯+10)=110,0002000 \times (1 + 2 + \dots + 10) = 110{,}000 tokens. With reuse, each turn prefills only its new 2,000 tokens, 20,000 in total. The gap grows quadratically with the number of turns.

Serving systems implement this by hashing fixed-size blocks of token prefixes, as vLLM does over its PagedAttention block tables, or with a radix tree over cached prefixes, as in SGLang's RadixAttention.

Exact token prefixes

A cache hit requires the new request's token IDs to match the cached ones token for token. Two prompts that read identically but tokenize differently share nothing after the first differing token. Multi-turn reuse therefore depends on token-level rendering: if the harness re-renders history from parsed messages and one token changes near the start, every block after it misses.

The trainer needs the same invariant. A prefix exact enough for a cache hit is exact enough for the recorded behavior probabilities to describe the context the trainer will evaluate. A prefix break costs both a recomputation at inference and a new training sample.

Replica affinity and cache capacity

A prefix cache lives in one engine's GPU memory, so a rollout's second turn benefits only if it is routed to the replica that served the first. Routers for RL traffic therefore pin a rollout's session to a replica, and may route a group's members together so they share the task prefix. Affinity trades against load balance: pinning everything to warm replicas can leave others idle.

Cache capacity is the other limit. Every in-flight rollout holds its context in the cache, and agent contexts grow as they run. When the pool of live rollouts outgrows the cache, the engine evicts prefixes it will soon need again and prefills them a second time. Admission control that watches cache usage (rollout generation) and offloading cold blocks to CPU memory or disk keep reuse from collapsing under load.

The cache after a weight update

Cached keys and values were computed by specific weights. When the trainer publishes a new policy while rollouts are in flight, the system has three coherent options.

  1. Flush and recompute. Clear the cache when weights change. Every in-flight rollout re-prefills its context under the new weights, so its history is consistent with the new policy, at the cost of recomputing everything in flight.
  2. Keep the cache and continue. Leave the cache in place and let in-flight rollouts keep decoding under the new weights. Tokens generated after the update come from new weights attending to keys and values produced by old weights, a hybrid that no single checkpoint reproduces. No compute is wasted.
  3. Namespace the cache by version. Add a salt to the cache key, so requests tagged with one version only hit blocks computed for that version. New rollouts then start from fresh prefills and never inherit stale state.

Option 2 keeps the logged data correct. The sampler still records the probability it used for every token, so the behavior probabilities μ\mu in the trace are exact, and the trainer's importance ratio compares its own probability against them. The hybrid shows up as extra trainer–inference mismatch on tokens sampled just after an update, the per-token discrepancy that token-level masking or ratio truncation is built to absorb. What the system gives up is a single policy that can be said to have generated the rollout. Options 2 and 3 combine: in-flight work continues on its existing cache, and new work is salted by version. Prime Intellect's RL at 1T scale report describes running this combination, with rollouts that mix tokens and cache entries from several policy versions.