Compaction, subagents, and recursive agents

Compaction replaces part of an agent's context with a shorter representation, usually a model-written summary, so a long task continues past the context window. A subagent is a separate model interaction that a parent agent starts for a bounded job, and whose result returns to the parent as an observation. A recursive agent applies subagents repeatedly, with the model choosing how to split a problem into nested calls. All three turn a rollout from one growing token sequence into a tree of model calls, and each choice they involve (when to summarize, what to keep, what to delegate) is something RL can train.

The common case is still linear: observation, turn, observation, turn. Branching pays off when a task needs more working context than one sequence holds, or when parts of it proceed independently.

Branches become training samples

The trace preserves shared history once. Training reads the relevant root-to-leaf paths without pretending that every event occurred in one linear conversation.

What compaction changes

Long agent tasks accumulate command output, file contents, plans, and repeated instructions. Eventually the full history is too expensive or too long to send again. The harness then asks the model for a summary, or writes one, and later requests start from that summary instead of the full history.

The earlier events still happened. The trace keeps them for analysis and training even though the next request contains only the compacted version, so the trace records both the interaction history and the context each model call saw. Because the compacted prompt no longer extends the earlier tokens, training starts a new sample at each compaction (token-level rendering).

Compaction determines which information survives, so it acts like an agent action even when a fixed rule triggers it. A summary that drops a failed test or an unresolved constraint makes the continuation look sensible while breaking the task. When the model writes the summary itself, the summary tokens were sampled from the policy and share in the rollout's reward, so RL improves them like any other action.

Learning when to compact

A fixed trigger, such as "summarize when 16k tokens of context remain", leaves only the summary's content to the policy. To learn the timing as well, the decision has to be a sampled action, typically a tool call, and the reward has to reflect what a long context costs.

Context-Folding gives the agent two tools: branch starts a subtask with its own working context, and return folds it away, leaving only a summary message in the main context. Its RL method, FoldGRPO, adds process penalties for tokens left unfolded once the main context passes half its limit, and for work done outside a branch's stated scope. With a 32K active context and at most 10 branches, the agent reaches 62.0% on BrowseComp-Plus and 58.0% on SWE-bench Verified, ahead of a baseline agent of the same model that keeps its full history in a 327K context.

MEM1 compacts every turn. The model rewrites an internal-state block, and the previous turn's block is pruned from the context, so memory stays constant however long the task runs. On a 16-objective multi-hop QA task, MEM1-7B performs 3.5× better than Qwen2.5-14B-Instruct with 3.7× less memory. CompactionRL keeps a fixed trigger (compaction fires when less than 10,240 tokens of context remain) and trains the summaries jointly with task execution, all under the rollout's reward. GLM-4.5-Air gains 7.0 points on SWE-bench Verified, to 66.8%.

Subagents and branches

A parent agent delegates a repository search, a literature review, or a test diagnosis. The subagent receives a context chosen by the parent or the harness, works through its own turns, and returns a result that becomes an observation for the parent.

task → parent turn → tool result → parent turn → ... → answer
                   └→ subagent request → subagent turns → finding ┘

The parent and subagent use the same model or different ones, share tools or not, and run in one runtime or several. "Subagent" names a relationship in the interaction, independent of how it is served.

What matters for training is provenance. The parent generated the decision to delegate and the instructions it sent. It did not generate the subagent's finding; it observed it. The subagent generated its own actions under its own contexts. Flattening all of these into one conversation loses the action boundaries and trains each policy on text another one wrote. A trace stored as a message graph keeps each branch separate, and each root-to-leaf branch becomes its own training sample.

Credit over branches

A successful answer depends on a subagent's search. The reward belongs to the whole episode, while the trainable tokens live on several branches. How the reward reaches them is a credit assignment choice.

The simplest rule gives every branch the episode's advantage. Suppose a group of four rollouts gets rewards 1,0,0,11, 0, 0, 1, and the first rollout used two subagents. With a group-mean baseline its advantage is +0.5+0.5, and all three of its branches (the parent and both subagents) receive +0.5+0.5 on every sampled token. This reinforces useful delegation. It also rewards a subagent whose finding the parent ignored, because the episode succeeded.

More local rules score a branch by its own output, or estimate whether the parent used it. They improve on the shared advantage only when the extra judgments are reliable (reward models and judges).

Parallel exploration raises a related question. When the harness samples four candidate plans and keeps one, the rejected plans are still policy behavior. They serve as negative examples, or they are excluded when selection was a search procedure outside what is being learned, and the trace needs enough structure to record which.

Learned parallel delegation

Delegation also buys wall-clock time: independent subtasks run concurrently. Training a policy to parallelize runs into a cost-shaped reward problem. Outcome reward alone gives no reason to parallelize, and a bonus for parallelism pays for spawning subagents whether or not the task decomposes.

Kimi K2.5's Agent Swarm, trained with parallel-agent RL (PARL), names both failures (Kimi Team). Serial collapse is the local optimum in which the orchestrator does everything as a single agent, and a reward for instantiating subagents counters it. That reward invites spurious parallelism, many subagents with no meaningful decomposition, which a second reward for the fraction of subagents that finish counters. Both auxiliary terms are annealed to zero during training, leaving the task outcome. Cost is measured in critical steps: per stage, the orchestrator's steps plus those of its longest-running subagent, so ten subagents running in parallel cost the same as the slowest one. The orchestrator is trained while the subagents are frozen checkpoints, which avoids assigning credit across co-trained policies. Agent Swarm cuts latency by up to 4.5× relative to single-agent baselines.

Recursive language models

A recursive language model (RLM) goes further. Instead of reading a huge input in one context, the model treats the input as part of an external environment, such as a variable in a Python REPL, examines it programmatically, and calls itself on pieces of it (Zhang, Kraska, and Khattab). The policy then chooses a computation strategy as well as actions in the task: how to split the input, what to delegate, and how to combine results.

Recursion avoids one enormous prefill, reaches more of a large state, and runs independent work in parallel. It also repeats context across branches, adds coordination overhead, and spreads an early misunderstanding to every child. At serving time, a tree of short calls behaves differently from one long request: branches run concurrently, each needs scheduling and builds its own cache, and a parent waiting on its slowest child inherits long-tail latency. A policy rewarded only for outcomes learns delegation patterns that score well and cost a lot, so recursive environments need a cost term such as critical steps or total tokens.