Credit assignment
Credit assignment is the problem of deciding which decisions in a rollout were responsible for its outcome. In a long agent rollout some actions matter, some are harmless, and some undo earlier progress, and a single final reward does not say which was which.
An advantage says how much better a rollout did than expected. Credit assignment sets how that signal is spread across the rollout's turns and tokens before the policy-gradient loss consumes it.
Rollout-level credit
The simplest method gives every sampled token in the rollout the same advantage. GRPO and RLOO do this by default: one reward, one baseline, one scalar broadcast over the whole trace.
Consider a three-turn coding rollout that ends with passing tests (reward 1) in a group whose mean reward is 0.5:
| Turn | What the agent did | Rollout-level advantage |
|---|---|---|
| 1 | Read the failing test and the relevant file | |
| 2 | Made an edit that broke an unrelated import | |
| 3 | Noticed the error, fixed the import and the bug |
The harmful edit in turn 2 is reinforced along with the good work. Across many rollouts this averages out, because behavior that tends to appear in successful rollouts is reinforced more often than behavior that tends to appear in failed ones. The cost is variance. Each token's advantage carries information about the whole rollout, so the longer the rollout and the more decisions it contains, the more samples are needed to separate the decisive actions from incidental ones.
Rollout-level credit has one strong property: it optimizes the outcome that was measured, with no intermediate proxy.
Turn-level credit with a value estimate
If the system can estimate the probability of eventual success from an intermediate state, it can credit each turn by how much it changed that estimate. With and value estimates after each turn,
where at the terminal state is the actual reward. Suppose at the start, after turn 1, after turn 2, and the rollout ends with reward :
| Turn | before | after | Turn advantage |
|---|---|---|---|
| 1 | 0.5 | 0.4 | |
| 2 | 0.4 | 0.2 | |
| 3 | 0.2 | 1.0 |
The turn advantages sum to , the rollout-level advantage, because the differences telescope to . What changes is the distribution: the fix in turn 3 receives most of the credit and the broken edit is pushed down. This is the end of GAE applied at turn granularity.
The values can come from a learned critic, as in PPO, or from Monte Carlo estimates: sample several continuations from each intermediate state and average their rewards. VinePPO takes the second route. It found that learned value networks on math reasoning barely beat a random baseline at ranking alternative steps, and that Monte Carlo estimates outperformed PPO on MATH and GSM8K despite the extra generation.
Step-level groups from shared states
A group baseline can be applied below the rollout level whenever several actions were taken from the same state. Rollouts in a group start from the same task and environment state, and agents that loop or backtrack revisit the same web page, room or game screen across rollouts. GiGPO (Feng et al.) finds these repeated anchor states after generation, groups the actions taken from each one, and compares each action's return from that step with the mean of its step-level group. A weighted step advantage is added to the ordinary rollout-level advantage, and no extra rollouts are generated. The authors report gains over GRPO of more than 12% on ALFWorld and more than 9% on WebShop at the same rollout budget.
The method depends on states that repeat verbatim and can be compared cheaply, as in text-game and web-shop environments. A shell whose output carries timestamps or process IDs rarely produces two identical states.
Tree rollouts at uncertain steps
The alternative is to create shared states on purpose: run part of a rollout, then sample several continuations from the same point, forming a tree. Continuations that share a prefix form a group for the decision right after the fork point, so their reward differences credit that decision rather than the whole rollout. Forking at every step is expensive, so the question is where to fork. ARPO (Dong et al.) observed that token entropy rises right after a tool result arrives. It samples some full rollouts, monitors entropy after each tool call, and samples extra partial rollouts from steps where entropy has risen past a threshold. Tokens on the shared prefix get a shared advantage and tokens after the fork get distinct ones. The authors report better results than rollout-level RL baselines on reasoning and search benchmarks with half the tool-call budget.
For agents, each fork point needs the environment state (files, databases, processes) restored or copied, which is the same cost that limits Monte Carlo value estimates. Environments that can snapshot a sandbox make this practical.
Process rewards
A process reward scores intermediate steps directly instead of estimating future success. It can be a verifier for a milestone the task already defines (a lemma checks, a failing test starts passing, a subgoal completes), or a learned process reward model (PRM) trained on step-level labels, as in Let's Verify Step by Step.
Process rewards give denser signal, which matters most when outcomes are rare. They also add intermediate proxies that the policy learns to produce whether or not they lead to success. A milestone is worth rewarding when it is trustworthy evidence of progress and the policy cannot satisfy it without making that progress.
Credit across branches and agents
Agent traces are not always one line. A parent agent may spawn subagents, compact its context, or run a verifier agent, producing a branched trace, and the training system has to decide which branches receive the parent's advantage. Broadcasting it to a subagent's work treats the subagent as part of one policy decision. Scoring subagent branches separately requires a reward for the subtask.
In multi-agent settings the same question appears across roles. A proposer and a solver trained together should each be compared with their own alternatives, and a group mean that mixes both roles, or both sides of a zero-sum game, blurs the signal (self-play).
Choosing a method
| Method | Signal source | Resolution | Main risk |
|---|---|---|---|
| Rollout-level | final reward | whole rollout | variance grows with length |
| Turn-level values | critic or Monte Carlo continuations | per turn | critic error or continuation cost |
| Step-level groups | actions from a revisited or forked state | per step | needs repeated or restorable states |
| Process reward | verifier or PRM on steps | per step | proxy exploitation |
| Branch-level | per-branch rewards | per subagent or role | needs subtask rewards |
A good default is the most local signal that can be obtained without changing what counts as success. When no such signal exists, rollout-level training still works, but rollout count, group size and task difficulty have to compensate for the weak credit.