Policy gradients

A policy gradient is the gradient of expected reward with respect to the parameters of a stochastic policy. REINFORCE estimates it from sampled rollouts by weighting the gradient of each sampled action's log-probability by the reward that followed, so actions that led to better outcomes become more likely.

PPO and GRPO, the most common RL algorithms for LLMs, build on this estimator. Their changes affect how the weight is computed or how far one update may move the policy; the gradient itself stays recognizable.

The objective

A policy πθ\pi_\theta with parameters θ\theta produces rollouts τ\tau by interacting with an environment. Each rollout receives a reward R(τ)R(\tau). The objective is expected reward:

J(θ)=Eτ∼πθ[R(τ)]=∑τpθ(τ) R(τ).J(\theta) = \mathbb{E}_{\tau\sim\pi_\theta}\left[R(\tau)\right] = \sum_\tau p_\theta(\tau)\,R(\tau).

The difficulty is that θ\theta appears in the sampling distribution, and the reward is typically a black box: a test suite, a verifier, a judge. There is no way to backpropagate through pytest.

The log-derivative trick

The identity ∇θp=p ∇θlog⁡p\nabla_\theta p = p\,\nabla_\theta \log p moves the gradient inside the expectation:

∇θJ=∑τ∇θpθ(τ) R(τ)=∑τpθ(τ) ∇θlog⁡pθ(τ) R(τ)=Eτ∼πθ ⁣[R(τ) ∇θlog⁡pθ(τ)].\nabla_\theta J = \sum_\tau \nabla_\theta p_\theta(\tau)\,R(\tau) = \sum_\tau p_\theta(\tau)\,\nabla_\theta \log p_\theta(\tau)\,R(\tau) = \mathbb{E}_{\tau\sim\pi_\theta}\!\left[R(\tau)\,\nabla_\theta \log p_\theta(\tau)\right].

The right-hand side is an expectation under the current policy, so it can be estimated by sampling. The reward only needs to be evaluated, never differentiated. This is also called the score-function estimator, since ∇θlog⁡pθ\nabla_\theta\log p_\theta is the score.

The environment drops out

For an agent, a rollout interleaves the policy's actions with the environment's observations. Its probability factors as

pθ(τ)=p(x)∏tπθ(at∣ht)  P(ot+1∣ht,at),p_\theta(\tau) = p(x)\prod_{t} \pi_\theta(a_t \mid h_t)\; P(o_{t+1} \mid h_t, a_t),

where xx is the task, ata_t the policy's action at turn tt, hth_t the history before it, and PP whatever the environment does: run a command, return a search result, respond as a user. Taking the log turns the product into a sum, and only the policy terms depend on θ\theta:

∇θlog⁡pθ(τ)=∑t∇θlog⁡πθ(at∣ht).\nabla_\theta \log p_\theta(\tau) = \sum_t \nabla_\theta \log \pi_\theta(a_t \mid h_t).

The environment's dynamics vanish from the gradient. The environment can be a sandboxed shell, a web browser, or another model, and it never needs to be modeled or differentiated. Its outputs still appear in the context hth_t that later actions condition on, but they contribute no gradient terms of their own. The policy-gradient loss implements this with a loss mask.

For a language model, each action is itself a sequence of tokens, so the sum runs over every sampled token yty_t in the rollout, each conditioned on its full context (x,y<t)(x, y_{<t}), observations included.

REINFORCE

Averaging over NN sampled rollouts gives REINFORCE (Williams, 1992):

g^=1N∑i=1NR(τi)∑t∈τi∇θlog⁡πθ(yt∣x,y<t).\hat g = \frac{1}{N}\sum_{i=1}^{N} R(\tau_i) \sum_{t \in \tau_i} \nabla_\theta \log \pi_\theta(y_t \mid x, y_{<t}).

This estimator is unbiased and very noisy. Two refinements are nearly always applied. Subtracting a baseline bb that does not depend on the sampled action leaves the expectation unchanged and, with a well-chosen baseline, reduces variance. This replaces RR with an advantage A=R−bA = R - b. And since an action cannot affect rewards that arrived before it, each token can be weighted by the reward that follows it (the reward-to-go). With the single terminal reward typical of LLM tasks, the reward-to-go at every token is RiR_i, so this second refinement changes nothing. If the baseline is also one number per rollout, as in GRPO, every token of rollout ii shares the advantage Ai=Ri−biA_i = R_i - b_i. A learned critic that estimates value at each position would instead give each token its own advantage.

A policy-gradient update

The gradient acts on an action that was actually sampled. Advantage determines the direction and strength of the update.

What one update does to the probabilities

Take a single decision with three candidate tokens and softmax probabilities π=(0.5,0.3,0.2)\pi = (0.5, 0.3, 0.2). The policy samples token 2 and the rollout receives advantage A=+1A = +1. For a softmax over logits zz, the gradient of the log-probability of the sampled token is

∇zlog⁡π(y=2)=e2−π=(−0.5,  0.7,  −0.2).\nabla_z \log \pi(y{=}2) = e_2 - \pi = (-0.5,\; 0.7,\; -0.2).

Gradient ascent with step size η\eta moves the logits by ηA (e2−π)\eta A\,(e_2-\pi): the sampled token's logit rises and the others fall in proportion to their current probability. With A=−1A = -1 every sign flips. With A=0A = 0 nothing moves, which is why uniform groups contribute no signal. Tokens that were already near probability 1 have ey−π≈0e_y - \pi \approx 0 and barely change, so most of the update lands on the uncertain decisions in a rollout.

In a full network the update goes through shared weights, so raising one sampled sequence's probability changes many other contexts too. A batch is a sum of such pushes, which can generalize, interfere, or concentrate the policy onto features shared across tasks.

In code

Automatic differentiation computes the estimator from a surrogate loss whose gradient equals −g^-\hat g, with advantages in place of rewards:

# logprobs: log π_θ(y_t | x, y_<t) for every token, shape [batch, seq]
# mask: 1 on tokens the policy sampled, 0 on prompt and observations
# advantages: one value per rollout, broadcast over its tokens
N = logprobs.shape[0]
loss = -(advantages[:, None] * logprobs * mask).sum() / N
loss.backward()

This sums each rollout's token contributions and averages over rollouts, as in g^\hat g. Many implementations instead divide by the number of sampled tokens. That token-mean normalization is a different estimator, and the choice of denominator changes what is optimized, as covered under length bias. The value of the loss is not the objective and does not track learning progress; only its gradient matters.

The on-policy assumption

The derivation samples τ\tau from πθ\pi_\theta, the same parameters being differentiated. Once the policy takes a step, the old rollouts come from a slightly different distribution. Reusing them, running several optimizer steps per batch, or generating asynchronously all break the assumption. On-policy and off-policy RL describes the distinction, and importance sampling and clipping are the corrections built on top of the plain gradient.