Teachers and privileged context

Self-distillation is on-policy distillation in which the teacher is the student model itself, conditioned on privileged context: information such as a demonstration or a verified solution that the student does not see. It gives a dense per-token signal on tasks that come with reference solutions but no reward function, without serving a second model.

A model that sees the answer, a worked example or a verified reasoning trace is a better policy than the same model without it, because in-context learning uses that information. Self-distillation turns the gap into a training signal. The student samples a rollout from the plain prompt, the same weights score those tokens with the privileged information prepended, and the student is trained toward the informed distribution.

Teacher targets on student-generated contexts

The sampled tokens come from the student. The target distribution is evaluated by the teacher at the same positions.

Worked example

Consider a math problem whose answer is 7. The student writes "the answer is" and then samples "9" with probability 0.30, where it gave "7" only 0.20. The teacher, the same model after reading a worked solution, scores that same "9" at 0.02. The per-token signal log⁡0.02−log⁡0.30≈−2.7\log 0.02 - \log 0.30 \approx -2.7 pushes the student away from its guess at the position where it went wrong. No reward function was involved: the signal came from the demonstration alone.

The update is the one on-policy distillation uses with an external teacher. The student's own tokens are scored, the teacher's log-probability minus the student's becomes a per-token advantage, and the objective is the reverse KL at the positions the student visited. Only the source of the teacher distribution differs.

Why a self-teacher

A self-teacher shares the student's tokenizer, style and habits, so the KL spends its signal on the difference the privileged context creates instead of on incidental differences between model families. It needs no second deployment. A target built on the student's own rollouts also stays nearer the starting model than demonstrations written by someone else, and forgetting tracks that distance.

What the teacher is shown

Two published methods differ mainly in the privileged input.

SDFT (Self-Distillation Fine-Tuning, Shenfeld et al.) conditions the teacher on an expert demonstration. It targets learning from demonstrations without the forgetting that SFT on those demonstrations causes, and reports higher new-task accuracy than SFT with much less degradation of prior capabilities, including when one model learns a sequence of skills.

OPSD (On-Policy Self-Distillation, Zhao et al.) conditions the teacher on a verified reasoning trace, with the student seeing only the question. It targets reasoning datasets that already contain ground-truth solutions. On math benchmarks it beat SFT on the same traces and matched GRPO while sampling one 1,024-token rollout per problem, where GRPO sampled eight rollouts of 16k tokens.

Where privileged teachers go wrong

  • Leakage. A teacher that knows the answer may assign high probability to stating it early and skipping the reasoning. The student then learns to assert conclusions it cannot yet derive. The privileged input has to help choose continuations without making the task trivial for the teacher.
  • A moving target. The teacher is the live policy, so it changes as training updates the weights. Its advice improves as the student improves, but the target is not stationary.
  • Distance from the student. Privileged context can make the teacher more correct while shifting its distribution toward text the student cannot produce unaided. Signal spent on those tokens teaches little.
  • Wrong demonstrations. A demonstration can be wrong, and a conditioned model can misread it. The student has to be evaluated on the capability itself.

An external teacher trades these problems for others: stronger predictions, but different tokenization and style and a second model to serve. A verifier can be layered on either kind to select the failing rollouts where the teacher's dense signal is spent.