Understanding On-Policy Distillation

A Mechanistic Interpretability Perspective via Sparse Crosscoders

Zichao Yu1 Qianshuo Ye2 Xu Wang1,3 Difan Zou1,3,

1School of Computing and Data Science, The University of Hong Kong 2Department of Computer Science and Technology, University of Cambridge 3Shenzhen Loop Area Institute

Corresponding author

TL;DR. This work reveals that on-policy distillation improves reasoning primarily by reshaping how students use existing representations, offering a mechanistic account of what stronger teachers actually teach.
Overview: OPD's training signal, the swap readout, and what changes in the student
Overview. (a) In OPD, the student writes a rollout and the teacher scores every token; the signal is strongest at decision tokens. (b) The swap readout places a student checkpoint in both student slots of the crosscoder, with the teacher's activation fixed, and reads the features the checkpoint uses. (c) OPD reweights the features the student shares with the teacher and acquires none of the teacher's own; the warm-up moves the student partly as OPD would and partly in ways OPD does not.

Findings at a glance

1

OPD creates no new features

Across three OPD settings, no feature is gained or lost, and over 98% of the student's frequently used features change their firing rate by less than 20%.

2

The teacher's own features stay with the teacher

The 7B teachers dominate 42 and 30 features. OPD does not pass them on: the student's share of their decoder norm is the same before and after OPD.

3

Large changes sit at decision tokens

Features for words such as Wait, Hmm, and So are strongly over-represented among the features OPD changes most, and the teacher disagrees with the student most at these words.

4

The SFT warm-up adds no features either

SFT on the teacher's own rollouts, which makes OPD more effective, keeps every feature shared and gives the student none of the teacher's own features.

5

It reweights the shared features in two ways

First, it already raises and lowers many of the features that OPD later raises and lowers, doing part of OPD's work in advance. Second, it changes features in ways that OPD alone would not, most notably those for the conversation format, the style of reasoning, and mathematical notation, and these changes persist through OPD.

6

This reweighting carries the warm-up's benefit

Imposing the warm-up's feature change on a directly distilled student, without changing any weights, recovers most of the warm-up's benefit; removing it from the warmed-up student removes most of it.

The swap readout

A sparse crosscoder learns one dictionary of features for the student before OPD, the student after OPD, and the teacher, so that the three can be compared feature by feature. Its joint encoding, however, gives a single code for all three models, and decoder-based attribution barely differs between two nearly identical students. Neither can tell how training changes when the student uses a feature.

The swap readout places the activation of one student checkpoint in both student slots and keeps the teacher's activation in the teacher slot. Because the two students' activations are nearly identical in the training data, the crosscoder determines only the sum of their encoders; with the same activation in both slots, the undetermined part cancels. The readout therefore reads each checkpoint on its own, including checkpoints the crosscoder never saw.

Four outcomes of the swap readout for one feature on one input
Reading the student before and after training on the same input gives, for every feature, one of four outcomes.
The swap readout on held-out reasoning passages
The swap readout on held-out reasoning text (JustRL). Shading shows a feature's activation; bold tokens start or stop firing after OPD. The Wait feature stops firing on some occurrences of Wait, and the maybe/perhaps feature starts to fire on hedging phrases, in both cases moving toward the teacher's firing count.

OPD reweights shared features rather than acquiring new ones

We distill a DeepSeek-R1-Distill-Qwen-1.5B student from three teachers that differ from it by RL, by scale, or by both, and train one crosscoder per setting on the base student, the OPD student, and the teacher.

SettingTeacherDiffers from the student by
JustRLJustRL-DeepSeek-1.5BRL
SkyworkSkywork-OR1-Math-7BRL and scale
R1-7BDeepSeek-R1-Distill-Qwen-7Bscale
MAS distributions and change in feature usage after OPD
(a–c) The OPD student dominates no feature, and its model attribution scores match those of the base student; the teachers' own features stay with the teachers. (d) After OPD, 98.2%, 99.6%, and 99.4% of frequently used features change their firing rate by less than 20%.

Large changes sit at decision tokens

Decision-token features make up 0.4% of all features but 12% of the 50 features OPD changes most under JustRL and 4% under Skywork. OPD moves these features toward the teacher's usage, and it is at these tokens that the teacher disagrees with the student most.

Decision-token features among the most-changed features
(a) Share of decision-token features among the most-changed features. (b) Under JustRL, OPD moves the decision-token features toward the teacher. (c) Teacher–student disagreement at the steps that emit each category of decision token, relative to the average step.

Why an SFT warm-up helps OPD

An SFT warm-up on the teacher's rollouts commonly precedes OPD and makes it more effective. Following Simple-OPD, we distill Qwen3-1.7B-Base from Qwen3-4B-Base-GRPO: with the warm-up, OPD recovers 46% of the teacher's advantage over the base student instead of 30%. A natural explanation is that the warm-up supplies new features that OPD cannot. It does not.

NRN distributions before and after the warm-up
In two-model crosscoders, the teacher's share of the decoder norm is distributed the same way before and after the warm-up (a, b), and every feature of the warmed-up student stays shared with the student before it (c).

Along and beyond OPD's direction

Instead, the warm-up reweights the shared features. To see how this relates to OPD, we plot, for every frequently used feature, its change in firing rate after the warm-up against its change after OPD, both relative to the base student. If the warm-up merely did part of OPD's work, the dots would lie on a line through the origin. We therefore split the warm-up's change into two parts: the part along OPD's direction, a scaled copy of OPD's change fitted by least squares, and the part beyond it, the scatter of the dots around this line.

0.46slope along OPD's direction: the warm-up already moves the features OPD moves, about half as far
56%of the warm-up's change lies beyond OPD's direction
0.69Spearman correlation of that part after the warm-up and after the warm-up and OPD: it persists

The part along OPD's direction does some of OPD's work in advance: of the 50 features OPD raises and lowers most, 88% and 92% already move the same way after the warm-up. The part beyond it is spread over many features; those it moves most concern the conversation format, the style of reasoning (features for Wait, Alternatively, and But fire less, and one for laying out a plan step by step fires more), and mathematical notation. A random shift of the same size produces no alignment (slope −0.03).

Warm-up change against OPD change for every feature
(a) Each dot is a feature: its change after the warm-up against its change after OPD. The fitted line (slope 0.46) is the part along OPD's direction, and the scatter around it is the part beyond; a random shift of the same size shows no alignment. (b) Mean change of the 50 features OPD raises and lowers most: the warm-up already moves them part of OPD's way.

A causal test: moving the reweighting between students

At every token, we decode the difference between the swap-readout codes of the student after and before the warm-up and add it to the residual stream of the directly distilled student, or subtract it from the student distilled after the warm-up, without changing any weights. The same change with the feature identities shuffled serves as a control.

StudentAIME24AIME25AMC23Avg.Δ [95% CI]
OPD10.47.537.818.6–
+ feature change11.27.145.621.3+2.7 [+0.5, +6.0]
+ shuffled change7.54.635.615.9−2.7 [−5.4, +0.3]
Warm-up + OPD12.59.244.121.9–
− feature change10.07.140.619.2−2.7 [−5.3, −0.4]
− shuffled change14.29.246.923.4+1.5 [−0.8, +4.1]

avg@8 (%) with 8 samples per problem. Adding the warm-up's feature change brings the directly distilled student close to the warmed-up one; removing it brings the warmed-up student close to direct OPD. The shuffled change does neither.

Takeaway. OPD behaves more like a reweighting of existing features than an acquisition of new ones: the student learns from the teacher how to use the features they already share.

What the features encode

Each card lists a feature's strongest held-out contexts in the student before training, shaded by the feature's activation, with its firing counts on the right.

Decision-token features that OPD reweights under JustRL
Decision-token features that OPD reweights under JustRL. They fire most strongly on the decision words themselves, and OPD moves each of them toward the teacher's firing count.
Features that the warm-up moves along OPD's direction
Features that the warm-up moves along OPD's direction (Qwen3). Six of the features OPD changes most. The warm-up moves each of them the way OPD does, by a smaller amount; they fire when the trace settles on an answer after a long search, names a solution method, or ends an attempt that does not work out, or on Wait and arithmetic within equations.
Features that the warm-up moves beyond OPD's direction
Features that the warm-up moves beyond OPD's direction (Qwen3). They concern the conversation format, the style of reasoning, and mathematical notation. Their change after the warm-up is not a scaled-down copy of OPD's: it appears where OPD barely moves them, runs against OPD, or goes further than OPD.

Citation

A BibTeX entry will be added when the paper is public.