On-policy distillation teaches reasoning skill, not new facts, paper finds

On-policy distillation (OPD) is known to strengthen the reasoning of language models. What a student model actually picks up from it, new factual knowledge or compositional skill for multi-step reasoning, was an open question, according to the paper's authors.
To separate the two, the authors built a controlled synthetic framework. It measures what the student can already do, then independently controls what the teacher adds: extra facts, compositional skill, or both.
Across four models from three families, reverse-KL OPD reliably transferred compositional skill across unseen reasoning structures. It transferred minimal factual knowledge.
The authors then decoupled the distillation recipe to find the source of this asymmetry. Replacing reverse KL with forward KL restored factual transfer, while student rollouts specifically improved the execution of multi-step reasoning.
Experiments on recent factual QA and competition mathematics showed a similar asymmetry under reverse-KL OPD: notable reasoning gains without factual memory expansion.
The authors' conclusion is that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.
Key facts
- In a controlled synthetic framework, reverse-KL on-policy distillation reliably transferred compositional skill across unseen reasoning structures but transferred minimal factual knowledge.
- The result held across four models from three families.
- Replacing reverse KL with forward KL restored factual transfer, while student rollouts specifically improved the execution of multi-step reasoning.
- Experiments on recent factual QA and competition mathematics showed a similar asymmetry: notable reasoning gains without factual memory expansion.
- The authors conclude that OPD teaches a model to organize and compose knowledge it already has, rather than expanding its parametric knowledge.
Why it matters
OPD is used to strengthen language-model reasoning, but it was unclear what the student really gains from it. This paper gives a specific answer: under reverse-KL OPD the student gets better at carrying out multi-step reasoning, not at knowing more facts. It also points to where the split comes from. The choice of divergence (reverse versus forward KL) governs factual transfer, while student rollouts drive the reasoning improvement.
Who it affects
Anyone who trains or distills language models with on-policy distillation, and anyone deciding what a distillation step can be expected to deliver. A student distilled this way should be expected to reason better over what it already knows, not to arrive with the teacher's extra facts.
How to use it
This is a research finding, not a tool. The practical reading from the abstract: if the goal is better multi-step reasoning, reverse-KL OPD transferred that skill reliably in the authors' tests. If the goal is to move factual knowledge from teacher to student, reverse KL transferred little, and swapping in forward KL restored factual transfer. No code, data or model release is mentioned in the source.
How solid is it
The evidence rests on a controlled synthetic framework run on four models from three families, backed by experiments on recent factual QA and competition mathematics that show a similar asymmetry. The source is the paper's abstract. The four models and three model families are not named, no numerical results such as accuracy or gain sizes are given, and the specific factual QA and competition mathematics benchmarks are not named.
Risks and caveats
The conclusion is stated for the tested setups: the synthetic framework plus factual QA and competition mathematics. The abstract does not say it was tested on all OPD recipes or at larger scale. The claim that OPD does not expand parametric knowledge is the authors' own reading of their results, and without reported numbers the size of the asymmetry cannot be judged from the abstract alone.
“on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses”
— Paper abstract, Hugging Face Papers 2610.09639