Distillation study finds on-policy rollouts matter less than KL direction and learning rate

On-policy learning, where a model trains on data it generated itself, has been argued to reduce catastrophic forgetting, produce sparser parameter updates and improve generalisation. A new paper questions how much of that credit belongs to the rollout policy. Its authors point out that existing comparisons between supervised fine-tuning and reinforcement learning change many factors at once, which makes the contribution of rollout policy hard to isolate.\n\nTo separate it out, they set up a controlled strong-to-weak distillation experiment. They vary three things independently: the rollout policy, the token-level KL direction and the learning rate. The experiments cover the Llama3 and Qwen2.5 model families and reasoning tasks in scientific, medical and arithmetic domains.\n\nThe result is what the authors call a nuanced picture. Rollout policy does not necessarily play a central role. Token-level KL direction shapes task performance and output coverage more clearly, while learning rate governs forgetting and update sparsity.\n\nThe authors explain this pattern through an analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum. Forward KL turns out to be remarkably robust to rollout policy: its performance stays stable and strong even as the rollout policy changes. Reverse KL is substantially more sensitive and favours student-generated rollouts.\n\nOn-policy data still helps in one place. It improves generalisation to harder variants of the Countdown arithmetic task under both KL directions. That advantage, however, does not reliably persist after subsequent RLVR (reinforcement learning with verifiable rewards).\n\nThe broader conclusions hold when gradient clipping is removed, when sampled KL estimators are used, and when training on tasks that need longer reasoning chains. Overall, the authors say their results challenge the view that on-policy rollouts are inherently preferable. Their value, they write, depends critically on the objective, the evaluation setting and the optimisation hyperparameters.
Key facts
- The study isolates rollout policy in a controlled strong-to-weak distillation setting, varying rollout policy, token-level KL direction and learning rate independently.
- Experiments span the Llama3 and Qwen2.5 model families and scientific, medical and arithmetic reasoning tasks.
- Token-level KL direction more clearly shapes task performance and output coverage; learning rate governs forgetting and update sparsity.
- Forward KL is robust to rollout policy, while reverse KL is substantially more sensitive and favours student-generated rollouts.
- On-policy data improves generalisation to harder Countdown variants under both KL directions, but the gain does not reliably persist after subsequent RLVR.
Why it matters
A common assumption is that training on a model's own samples is better than training on someone else's, with benefits for forgetting, update sparsity and generalisation. This paper tests that assumption with the other variables held apart, and finds that rollout policy is not necessarily central. If the finding holds, the choice of KL direction and learning rate may deserve more attention than the on-policy versus off-policy debate.
Who it affects
Researchers and engineers who distil large models into smaller ones, and anyone choosing between supervised fine-tuning and reinforcement learning style post-training. The experiments use Llama3 and Qwen2.5 families, so teams working with those model families have the most direct evidence.
How to use it
The abstract points to practical levers. Forward KL keeps performance stable and strong whatever the rollout policy, so it is the safer choice when generating student rollouts is costly or inconvenient. Reverse KL favours student-generated rollouts, so pairing it with off-policy data is the less promising combination. Learning rate is the knob tied to forgetting and update sparsity. If the target is harder Countdown-style variants, on-policy data helped under both KL directions.
How solid is it
The setup is controlled, and the authors report that the broader conclusions survive removing gradient clipping, using sampled KL estimators and training on tasks with longer reasoning chains. They also back the pattern with an analysis of KL gradients and a continuous student-teacher rollout-policy spectrum. The abstract itself gives no numerical results, model sizes or teacher-student pairs, so the size of the effects cannot be judged from it.
Risks and caveats
The on-policy advantage on harder Countdown variants does not reliably persist after subsequent RLVR, so the picture is mixed rather than a blanket dismissal of on-policy data. The authors say the value of on-policy rollouts depends on the objective, evaluation setting and optimisation hyperparameters, so the conclusions should not be assumed to carry over to other setups. The tasks studied are scientific, medical and arithmetic reasoning in two model families.
“our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters”
— From the paper's abstract