LSPD brings RL-style exploration to on-policy distillation of LLMs

LSPD brings RL-style exploration to on-policy distillation of LLMs

The paper studies on-policy distillation (OPD), a way of training a student language model against a teacher, through the lens of reinforcement learning. Its first step is theoretical: the authors establish a connection between the reverse-KL objective used in OPD and KL-regularized policy optimization.

Building on that link, they introduce Least-Square Policy Distillation (LSPD). It is described as an RL-inspired framework that brings two ideas from value-based RL into policy distillation: optimistic exploration and off-policy data reuse. The authors say LSPD preserves policy diversity through exploration, and improves rollout efficiency by repeatedly learning from trajectories that were collected earlier.

On the theory side, the analysis connects LSPD to optimistic value-based learning. It shows that the idealized formulation of LSPD achieves a regret bound of O(log K) under online exploration. This is a theoretical result about the idealized version, not a measured one.

On the empirical side, the authors report that LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with an average gain of +1.59 points in Avg@16 (absolute points, not a relative percentage). They also ran Pass@k evaluations up to k=64 and found that LSPD better preserves policy diversity, achieving stronger performance as k grows. Finally, the fully off-policy variant reaches performance comparable to vanilla OPD using only the first 25% of rollout batches. Comparable, not better.

The authors sum this up as an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.

Key facts

  • The paper links the reverse-KL objective in on-policy distillation (OPD) to KL-regularized policy optimization.
  • Least-Square Policy Distillation (LSPD) adds optimistic exploration and off-policy data reuse from value-based RL to policy distillation.
  • Across six mathematical reasoning benchmarks and diverse teacher-student settings, LSPD averages +1.59 points in Avg@16 over existing distillation baselines.
  • In Pass@k evaluations up to k=64, LSPD keeps improving as k grows, which the authors read as better preserved policy diversity.
  • The fully off-policy variant matches vanilla OPD's performance using only the first 25% of rollout batches; the O(log K) regret bound applies to the idealized formulation.

Why it matters

Distilling a strong teacher into a smaller student is a common route to better reasoning models, and on-policy distillation is one form of it. This paper offers a way to read OPD as a reinforcement-learning problem, tying its reverse-KL objective to KL-regularized policy optimization. That framing lets the authors borrow tools from value-based RL: optimistic exploration to keep the student's outputs diverse, and off-policy reuse of old trajectories to make each rollout go further. The stated goal is more effective and more rollout-efficient distillation.

Who it affects

The work is aimed at researchers and engineers who distill reasoning ability into language models, especially for mathematical reasoning, which is where the evaluation is done. It also speaks to anyone who cares about the cost of generating rollouts during distillation, since the efficiency claim is framed in terms of rollout batches.

How to use it

The source is a short abstract, so it gives no implementation recipe. The practical idea it describes is to keep rollouts already collected and learn from them repeatedly, while using optimistic exploration to avoid collapsing the student's diversity. It mentions no code release, compute cost or wall-clock savings, so a team would need the full paper to try the method.

How solid is it

The claims come from the authors' own abstract. The evidence has two parts: a theoretical analysis and an empirical comparison on six mathematical reasoning benchmarks in diverse teacher-student settings. The average gain of +1.59 points in Avg@16 is modest in size, and the source gives no per-benchmark results or baseline names, only that average. The Pass@k results up to k=64 support the diversity argument. The efficiency result is a comparison to vanilla OPD at parity, not an improvement.

Risks and caveats

The O(log K) regret bound is for the idealized formulation under online exploration; the source does not say whether it holds for the practical algorithm. No teacher or student model names or sizes are given, and the six benchmarks are not named, so it is hard to judge how far the gains generalize beyond math reasoning. The efficiency claim is stated only in terms of rollout batches, with no compute cost or wall-clock savings mentioned. The 25% figure describes the fully off-policy variant reaching comparable performance, not surpassing vanilla OPD.