DN-MOPD rescales teacher feedback to improve multi-teacher distillation on Qwen3.5

DN-MOPD rescales teacher feedback to improve multi-teacher distillation on Qwen3.5

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions. Users, though, need one model that has all of these skills. Multi-teacher on-policy distillation (MOPD) is one way to merge them: the specialists teach a single student. The student answers each prompt, and the specialist for that prompt's domain gives feedback on every token.

The paper points to a gap in this setup. The routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, the authors find that MOPD's student does not beat one taught by the best single specialist, and it gains little of the mathematics specialist's advantage.

The authors trace this to unbalanced feedback. Instruction-following feedback is several times more spread out than mathematics feedback, and it dominates the student's updates. The source gives no exact factor, only "several times".

The proposed fix is Domain-Normalized MOPD (DN-MOPD). It keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits. It also recovers most of the lost mathematics gain.

The authors also ran controls with fixed domain weights. These show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone. Fixed weights close to those DN-MOPD measures perform comparably. The conclusion: combining specialists requires deciding not only which one teaches, but also how strongly its feedback counts.

Key facts

  • MOPD merges RL-trained specialists into one student: the student answers each prompt and the specialist for that prompt's domain gives feedback on every token.
  • In Qwen3.5 models at three sizes, MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage.
  • Instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates.
  • DN-MOPD keeps the routing but rescales each domain's feedback by its measured spread, and improves the average over MOPD on six public benchmarks at every size, across three seeds and two answer-length limits.
  • Fixed-weight controls show the gain comes mainly from turning down instruction-following feedback, and fixed weights close to DN-MOPD's measured ones perform comparably.

Why it matters

Turning several RL-trained specialists into one general model is a practical need, since users want one model with mathematics, coding and instruction-following skills together. The paper argues that routing prompts to the right teacher is not enough. If one domain's feedback is far more spread out than another's, it drowns out the rest, and the merged student loses much of what a specialist could offer. The point of the work is that the strength of each teacher's feedback is a separate design choice from which teacher is used.

Who it affects

Teams that train domain specialists with reinforcement learning and then merge them into a single model through multi-teacher on-policy distillation. The experiments were run on Qwen3.5 models at three sizes, so the findings speak most directly to that family of setups.

How to use it

The recipe in the abstract is simple to state: keep the existing routing of prompts to domain specialists, measure the spread of each domain's feedback, and rescale that feedback accordingly. The controls also suggest a cheaper route: fixed weights close to those DN-MOPD measures perform comparably. The source does not mention any code or model release.

How solid is it

The evidence rests on the authors' own reported results: a higher average score than MOPD on six public benchmarks at every tested size, across three random seeds and under two answer-length limits. Fixed-weight controls back up the explanation that the gain comes mainly from turning down instruction-following feedback. The source gives no numeric benchmark scores or size of the improvement over MOPD, so the magnitude cannot be judged from the text.

Risks and caveats

The three Qwen3.5 model sizes, the six benchmarks and the two answer-length limits are not named in the source, which limits how far the results can be checked or generalised. The exact factor by which instruction-following feedback is more spread out than mathematics feedback is not given, and the spread measure is not specified beyond "measured spread". The finding that fixed weights perform comparably also means that the adaptive measurement itself may not be the only route to the gain.

“Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.”

— Paper abstract, Hugging Face Papers