PMOPD protects task subspaces in multi-teacher distillation, beating MOPD on Qwen2.5-7B and Llama-3.1-8B

PMOPD protects task subspaces in multi-teacher distillation, beating MOPD on Qwen2.5-7B and Llama-3.1-8B

Multi-teacher on-policy distillation (MOPD) is, in the paper's words, a popular post-training paradigm for integrating specialized capabilities in frontier language models. The authors argue that earlier on-policy distillation work has mostly focused on single-task distillation: objective design, distillation scope and how the teacher signal is built. MOPD has a different problem. It must pack several capabilities into one shared set of parameters, and that produces what the paper calls a capability seesaw, where improving one domain suppresses capabilities acquired from another.

The authors' starting observation concerns the geometry of the updates. They find that parameter updates from different tasks rapidly concentrate in their own low-dimensional subspaces during MOPD. They say this gives a direct geometric basis for identifying and controlling cross-task interference.

On that basis they propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation). It builds subspace memories from the cumulative parameter displacements of the different tasks. It then projects both gradients and optimizer updates to remove the components that interfere with protected task directions.

Two further pieces go with it. One is a lightweight conflict probe that characterizes how tasks interact and guides the order in which tasks are trained. The other is a cycling strategy that balances subspace estimation against timely revisiting of each task.

The experiments cover representative Code, Reason and Math tasks. The authors report that PMOPD improves every evaluated capability over MOPD. The average score across the three tasks rises by 2.54 points on Qwen2.5-7B and by 2.09 points on Llama-3.1-8B. They conclude that these consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.

Key facts

  • PMOPD (Projection-based Multi-Teacher On-Policy Distillation) targets the capability seesaw in MOPD, where improving one domain suppresses capabilities acquired from another.
  • The authors find that parameter updates from different tasks rapidly concentrate in their own low-dimensional subspaces during MOPD.
  • The method builds subspace memories from each task's cumulative parameter displacement and projects both gradients and optimizer updates to remove components that interfere with protected task directions.
  • A lightweight conflict probe guides task ordering, and a cycling strategy balances subspace estimation with timely task revisitation.
  • Average score across Code, Reason and Math rises by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B over MOPD, with every evaluated capability improved.

Why it matters

Multi-teacher on-policy distillation is used to fold several specialist capabilities into one model, and the paper says it has become a popular post-training paradigm for frontier language models. The weak spot it names is the capability seesaw: gains in one domain come at the cost of another. PMOPD attacks that with an explicit geometric mechanism, protecting the parameter directions each task has used, rather than only changing the objective or the teacher signal as earlier single-task OPD work did.

Who it affects

The work is aimed at teams that post-train language models by distilling from several specialized teachers into one student, for example a code teacher, a reasoning teacher and a math teacher. The reported results are on two student models: Qwen2.5-7B and Llama-3.1-8B.

How to use it

The recipe, as described, has four parts. Keep a subspace memory for each task, built from that task's cumulative parameter displacement. Project both gradients and optimizer updates to strip components that interfere with protected task directions. Use the lightweight conflict probe to decide task ordering. Use the cycling strategy to revisit tasks in time while still estimating subspaces. No code release or publication venue is mentioned.

How solid is it

The claims come from the authors' own experiments on Code, Reason and Math tasks with two models. The reported gains are absolute score points on the three-task average (2.54 on Qwen2.5-7B, 2.09 on Llama-3.1-8B), not relative percentages. No benchmark names, per-task scores or baseline absolute scores are given; only the average gains across the three tasks. The abstract does not say whether the gains are statistically significant or how many runs were made.

Risks and caveats

Only two models are named (Qwen2.5-7B and Llama-3.1-8B); results on larger models are not reported. No compute cost, training time or memory overhead of PMOPD is stated. The number of teachers, the dimensionality of the subspaces and the cycling schedule are not given. The claim that the approach is transferable is the authors' own conclusion from these two models.