Δ-MOPD distills several teachers by transferring logit shifts, not endpoint policies

Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain, and their signals are merged into a single target. In routed-domain distillation, prompts from different domains are assigned to the matching specialist teacher. The authors say both settings usually transfer each teacher's endpoint policy, which mixes what post-training changed with preferences the teacher inherited from its base model.
The paper introduces Δ-MOPD. Instead of the endpoint policy, it transfers each teacher's teacher-minus-base logit shift, re-anchored at the student's frozen initialization. The authors compare it with endpoint supervision in both settings while holding teacher selection fixed, so the only thing that changes is how the target is built.
They first point to a mechanism that they say impedes endpoint transfer: the inherited base pull can exceed the post-training shift. Removing that pull reduces the teacher-term norm ratio and the target-student KL.
The results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, Δ-MOPD exceeds endpoint composition by 4.11 points on Math and 1.95 points on a five-benchmark measure. With two composed teachers, it only matches endpoint accuracy.
Under phased routing, Δ-MOPD achieves higher mean performance in both phase orders, and it reduces the observed gap between the orders from 10.50 to 6.42 points. Under interleaved routing, where each update involves a single teacher, the two targets perform comparably. The authors read the phased results as supporting evidence that the benefit may extend to signals accumulated across training phases. Their conclusion is that target construction is an independent design axis in MOPD, complementary to teacher selection.
Key facts
- Δ-MOPD transfers each teacher's teacher-minus-base logit shift, re-anchored at the student's frozen initialization, instead of the teacher's endpoint policy.
- With three composed teachers, Δ-MOPD beats endpoint composition by 4.11 Math points and 1.95 five-benchmark points; with two teachers it only matches endpoint accuracy.
- Under phased routing it scores higher in both phase orders and narrows the observed order gap from 10.50 to 6.42 points.
- Under interleaved routing, where each update involves one teacher, the shift and endpoint targets perform comparably.
- The authors say inherited base pull can exceed the post-training shift, which is the mechanism that impedes endpoint transfer.
Why it matters
Multi-teacher distillation usually copies what each teacher does at the end of training. The authors argue that this endpoint mixes two things: what post-training added, and preferences the teacher already had from its base model. Δ-MOPD tries to separate them by transferring only the shift. The paper's broader claim is that how the target is constructed is a design choice of its own in MOPD, independent of which teachers are picked.
Who it affects
The work concerns people who distill a student from several teachers, either by combining their scores on the same rollout or by routing prompts from different domains to specialists. The source does not name any lab, model or product.
How to use it
The source describes a change to the distillation target, not a new pipeline: swap the teacher's endpoint policy for its teacher-minus-base logit shift, anchored at the student's frozen initialization. The reported gains came where teacher signals are combined at a state, with three teachers. With two teachers, or with interleaved routing where each update involves one teacher, the source reports no advantage over endpoint targets. No code release is mentioned.
How solid is it
This is an abstract-level account, and the authors themselves word the claims cautiously: the results "suggest" shift targets help when signals are combined, and the phased results are "supporting evidence" that the benefit "may" extend across training phases. The reported numbers are point differences (4.11, 1.95, and a gap of 10.50 versus 6.42 points). The source does not give baseline absolute scores, model names or sizes, or the names of the five benchmarks beyond Math, so the size of the gains in context cannot be judged from it.
Risks and caveats
The advantage is not general. With two composed teachers Δ-MOPD only matches endpoint accuracy, and under interleaved routing the two targets are comparable. The 10.50 to 6.42 figure is an observed order gap under phased routing, which shrinks but does not disappear. The source does not say which models or benchmarks were used, so how far the results carry to other setups is unknown.
“Target construction is thus an independent design axis in MOPD, complementary to teacher selection.”
— Δ-MOPD paper abstract