Smaller frozen models make better rejects for preference distillation

Preference distillation usually treats a teacher's response as the preferred answer and the student's own response as the rejected one. That setup rests on two assumptions: that self-generated failures are the most informative negatives, and that rejects must come from a model at least as large as the student, which makes generation costly at scale. The authors of this paper report that neither assumption holds.
Across students from 7B to 72B parameters, smaller frozen models generate rejects with less inference compute, yet the students trained on them end up stronger than students trained on self-generated rejects. The result holds before and after sequence-level knowledge distillation, and it was observed on code generation and mathematical reasoning.
To explain this, the authors derive a finite-horizon utility bound for Direct Preference Optimization (DPO) in a linearized feature model. The bound characterizes which reject distributions are favorable and motivates three interventions.
First, mixing rejects from smaller and student-scale models improves performance as the smaller model's share increases. Second, rejects that are reassigned to other prompts and have their code tokens shuffled still outperform length-matched gibberish, which the authors take to show that task structure contributes to a reject's usefulness. Third, selecting candidates with lower likelihood under the reference policy improves net transfer when higher-likelihood candidates provide less useful contrast. Lower-likelihood selections outperform higher-likelihood ones for every source.
The authors conclude that effective rejects preserve task structure while limiting coupling to the reference policy, and that smaller frozen models can provide such rejects at low cost.
Key facts
- Rejects from smaller frozen models train stronger students than self-generated rejects, using less inference compute, on code generation and mathematical reasoning.
- The effect was observed across student sizes from 7B to 72B, before and after sequence-level knowledge distillation.
- The authors derive a finite-horizon utility bound for DPO in a linearized feature model and use it to motivate three interventions.
- Rejects with shuffled code tokens, reassigned to other prompts, still beat length-matched gibberish, which points to task structure as a source of reject utility.
- Lower-likelihood candidates under the reference policy outperform higher-likelihood ones for every source.
Why it matters
The common recipe for preference distillation pairs the teacher's answer with the student's own failed answer. The paper challenges two assumptions behind it: that self-generated failures are the most informative negatives, and that rejects must come from a model at least as large as the student. If smaller frozen models can supply rejects that train stronger students with less inference compute, producing negative examples at scale gets cheaper.
Who it affects
The finding is aimed at teams that distill a stronger model's behavior into a student using preference data such as DPO. It was tested on students from 7B to 72B parameters, on code generation and mathematical reasoning. Those are the settings where the claim has been shown.
How to use it
The abstract points to three levers. Mix rejects from smaller models with rejects from student-scale models, since performance improves as the smaller model's share increases. Prefer candidates with lower likelihood under the reference policy, which outperformed higher-likelihood ones for every source. Keep the task structure of the rejects intact, because rejects reassigned to other prompts with shuffled code tokens still beat length-matched gibberish.
How solid is it
This rests on the abstract of a paper alone. The reported results are comparative and no numerical results, such as accuracy gains or compute savings, are given. The theory is a finite-horizon utility bound derived in a linearized feature model, not for real full-scale networks. The three interventions test the bound's predictions empirically. The authors word the conclusion cautiously: the results "suggest" that effective rejects preserve task structure and limit coupling to the reference policy.
Risks and caveats
The abstract does not name the smaller models, teacher models, datasets or benchmarks, and does not state how small the reject-generating models were. The amount of inference compute saved is not quantified. Evidence covers code generation and mathematical reasoning only, so it should not be assumed to carry over to other domains. No release of code, models or data is mentioned.
“We find neither assumption holds”
— From the paper's abstract