Where-OPD: self-distillation on synthetic scenes lifts MLLM perception

A paper introduces Where-OPD, a form of on-policy self-distillation for multimodal large language models (MLLMs). On-policy self-distillation has recently become an effective way to improve language-model reasoning: a student is supervised by a frozen or EMA version of itself that receives privileged information. The authors say applying the idea to MLLMs remains largely unexplored.
They describe what exists so far. Recent approaches give the teacher privileged visual information, such as an image crop that corresponds to the question, to sharpen fine-grained perception. According to the authors, the gains from those methods are confined to tasks that benefit from visual zooming, and the methods need either human-annotated grounding data or external teacher models.
Where-OPD changes what the teacher is given. Instead of a crop, the teacher receives textual, spatially grounded guidance that identifies the visual elements relevant to a query. To produce that guidance without human labelling, the authors use procedurally generated scenes, where object identities and spatial coordinates are available automatically. They say this makes post-training scalable and annotation-free.
The teacher uses the spatial guidance to locate and integrate evidence from several relevant image regions. The student sees only the image and the question, and learns to reproduce the teacher's resulting behavior.
The authors report that the approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. The headline result concerns transfer: post-training uses only synthetic scenes, yet the improvements carry over to real-world perception benchmarks. The single quantified figure is a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench and MME-RealWorld. The authors conclude that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. A project page is listed at https://github.com/sirkosophia/Where-OPD.
Key facts
- Where-OPD is on-policy self-distillation for MLLMs: a teacher version of the model gets textual, spatially grounded hints about which visual elements matter, and the student learns from the image and question alone.
- Training data are procedurally generated scenes whose object identities and coordinates come for free, so no human annotation or external teacher model is needed.
- The authors report consistent gains on counting, document and chart understanding benchmarks across multiple models.
- Post-training on synthetic scenes only still yields a 3.23-point average gain across CVBench, V*, ZoomBench, BLINK, HR-Bench and MME-RealWorld.
- The authors contrast the method with earlier crop-based approaches, whose gains they say are confined to tasks that benefit from visual zooming.
Why it matters
Fine-grained perception, such as counting objects or reading charts and documents, is a known weak spot for multimodal models. Self-distillation has helped language-model reasoning, and this paper tries to carry it over to vision. Its twist is the kind of privileged information the teacher gets: text that says where the relevant things are, rather than an image crop. The authors say this avoids the limits of crop-based methods, whose gains they describe as confined to zoom-friendly tasks and dependent on human-annotated grounding data or external teachers. If the reported transfer from synthetic scenes to real benchmarks holds up, training signals could be generated cheaply and at scale.
Who it affects
Mainly researchers and engineers who post-train multimodal LLMs and care about perception benchmarks such as CVBench, V*, ZoomBench, BLINK, HR-Bench and MME-RealWorld. Teams that lack annotated grounding data may find the annotation-free setup of most interest.
How to use it
This is a research method rather than a product. The authors list a project page at https://github.com/sirkosophia/Where-OPD. The recipe as described: generate procedural scenes with known object identities and coordinates, give a teacher copy of the model text guidance about the relevant elements, and train the student to match the teacher's behavior from the image and question alone.
How solid is it
The claims come from the authors' own description of their work, and the text carries a single quantified result: a 3.23-point gain in average performance (absolute points) across six real-world benchmarks after post-training only on synthetic scenes. Per-benchmark scores and baseline values are not given, and it is not stated whether the 3.23-point gain applies to every model, to one model, or is an average over models. The size of the improvements on counting, document and chart benchmarks is not quantified. The specific MLLMs used are not named.
Risks and caveats
With only an aggregate gain and no named models or baselines in the text, the practical size of the effect is hard to judge. The claim of broader perceptual capability and substantial synthetic-to-real transfer is the authors' own conclusion. Training compute, dataset size and the number of synthetic scenes are not stated, so the cost of reproducing the result is unclear.
“These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation”
— From the paper's abstract