GPD distills 3D geometry cues into RGB-only vision-language models for spatial reasoning

GPD distills 3D geometry cues into RGB-only vision-language models for spatial reasoning

Spatial reasoning remains a persistent weakness of vision-language models (VLMs), the authors write, because RGB inputs do not directly provide geometric evidence. They name two existing remedies and their costs: injecting 3D into the model at inference, which brings architecture and latency costs, or training with outcome rewards that supervise only the final answer.

The paper starts from the observation that spatial errors originate in perception. A misjudged depth or direction can be corrected only by the scene's true geometry, and the 3D-scanned sources of spatial training corpora already provide that geometry.

The proposed method is GPD, short for Geometry-Privileged Distillation. It makes geometric evidence the privilege in on-policy self-distillation (OPSD). For each question, depth, semantic and bird's-eye-view (BEV) cues are rendered as compact text and routed to the teacher alongside the reference answer. A privileged KL term, applied only to incorrect trajectories, augments GRPO. The deployed model remains RGB-only, so nothing extra is needed at inference.

On the 4B backbone, GPD reaches 57.1 on VSI-Bench and a 37.6 average across MindCube, SPARBench, MMSI-Bench and ViewSpatial. The authors report that this outperforms both GRPO and answer-privileged OPSD across spatial reasoning benchmarks.

The ablations, as the authors describe them, confirm three things: 3D privilege and answer privilege complement each other; question-conditioned routing of the cues beats injecting the full context; and restricting distillation to incorrect trajectories helps.

Key facts

  • GPD (Geometry-Privileged Distillation) gives the teacher in on-policy self-distillation depth, semantic and bird's-eye-view cues rendered as compact text, alongside the reference answer.
  • A privileged KL applied only to incorrect trajectories augments GRPO, and the deployed model remains RGB-only.
  • On the 4B backbone, GPD reports 57.1 on VSI-Bench and a 37.6 average across MindCube, SPARBench, MMSI-Bench and ViewSpatial.
  • The authors say it outperforms both GRPO and answer-privileged OPSD across spatial reasoning benchmarks.
  • Ablations favour question-conditioned routing over full-context injection and favour restricting distillation to incorrect trajectories.

Why it matters

Spatial reasoning is a known weak spot of VLMs, since a plain RGB image does not directly carry geometric evidence. The usual fixes carry a price: 3D inputs at inference add architecture and latency costs, and outcome-only rewards supervise just the final answer. GPD moves the 3D information to training time. The geometry helps the teacher, and the model that ships still takes only RGB. The paper also targets where the errors begin, in perception, rather than only rewarding correct final answers.

Who it affects

Researchers and engineers who train VLMs for spatial reasoning and evaluate on benchmarks such as VSI-Bench, MindCube, SPARBench, MMSI-Bench and ViewSpatial. It is also relevant to anyone who wants spatial gains without adding a 3D module or extra latency to the deployed model.

How to use it

The abstract describes a training recipe rather than a product. For each training question, render depth, semantic and BEV cues as compact text and give them to the teacher together with the reference answer. Apply a privileged KL only to incorrect trajectories and add it to GRPO. The deployed model stays RGB-only. The geometry comes from the 3D-scanned sources of spatial training corpora. No code, model or data release is mentioned in the source.

How solid is it

The material is the paper's abstract, so every result is the authors' own report. The headline numbers, 57.1 on VSI-Bench and a 37.6 average over four other benchmarks, are given for the 4B backbone. The authors say ablations support their design choices. No baseline scores for GRPO or answer-privileged OPSD are given, so the size of the improvement is not stated, and no units or score scale are stated for 57.1 and 37.6.

Risks and caveats

Results are reported only on the 4B backbone, and the abstract does not say which model family that is; no results are given for other backbones. No latency, compute or training-cost figures are given, so the cost of producing and routing the privileged cues is not quantified. The geometric cues come from the 3D-scanned sources of spatial training corpora, so the method depends on that geometry being available at training time. Without baseline numbers, claims of outperforming GRPO and answer-privileged OPSD cannot be sized from the abstract alone.

“a misjudged depth or direction can be corrected only by the scene's true geometry, which the 3D-scanned sources of spatial training corpora already provide”

— Paper abstract