SCOPD self-distillation narrows the accuracy loss from visual token pruning

SCOPD self-distillation narrows the accuracy loss from visual token pruning

Reasoning vision-language models (VLMs) read images and video as long sequences of visual tokens, which makes inference expensive. One common fix is training-free token pruning, which drops many of those tokens. The catch is that aggressive pruning can sharply degrade performance, and that drop is often blamed on the irreversible loss of task-relevant visual information.

The paper says that explanation is incomplete. In a fixed-context Pass@K analysis, the authors sample repeatedly from the same pruned visual representation. That repeated sampling recovers many examples that greedy decoding missed. Their reading is that useful visual evidence can remain accessible in the pruned input, but the model uses it unreliably. They call this the representation-utilization gap.

Building on that observation, they introduce SCOPD, a sparse-context on-policy self-distillation framework. A student model generates reasoning trajectories from the pruned visual tokens. A privileged teacher that sees the full context then supervises the same on-policy prefixes. According to the authors, SCOPD needs no ground-truth responses, no architectural changes and no additional inference-time computation.

They also introduce SCOPD+. It uses a small visual-budget intervention to find the response positions that are sensitive to the visual input, and then distills selectively at those positions.

The headline result is at 10% visual-token retention, measured across 13 benchmarks as the share of the unpruned model's performance that is kept. The Vanilla model retains 86.37%. SCOPD raises that to 90.49%, and SCOPD+ to 92.43%. These are shares of unpruned performance, not absolute benchmark scores. The authors conclude that, across token budgets, benchmarks and pruning operators, efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.

Key facts

  • Pruning visual tokens cuts VLM inference cost, but aggressive pruning can sharply hurt performance; the authors say the usual blame on lost visual information is incomplete.
  • In a fixed-context Pass@K analysis, repeated sampling from the same pruned representation recovers many examples greedy decoding missed. The authors call this the representation-utilization gap.
  • SCOPD has a student reason from pruned visual tokens while a full-context teacher supervises the same on-policy prefixes, with no ground-truth responses, architecture changes or extra inference-time compute.
  • SCOPD+ adds a small visual-budget intervention to find visually sensitive response positions and distill only those.
  • At 10% visual-token retention across 13 benchmarks, retained performance goes from 86.37% (Vanilla) to 90.49% (SCOPD) and 92.43% (SCOPD+).

Why it matters

Visual tokens are a large part of what makes reasoning VLMs costly to run, and pruning them is an appealing shortcut. This paper challenges the standard explanation for why heavy pruning hurts. If the information is often still there but used unreliably, then the loss can be partly repaired by training the model to use what remains, rather than only by choosing better tokens to keep. The authors frame this as the representation-utilization gap.

Who it affects

Researchers and engineers working on efficient inference for vision-language models, especially those who apply token pruning to image or video reasoning models. The method targets the training side, so it is relevant to anyone who can fine-tune the model they deploy.

How to use it

The abstract describes the recipe, not a release. In SCOPD, the student generates reasoning from pruned visual tokens and a full-context teacher supervises the same prefixes. SCOPD+ adds a small visual-budget intervention to pick the response positions worth distilling. No code or model release is mentioned.

How solid is it

The evidence here is the paper's abstract. The headline figures are clear: at 10% visual-token retention across 13 benchmarks, retained performance rises from 86.37% to 90.49% with SCOPD and 92.43% with SCOPD+. The authors say results hold across token budgets, benchmarks and pruning operators. The abstract names no authors or institutions, no model names or sizes, and not the 13 benchmarks. It gives no results for token budgets other than 10%.

Risks and caveats

The figures are shares of unpruned performance, not absolute benchmark scores, so they do not show how good the models are in raw terms. The abstract does not say how much inference cost or speed is saved, nor how much training compute SCOPD needs. It also does not say which pruning operators were tested. The claims about the representation-utilization gap and the method's benefits are the authors' own.

“We call this the representation-utilization gap.”

— From the paper's abstract