SpaceCast-Bench: best of 21 vision-language models scores 58.0% on predictive spatial reasoning vs 87.2% for humans

A paper introduces SpaceCast-Bench, a benchmark for what the authors call predictive spatial reasoning in vision-language models. Their argument is that existing spatial reasoning benchmarks mainly test spatial perception, meaning reading off relations that are already visible in the input. Real-world spatial intelligence, they say, demands more: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. The authors describe SpaceCast-Bench as the first benchmark to directly and diagnostically evaluate this capability.
The benchmark is built around an observe-transform-infer framework. It has 3,862 questions drawn from 182 real-world scenes, spanning 16 task types at three levels: static perception, local prediction, and global prediction. The levels progressively require scene understanding, spatial state updating, and relational inference over unobserved outcomes.
The authors evaluated 21 models and report a stark gap. The strongest model reaches only 58.0%, against 87.2% human performance. Spatially specialized models remain near random chance.
Controlled analyses add two findings. Bridge views are critical for integrating distributed observations. And explicit 3D evidence benefits models more reliably than generated outcome images or videos.
The authors also show that the benchmark's approach can be used for training. Fine-tuning on their programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7%, with macro-average gains reported across six out-of-domain benchmarks.
Key facts
- SpaceCast-Bench has 3,862 questions from 182 real-world scenes, across 16 task types at three levels: static perception, local prediction and global prediction.
- Of 21 evaluated models, the strongest reaches 58.0% against 87.2% human performance; spatially specialized models remain near random chance.
- Bridge views are critical for integrating distributed observations, and explicit 3D evidence helps models more reliably than generated outcome images or videos.
- Fine-tuning Qwen3-VL-4B on the authors' programmatically generated data lifts it from 34.0% to 65.7%, with macro-average gains across six out-of-domain benchmarks.
Why it matters
The authors argue that most spatial reasoning benchmarks check only whether a model can read relations already visible in an image. SpaceCast-Bench instead asks a model to build a scene from observations, work out how an intervention would change it, and infer an outcome it never sees. The authors call it the first benchmark to directly and diagnostically evaluate that capability. The headline result is a wide gap: 58.0% for the strongest of 21 models against 87.2% for humans.
Who it affects
Researchers building and evaluating vision-language models and spatial reasoning systems are the direct audience, since the benchmark targets a skill that existing tests mainly leave out. The result about spatially specialized models staying near random chance is relevant to anyone working on that class of model. The fine-tuning result concerns users of small open models such as Qwen3-VL-4B.
How to use it
The text gives no code or dataset link, so access is not described. What it does point to is a method: the authors generated training data programmatically, and fine-tuning Qwen3-VL-4B on it raised its score from 34.0% to 65.7%. Their controlled analyses also suggest what to supply to a model: bridge views help it integrate distributed observations, and explicit 3D evidence helps more reliably than generated outcome images or videos.
How solid is it
These are the authors' own claims, taken from a paper abstract. The benchmark is sizeable, with 3,862 questions across 182 real-world scenes and 16 task types, and it covers 21 models. The text does not name the models, including which one scored 58.0%, and does not state the metric behind the percentages. It also gives no details of the human evaluation, so the 87.2% human figure cannot be assessed beyond the number itself. The names of the six out-of-domain benchmarks and the size of the macro-average gain are not given either.
Risks and caveats
The claim of being the first benchmark of its kind is the authors' own. The fine-tuning gain from 34.0% to 65.7% is reported on the authors' own benchmark, with data generated by their own method, so it shows what the method does on this test; the out-of-domain gains are reported only as macro-averages over six benchmarks, with no figures given. The statement that specialized models stay near random chance is not backed by a number in the text.