SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
Abstract
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest model reaches only 58.0% against 87.2% human performance, while spatially specialized models remain near random chance. Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos. Fine-tuning on our programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7% with macro-average gains across six out-of-domain benchmarks.
Community
This paper introduces SpaceCast-Bench, a benchmark for predictive spatial reasoning that evaluates whether vision-language models can infer unseen outcomes of spatial transformations, with 3,862 questions across 16 task types covering static perception, local prediction, and global prediction.
š» Code: https://github.com/ZJU-REAL/SpaceCast-Bench
š¤ Dataset: https://huggingface.co/datasets/hongxingli/SpaceCast-Bench
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models (2026)
- Uruqi: Learning Spatial Cognition from Visual Experience (2026)
- From Reasoning Failures to Composable Video Spatial Intelligence (2026)
- Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models (2026)
- Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes (2026)
- MVVBench: Benchmarking 4D Reasoning in Vision-Language Models (2026)
- Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper