Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
Abstract
FactoSR improves vision-language model spatial reasoning by decomposing 3D and temporal recovery into factorized geometric sub-objectives optimized via reinforcement learning.
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reasoning demands the recovery of latent 3D geometry and temporal continuity. To conquer this high-dimensional complexity, we advocate a shift from monolithic learning to a ``divide and conquer'' paradigm. We present FactoSR, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection. At its core, FactoSR decomposes the monolithic problem of world-consistent reasoning into three orthogonal, geometric sub-objectives: planar correspondence (XY), depth consistency (Z), and temporal reversibility (T). By optimizing these verifiable constraints within a unified policy learning mechanism, we effectively transform an ill-posed projection recovery problem into a series of tangible reasoning steps. Extensive evaluations on multi-view and video benchmarks demonstrate that this elegant decomposition yields substantial gains in 3D and 4D reasoning, achieving a 5.9% boost on VSI-Bench and 4.5% on All-Angles-Bench. Our findings suggest that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.
Community
๐ Excited to share FactoSR: Unfold The World โ Factorize 4D Properties in Reinforcing Spatial Reasoning!
We factorize 4D spatial reasoning into XY (plane), Z (depth), and T (time) with verifiable rewards, pushing VLMs beyond flat 2D understanding toward more world-aware spatial reasoning. FactoSR brings +5.9% on VSI-Bench and +4.5% on All-Angles-Bench.
๐ Paper: https://arxiv.org/abs/2609.03729
๐ป Code: https://github.com/ZimaBlue-WAM/FactoSR
Comments and feedback are very welcome! ๐
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward (2026)
- ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning (2026)
- Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models (2026)
- GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding (2026)
- Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models (2026)
- SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models (2026)
- Hy-Embodied-VLM-1.0: Efficient Physical-World Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.03729 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper