Researchers present SpatialBlock, an approach for improving how large vision-language models understand 3D spatial relationships from 2D images. The method introduces a synthetic dataset of 15,000 block-stacking problems designed to teach foundational spatial skills through structured manipulation tasks. Models trained on this synthetic data significantly outperform baselines and generalize to real-world spatial reasoning tasks, demonstrating that compact, purpose-built synthetic training data can meaningfully enhance spatial reasoning capabilities in vision-language models.
