vix.ing · top · new · best · stats

Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation

2025/12/02 by Zhigang Xiao, Yiwei Zhao, Xiao, Zeqi +13
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Context (archaeology) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Generative grammar #Generative model #Modalities #Multimodal Machine Learning Applications #Object (grammar) #Robot Manipulation and Learning #Simple (philosophy) #Spatial contextual awareness #Video tracking

paper · pdf · doi:10.48550/arxiv.2512.03040

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/12/02 · openalex created_date 2025/12/04 · openalex updated_date 2026/08/05

Abstract

We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video diffusion models conditioned solely on video-based scene context can perform complex spatial tasks. We validate on two tasks: scene navigation - following camera-pose instructions while remaining consistent with 3D geometry of the scene, and object grounding - which requires semantic localization, instruction following, and planning. Both tasks use video-only inputs, without auxiliary modalities such as depth or poses. With simple yet effective design choices in the framework and data curation, Video4Spatial demonstrates strong spatial understanding from video context: it plans navigation and grounds target objects end-to-end, follows camera-pose instructions while maintaining spatial consistency, and generalizes to long contexts and out-of-domain environments. Taken together, these results advance video generative models toward general visuospatial reasoning.

Citations

Related