COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
2025/12/04 by Zhang, Zefeng, Hao, Xiangzhao, Tang, Hengzhu +8
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2512.04563
Abstract
Visual Spatial Reasoning is crucial for enabling Multimodal Large Language Models (MLLMs) to understand object properties and spatial relationships, yet current models still struggle with 3D-aware reasoning. Existing approaches typically enhance either perception, by augmenting RGB inputs with auxiliary modalities such as depth and segmentation, or reasoning, by training on spatial VQA datasets and applying reinforcement learning, and thus treat these two aspects in isolation. In this work, we investigate whether a unified MLLM can develop an intrinsic ability to enhance spatial perception and, through adaptive interleaved reasoning, achieve stronger spatial intelligence. We propose COOPER, a unified MLLM that leverages depth and segmentation as auxiliary modalities and is trained in two stages to acquire auxiliary modality generation and adaptive, interleaved reasoning capabilities. COOPER achieves an average 6.91% improvement in spatial reasoning while maintaining general performance. Moreover, even a variant trained only for auxiliary modality generation attains a 7.92% gain on distance and size estimation, suggesting that learning to generate auxiliary modalities helps internalize spatial knowledge and strengthen spatial understanding.
Citations
- MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- RewardDance: Reward Scaling in Visual Generation
- Reinforced Visual Perception with Tools
- Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Thyme: Think Beyond Images
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Qwen-Image Technical Report
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
- SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- Emerging Properties in Unified Multimodal Pretraining
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- Seed1.5-VL Technical Report
- Flow-GRPO: Training Flow Matching Models via Online RL
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
- Improved Visual-Spatial Reasoning via R1-Zero-Like Training
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
- Mind with Eyes: from Language Reasoning to Multimodal Reasoning
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
- Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
- Liquid: Language Models are Scalable and Unified Multi-modal Generators
- One Diffusion to Generate Them All
- HourVideo: 1-Hour Video-Language Understanding
- GPT-4o System Card
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Emu3: Next-Token Prediction is All You Need
- OmniGen: Unified Image Generation
- Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
- SpatialBot: Precise Spatial Understanding with Vision Language Models
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- MMBench: Is Your Multi-modal Model an All-around Player?
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
- DDP: Diffusion Model for Dense Visual Prediction
- Unleashing Text-to-Image Diffusion Models for Visual Perception
- Scalable Diffusion Models with Transformers
- RT-1: Robotics Transformer for Real-World Control at Scale
- Flow Matching for Generative Modeling
- Building Normalizing Flows with Stochastic Interpolants
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- Vision Transformers for Dense Prediction
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
- Neural Discrete Representation Learning
- Semantic Understanding of Scenes through the ADE20K Dataset
- LATTE: Learning to Think with Vision Specialists
Related