Cambrian-S: Towards Spatial Supersensing in Video
2025/11/06 by Shusheng Yang, Yang, Shusheng, Jihan Yang +23 · 14 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2511.04670
openalex publication_date 2025/11/06 · openalex created_date 2025/11/08 · openalex updated_date 2026/07/28
Abstract
We argue that progress in true multimodal intelligence calls for a shift from reactive, task-driven systems and brute-force long context towards a broader paradigm of supersensing. We frame spatial supersensing as four stages beyond linguistic-only understanding: semantic perception (naming what is seen), streaming event cognition (maintaining memory across continuous experiences), implicit 3D spatial cognition (inferring the world behind pixels), and predictive world modeling (creating internal models that filter and organize information). Current benchmarks largely test only the early stages, offering narrow coverage of spatial cognition and rarely challenging models in ways that require true world modeling. To drive progress in spatial supersensing, we present VSI-SUPER, a two-part benchmark: VSR (long-horizon visual spatial recall) and VSC (continual visual spatial counting). These tasks require arbitrarily long video inputs yet are resistant to brute-force context expansion. We then test data scaling limits by curating VSI-590K and training Cambrian-S, achieving +30% absolute improvement on VSI-Bench without sacrificing general capabilities. Yet performance on VSI-SUPER remains limited, indicating that scale alone is insufficient for spatial supersensing. We propose predictive sensing as a path forward, presenting a proof-of-concept in which a self-supervised next-latent-frame predictor leverages surprise (prediction error) to drive memory and event segmentation. On VSI-SUPER, this approach substantially outperforms leading proprietary baselines, showing that spatial supersensing requires models that not only see but also anticipate, select, and organize experience.
Citations
- Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- NeMo: Needle in a Montage for Video-Language Understanding
- MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
- Scaling RL to Long Videos
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- MindCube: Spatial Mental Modeling from Limited Views
- Whole-Body Conditioned Egocentric Video Prediction
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- SmolVLM: Redefining small and efficient multimodal models
- SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
- TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding
- Scaling Language-Free Visual Representation Learning
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
- Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
- VGGT: Visual Geometry Grounded Transformer
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
- LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant
- EgoLife: Towards Egocentric Life Assistant
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- Intuitive physics understanding emerges from self-supervised pretraining on natural videos
- VideoRoPE: What Makes for Good Video Rotary Position Embedding?
- Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
- Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- Qwen2.5 Technical Report
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Apollo: An Exploration of Video Understanding in Large Multimodal Models
- Navigation World Models
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
- HourVideo: 1-Hour Video-Language Understanding
- TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
- GPT-4o System Card
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
- TIPS: Text-Image Pretraining with Spatial awareness
- Does Spatial Cognition Emerge in Frontier Models?
- AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
- LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- LLaVA-OneVision: Easy Visual Task Transfer
- The Unbearable Slowness of Being: Why do we live at 10 bits/s?
- Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
- SAM 2: Segment Anything in Images and Videos
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- Long Context Transfer from Language to Vision
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- SpatialBot: Precise Spatial Understanding with Vision Language Models
- VideoLLM-online: Online Video Large Language Model for Streaming Video
- GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
- Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
- OpenVLA: An Open-Source Vision-Language-Action Model
- VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
- StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
- Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
- Vript: A Video Is Worth Thousands of Words
- TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- LocCa: Visual Pretraining with Location-aware Captioners
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- VideoMamba: State Space Model for Efficient Video Understanding
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- V-IRL: Grounding Virtual Intelligence in Real Life
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- Simple Hierarchical Planning with Diffusion
- COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Text-Conditioned Resampler For Long Form Video Understanding
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- SILC: Improving Vision Language Pretraining with Self-Distillation
- Learning Interactive Real-World Simulators
- Improved Baselines with Visual Instruction Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Qwen Technical Report
- ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MMBench: Is Your Multi-modal Model an All-around Player?
- Lost in the Middle: How Language Models Use Long Contexts
- Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Perception Test: A Diagnostic Benchmark for Multimodal Video Models
- VideoChat: Chat-Centric Video Understanding
- Visual Instruction Tuning
- Sigmoid Loss for Language Image Pre-Training
- GPT-4 Technical Report
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- LLaMA: Open and Efficient Foundation Language Models
- Connecting Vision and Language with Video Localized Narratives
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- ProcTHOR: Large-Scale Embodied AI Using Procedural Generation
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Predictive Coding: Towards a Future of Deep Learning beyond Backpropagation?
- ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
- Masked Autoencoders Are Scalable Vision Learners
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- GSPMD: General and Scalable Parallelization for ML Computation Graphs
- Learning Transferable Visual Models From Natural Language Supervision
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
- DocVQA: A Dataset for VQA on Document Images
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Conv-Linformer: Boosting Linformer's Performance with Convolution in Small-Scale Settings
- Language Models are Few-Shot Learners
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips
- World Models
- Towards Automatic Learning of Procedures from Web Instructional Videos
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- Language Modeling with Gated Convolutional Networks
- Gaussian Error Linear Units (GELUs)
- A Diagram Is Worth A Dozen Images
- Nonverbal expectancy violations: Model elaboration and application to immediacy behaviors
- Gemini Robotics: Bringing AI into the Physical World
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Cited by
Related