LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
2025/11/24 by Wang, Shuai, Zhang, Daoan, Bai, Tianyi +3 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2511.19261
Abstract
Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-of-the-art VLMs still struggle to understand 3D space and long videos, although they are powerful in typical vision-language tasks. Current methods often rely on specialized architectural designs to improve performance for 3D tasks and video understanding tasks separately. In contrast, we propose LAST, short for LeArn to Think in Space and Time, to jointly improve 3D spatial and long video understanding for general VLMs with only a set of 2D images as inputs. LAST makes VLMs think in space and time rather than only with text before giving the final answer, building visual thinking trajectories in 3D space and temporal dimension. We demonstrate the effectiveness of LAST in two scenarios: 1) zero-shot, where we directly prompt proprietary models; and 2) fine-tuning general VLMs with data that include thinking trajectories in 3D space and time. We show that LAST brings substantial gains in various benchmarks, including 3 spatial understanding, 4 video understanding, and 3 image understanding tasks. Notably, 15.8% gains on EgoSchema with GPT-4o in a zero-shot manner and 8.3 gains on VSI-Bench compared with Qwen2.5-VL-7B.
Citations
- 3D Aware Region Prompted Vision Language Model
- CAViAR: Critic-Augmented Video Agentic Reasoning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
- Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
- Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
- Video-R1: Reinforcing Video Reasoning in MLLMs
- VGGT: Visual Geometry Grounded Transformer
- Adaptive Keyframe Sampling for Long Video Understanding
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
- 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Apollo: An Exploration of Video Understanding in Large Multimodal Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
- RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
- GPT-4o System Card
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
- Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Training-free Video Temporal Grounding using Large-scale Pre-trained Models
- Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
- LLaVA-OneVision: Easy Visual Task Transfer
- Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
- SAM 2: Segment Anything in Images and Videos
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- SpatialBot: Precise Spatial Understanding with Vision Language Models
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
- LVBench: An Extreme Long Video Understanding Benchmark
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Understanding Long Videos with Multimodal Language Models
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- Language Repository for Long Video Understanding
- VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
- VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- Memory Consolidation Enables Long-Context Video Understanding
- CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
- A Simple LLM Framework for Long-Range Video Question-Answering
- V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
- Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
- A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
- VILA: On Pre-training for Visual Language Models
- LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
- Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
- LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Lost in the Middle: How Language Models Use Long Contexts
- Self-Chained Image-Language Model for Video Localization and Question Answering
- Visual Instruction Tuning
- Verbs in Action: Improving verb understanding in video-language models
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- ViperGPT: Visual Inference via Python Execution for Reasoning
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- PaLM-E: An Embodied Multimodal Language Model
- LLaMA: Open and Efficient Foundation Language Models
- Multimodal Chain-of-Thought Reasoning in Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Visual Programming: Compositional visual reasoning without training
- SQA3D: Situated Question Answering in 3D Scenes
- Flamingo: a Visual Language Model for Few-Shot Learning
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- ScanQA: 3D Question Answering for Spatial Scene Understanding
- NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
- Learning Transferable Visual Models From Natural Language Supervision
- Language Models are Few-Shot Learners
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- CIDEr: Consensus-based Image Description Evaluation
- 3D-LLM: Injecting the 3D World into Large Language Models
Cited by
Related