LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
2025/11/25 by Man, Yunze, Wang, Shihao, Zhang, Guowen +7
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2511.20648
Abstract
To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models (VLMs) excel at open-ended 2D description and grounding, yet multi-object 3D detection remains largely missing from the VLM toolbox. We present LocateAnything3D, a VLM-native recipe that casts 3D detection as a next-token prediction problem. The key is a short, explicit Chain-of-Sight (CoS) sequence that mirrors how human reason from images: find an object in 2D, then infer its distance, size, and pose. The decoder first emits 2D detections as a visual chain-of-thought, then predicts 3D boxes under an easy-to-hard curriculum: across objects, a near-to-far order reduces early ambiguity and matches ego-centric utility; within each object, a center-from-camera, dimensions, and rotation factorization ranks information by stability and learnability. This VLM-native interface preserves open-vocabulary and visual-prompting capability without specialized heads. On the challenging Omni3D benchmark, our model achieves state-of-the-art results, with 49.89 AP3D, surpassing the previous best by +15.51 absolute improvement even when the baseline is given ground-truth 2D boxes. It also generalizes zero-shot to held-out categories with strong robustness. By turning 3D detection into a disciplined next-token problem, LocateAnything3D offers a practical foundation for models to perceive in 3D.
Citations
- ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- MolmoAct: Action Reasoning Models that can Reason in Space
- Where, What, Why: Towards Explainable Driver Attention Prediction
- Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
- From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
- Detect Anything 3D in the Wild
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
- RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
- Qwen2.5-VL Technical Report
- UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
- Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
- SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D Annotations
- EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
- Open Vocabulary Monocular 3D Object Detection
- RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
- LLaVA-Critic: Learning to Evaluate Multimodal Models
- Emu3: Next-Token Prediction is All You Need
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
- YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
- Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
- SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models
- Building and better understanding vision-language models: insights and future directions
- Global-Local Collaborative Inference with LLM for Lidar-Based Open-Vocabulary Detection
- Robotic Control via Embodied Chain-of-Thought Reasoning
- ColPali: Efficient Document Retrieval with Vision Language Models
- Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
- SpatialBot: Precise Spatial Understanding with Vision Language Models
- MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
- WeatherQA: Can Multimodal Language Models Reason about Severe Weather?
- WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
- RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
- OpenVLA: An Open-Source Vision-Language-Action Model
- Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
- TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
- MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Language-Image Models with 3D Understanding
- What matters when building vision-language models?
- Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Advancing LLM Reasoning Generalists with Preference Trees
- TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios
- OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
- Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
- Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
- OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
- ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
- Orca-Math: Unlocking the potential of SLMs in Grade School Math
- Visually Dehallucinative Instruction Generation: Know What You Don't Know
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- FM-OV3D: Foundation Model-based Cross-modal Knowledge Blending for Open-Vocabulary 3D Detection
- DriveLM: Driving with Graph Visual Question Answering
- Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
- OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection
- Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning
- How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs
- MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
- To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
- Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
- LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
- CogVLM: Visual Expert for Pretrained Language Models
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- Uni3DETR: Unified 3D Detection Transformer
- UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
- CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
- Object2Scene: Putting Objects in Context for Open-Vocabulary 3D Detection
- EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding
- Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents
- Embodied Task Planning with Large Language Models
- LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
- VisText: A Benchmark for Semantically Rich Chart Captioning
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations
- ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
- TheoremQA: A Theorem-driven Question Answering dataset
- Generalized Planning in PDDL Domains with Pretrained Large Language Models
- ICDAR 2023 Competition on Hierarchical Text Detection and Recognition
- CHIC: Corporate Document for Visual question Answering
- Visual Instruction Tuning
- PDFVQA: A New Dataset for Real-World VQA on PDF Documents
- Open-Vocabulary Point-Cloud Object Detection without 3D Annotation
- Sigmoid Loss for Language Image Pre-Training
- GPT-4 Technical Report
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models
- UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
- Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning
- Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion
- MapQA: A Dataset for Question Answering on Choropleth Maps
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
- ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
- Code as Policies: Language Model Programs for Embodied Control
- CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
- Toward Understanding WordArt: Corner-Guided Transformer for Scene Text Recognition
- Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
- Open-Vocabulary 3D Detection via Image-level Class and Debiased Cross-modal Contrastive Learning
- MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data
- A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- MonoDETR: Depth-guided Transformer for Monocular 3D Object Detection
- MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Chart-to-Text: A Large-Scale Benchmark for Chart Summarization
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Detecting Twenty-thousand Classes using Image-level Supervision
- PointCLIP: Point Cloud Understanding by CLIP
- OCR-free Document Understanding Transformer
- ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
- IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
- Probabilistic and Geometric Depth: Detecting Objects in Perspective
- TextStyleBrush: Transfer of Text Aesthetics from a Single Example
- TextStyleBrush: Transfer of Text Aesthetics From a Single Example
- ImVoxelNet: Image to Voxels Projection for Monocular and Multi-View General-Purpose 3D Object Detection
- Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
- InfographicVQA
- FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection
- MonoRUn: Monocular 3D Object Detection by Reconstruction and Uncertainty Propagation
- IAFA: Instance-aware Feature Aggregation for 3D Object Detection from a Single Image
- Learning Transferable Visual Models From Natural Language Supervision
- Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild\n with Pose Annotations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
- Google Landmarks Dataset v2 -- A Large-Scale Benchmark for Instance-Level Recognition and Retrieval
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question Answering
- SMOKE: Single-Stage Monocular 3D Object Detection via Keypoint Estimation
- ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language
- Connecting Vision and Language with Localized Narratives
- PlotQA: Reasoning over Scientific Plots
- ICDAR 2019 Competition on Large-scale Street View Text with Partial Labeling -- RRC-LSVT
- ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text (RRC-ArT)
- ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard
- SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition
- LVIS: A Dataset for Large Vocabulary Instance Segmentation
- Scene Text Visual Question Answering
- FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents
- Towards VQA Models That Can Read
- Objects as Points
- nuScenes: A multimodal dataset for autonomous driving
- RAVEN: A Dataset for Relational and Analogical Visual rEasoNing
- Learning to Describe Differences Between Pairs of Similar Images
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- VizWiz Grand Challenge: Answering Visual Questions from Blind People
- DVQA: Understanding Data Visualizations via Question Answering
- Simple and Effective Multi-Paragraph Reading Comprehension
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- Modeling Context in Referring Expressions
- A Diagram Is Worth A Dozen Images
- COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images
- Generation and Comprehension of Unambiguous Object Descriptions
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for\n Richer Image-to-Sentence Models
- Exploring Models and Data for Image Question Answering
- ImageNet Large Scale Visual Recognition Challenge
- ImageNet Large Scale Visual Recognition Challenge
- Microsoft COCO: Common Objects in Context
- Vision meets robotics: The KITTI dataset
- Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
- MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
- xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
- 3D-LLM: Injecting the 3D World into Large Language Models
- EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
- RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks
- WizardLM: Empowering large pre-trained language models to follow complex instructions
Related