Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
2025/12/04 by Haobo Yuan, Yuan, Haobo, Sun, Yueyi +16
Computer Science · #Advanced Graph Neural Networks #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2512.05091
openalex publication_date 2025/12/04 · openalex created_date 2025/12/06 · openalex updated_date 2026/07/28
Abstract
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved performance on tasks such as visual grounding and visual question answering. However, the reasoning processes of these models remain largely opaque; they typically output only final predictions without revealing the intermediate steps or fine-grained evidence (e.g., pixels, locations) that lead to the result. This contrasts with human intelligence, which naturally operates through a chain of visual reasoning. To address this limitation, we introduce the Visual Reasoning Tracer (VRT) task, which requires models to not only localize the target object but also explicitly predict the intermediate objects that form the reasoning path. To advance research in this area, we contribute: (1) VRT-Bench, a human-annotated benchmark for evaluating visual reasoning; (2) a new metric for assessing the quality of reasoning traces; and (3) VRT-80k, a large-scale dataset for reasoning model training. Our experiments reveal that while existing models often produce the correct final output, they struggle to ground their intermediate reasoning. In contrast, models trained on VRT-80k achieve substantial improvements in tracing the reasoning path.
Citations
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
- Describe Anything: Detailed Localized Image and Video Captioning
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
- MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning
- MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
- Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
- Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
- MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
- Qwen2.5-VL Technical Report
- Pixel-Level Reasoning Segmentation via Multi-turn Conversations
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
- ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
- OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
- On The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Can GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugs
- VISA: Reasoning Video Object Segmentation via Large Language Models
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
- TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
- Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
- Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Osprey: Pixel Understanding with Visual Instruction Tuning
- GSVA: Generalized Segmentation via Multimodal Large Language Models
- LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
- Aligning and Prompting Everything All at Once for Universal Visual Perception
- Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
- Llemma: An Open Language Model For Mathematics
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- Thinking Like an Expert:Multimodal Hypergraph-of-Thought (HoT) Reasoning to boost Foundation Modals
- LISA: Reasoning Segmentation via Large Language Model
- GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- LeanDojo: Theorem Proving with Retrieval-Augmented Language Models
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- Recognize Anything: A Strong Image Tagging Model
- Let's Verify Step by Step
- Visual Instruction Tuning
- Micrograph segmentations for DDEVD
- Toolformer: Language Models Can Teach Themselves to Use Tools
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- End-to-End Object Detection with Transformers
- Panoptic Segmentation
- Microsoft COCO: Common Objects in Context
- OpenAI o1 System Card
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
Related