Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
2025/10/02 by Li, Xuchen, Li, Xuzhao, Gao, Jiahui +3 · 3 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.01681
Abstract
Vision-Language Models (VLMs) excel at many multimodal tasks, yet they frequently struggle with tasks requiring precise understanding and handling of fine-grained visual elements. This is mainly due to information loss during image encoding or insufficient attention to critical regions. Recent work has shown promise by incorporating pixel-level visual information into the reasoning process, enabling VLMs to access high-resolution visual details during their thought process. However, this pixel-level information is often overused, leading to inefficiency and distraction from irrelevant visual details. To address these challenges, we propose the first framework for adaptive pixel reasoning that dynamically determines necessary pixel-level operations based on the input query. Specifically, we first apply operation-aware supervised fine-tuning to establish baseline competence in textual reasoning and visual operations, then design a novel rollout-guided reinforcement learning framework relying on feedback of the model's own responses, which enables the VLM to determine when pixel operations should be invoked based on query difficulty. Experiments on extensive multimodal reasoning benchmarks show that our model achieves superior performance while significantly reducing unnecessary visual operations. Impressively, our model achieves 73.4% accuracy on HR-Bench 4K while maintaining a tool usage ratio of only 20.1%, improving accuracy and simultaneously reducing tool usage by 66.5% compared to the previous methods.
Citations
- Reinforced Visual Perception with Tools
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Simple o3: Towards Interleaved Vision-Language Reasoning
- Thyme: Think Beyond Images
- CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- PixelThink: Towards Efficient Chain-of-Pixel Reasoning
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- RVTBench: A Benchmark for Visual Reasoning Tasks
- DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- Video-R1: Reinforcing Video Reasoning in MLLMs
- Gemma 3 Technical Report
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching
- Qwen2.5-VL Technical Report
- OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
- Socratic Questioning: Learn to Self-guide Multimodal Reasoning in the Wild
- ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
- How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
- GPT-4o System Card
- FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
- DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
- LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
- Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
- LLaVA-OneVision: Easy Visual Task Transfer
- TokenPacker: Efficient Visual Projector for Multimodal LLM
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
- Instruction-Guided Visual Masking
- ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
- DTLLM-VLT: Diverse Text Generation for Visual Language Tracking Based on LLM
- InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
- MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
- CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
- V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
- Honeybee: Locality-enhanced Projector for Multimodal LLM
- Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
- Dolphins: Multimodal Language Model for Driving
- Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training
- mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
- DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
- Multimodal Chain-of-Thought Reasoning in Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning
- Flamingo: a Visual Language Model for Few-Shot Learning
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- InfographicVQA
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Cited by
Related