Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
2025/05/21 by Haozhe Wang, A.W.Y. Su, Wang, Haozhe +7 · 87 citations
Computer Science · Psychology · #Visual Attention and Saliency Detection #Reinforcement Learning in Robotics #Psychological and Educational Research Studies
paper · pdf · doi:10.48550/arxiv.2505.15966
Abstract
Chain-of-thought reasoning has significantly improved the performance of Large Language Models (LLMs) across various domains. However, this reasoning process has been confined exclusively to textual space, limiting its effectiveness in visually intensive tasks. To address this limitation, we introduce the concept of reasoning in the pixel-space. Within this novel framework, Vision-Language Models (VLMs) are equipped with a suite of visual reasoning operations, such as zoom-in and select-frame. These operations enable VLMs to directly inspect, interrogate, and infer from visual evidences, thereby enhancing reasoning fidelity for visual tasks. Cultivating such pixel-space reasoning capabilities in VLMs presents notable challenges, including the model's initially imbalanced competence and its reluctance to adopt the newly introduced pixel-space operations. We address these challenges through a two-phase training approach. The first phase employs instruction tuning on synthesized reasoning traces to familiarize the model with the novel visual operations. Following this, a reinforcement learning (RL) phase leverages a curiosity-driven reward scheme to balance exploration between pixel-space reasoning and textual reasoning. With these visual operations, VLMs can interact with complex visual inputs, such as information-rich images or videos to proactively gather necessary information. We demonstrate that this approach significantly improves VLM performance across diverse visual reasoning benchmarks. Our 7B model, \model, achieves 84% on V* bench, 74% on TallyQA-Complex, and 84% on InfographicsVQA, marking the highest accuracy achieved by any open-source model to date. These results highlight the importance of pixel-space reasoning and the effectiveness of our framework.
Cited by
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
- LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
- Latent Implicit Visual Reasoning
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- AdaTooler-V: Adaptive Tool-Use for Images and Videos
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- CogDoc: Towards Unified thinking in Documents
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- Rethinking Chain-of-Thought Reasoning for Videos
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
- Training Multi-Image Vision Agents via End2End Reinforcement Learning
- Thinking with Programming Vision: Towards a Unified View for Thinking with Images
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
- Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
- Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
- Thinking in 360°: Humanoid Visual Search in the Wild
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
- Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
- DeepEyesV2: Toward Agentic Multimodal Model
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- VAR: Visual Attention Reasoning via Structured Search and Backtracking
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
- A Survey on Agentic Multimodal Large Language Models
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
- Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
- VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
- Latent Visual Reasoning
- ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- Reverse-Engineered Reasoning for Open-Ended Generation
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
- Simple o3: Towards Interleaved Vision-Language Reasoning
- Reinforcement Learning for Large Model: A Survey
- Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
- PyVision: Agentic Vision with Dynamic Tooling
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- Perception-Aware Policy Optimization for Multimodal Reasoning
Related