ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
2025/10/28 by Zhang, Juntian, Jin, Song, Cheng, Chuanqi +8 · 3 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.24285
Abstract
The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging due to the scarcity of high-quality data and the limitations of existing methods: supervised fine-tuning (SFT) often compromises general capabilities, while reinforcement fine-tuning (RFT) prioritizes textual reasoning over visual perception. To bridge this gap, we propose a novel two-stage task that structures visual perception learning as a coarse-to-fine progressive process. Based on this task formulation, we develop ViPER, a self-bootstrapping framework specifically designed to enable iterative evolution through self-critiquing and self-prediction. By synergistically integrating image-level and instance-level reconstruction with a two-stage reinforcement learning strategy, ViPER establishes a closed-loop training paradigm, where internally synthesized data directly fuel the enhancement of perceptual ability. Applied to the Qwen2.5-VL family, ViPER produces the Qwen-Viper series. With an average gain of 1.7% on seven comprehensive benchmarks spanning various tasks and up to 6.0% on fine-grained perception, Qwen-Viper consistently demonstrates superior performance across different vision-language scenarios while maintaining generalizability. Beyond enabling self-improvement in perceptual capabilities, ViPER provides concrete evidence for the reciprocal relationship between generation and understanding, a breakthrough to developing more autonomous and capable VLMs.
Citations
- Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- AME: Aligned Manifold Entropy for Robust Vision-Language Distillation
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- Qwen-Image Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- OmniGen2: Exploration to Advanced Multimodal Generation
- HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- Qwen3 Technical Report
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Qwen2.5-VL Technical Report
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
- Active Data Curation Effectively Distills Large-Scale Multimodal Models
- GPT-4o System Card
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- LLaVA-OneVision: Easy Visual Task Transfer
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
- Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
- Chain of Thoughtlessness? An Analysis of CoT in Planning
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
- The Revolution of Multimodal Large Language Models: A Survey
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- Language Models as Black-Box Optimizers for Vision-Language Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Towards Faithful Model Explanation in NLP: A Survey
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Proximal Policy Optimization Algorithms
Cited by
Related