vix.ing · top · new · best · stats · spec

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

2025/06/10 by Sunil Kumar, Kumar, Sunil, Bowen Zhao +5 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Constraint Satisfaction and Optimization #Data Visualization and Analytics #FOS: Computer and information sciences #Machine Learning (cs.LG)

paper · pdf · doi:10.48550/arxiv.2506.14821

openalex publication_date 2025/06/10 · openalex created_date 2025/10/19 · openalex updated_date 2026/07/28

Abstract

Despite tremendous recent advances in large model reasoning ability, vision-language models (VLMs) still struggle with detailed visual reasoning, especially when compute resources are limited. To address this challenge, we draw inspiration from methods like Deepseek-r1 for VLMs and train smaller-scale models with Group Relative Policy Optimization (GRPO) to use external tools such as zoom. The greatest benefit is obtained with a combination of GRPO learning, a simple reward structure, a simplified tool-calling interface, allocating additional tokens to the result of the tool call, and a training data mix that over-represents visually difficult examples. Compared to similarly-sized baseline models, our method achieves better performance on some visual question-answering (VQA) tasks, thanks to the detailed visual information gathered from the external tool.

Citations

Cited by

Related