GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
2025/04/14 by Luo, Run, Lu Wang, Wang, Lu +8 · 233 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.2504.10458
openalex publication_date 2025/04/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Existing efforts in building Graphical User Interface (GUI) agents largely rely on the training paradigm of supervised fine-tuning on Large Vision-Language Models (LVLMs). However, this approach not only demands extensive amounts of training data but also struggles to effectively understand GUI screenshots and generalize to unseen interfaces. The issue significantly limits its application in real-world scenarios, especially for high-level tasks. Inspired by Reinforcement Fine-Tuning (RFT) in large reasoning models (e.g., DeepSeek-R1), which efficiently enhances the problem-solving capabilities of large language models in real-world settings, we propose \name, the first reinforcement learning framework designed to enhance the GUI capabilities of LVLMs in high-level real-world task scenarios, through unified action space rule modeling. By leveraging a small amount of carefully curated high-quality data across multiple platforms (including Windows, Linux, MacOS, Android, and Web) and employing policy optimization algorithms such as Group Relative Policy Optimization (GRPO) to update the model, \name achieves superior performance using only 0.02% of the data (3K vs. 13M) compared to previous state-of-the-art methods like OS-Atlas across eight benchmarks spanning three different platforms (mobile, desktop, and web). These results demonstrate the immense potential of reinforcement learning based on unified action space rule modeling in improving the execution capabilities of LVLMs for real-world GUI agent tasks.
Citations
Cited by
- LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
- Scaling GUI Agents with Visual State Transitions
- RollArt: Disaggregated Multi-Task Agentic RL Training at Scale
- ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
- MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- Using GUI Agent for Electronic Design Automation
- MVP: Multiple View Prediction Improves GUI Grounding
- From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- HiconAgent: History Context-aware Policy Optimization for GUI Agents
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
- Calibrated Multimodal Representation Learning with Missing Modalities
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- Grounding Computer Use Agents on Human Demonstrations
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
- GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
- ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
- AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
- A Survey on Agentic Multimodal Large Language Models
- WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
- Auto-scaling Continuous Memory for GUI Agent
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
- UI-UG: A Unified MLLM for UI Understanding and Generation
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
- RISK: A Framework for GUI Agents in E-commerce Risk Management
- ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
- Mano Technical Report
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- InfraMind: A Novel Exploration-based GUI Agentic Framework for Mission-critical Industrial Management
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration
- Towards Secure and Explainable Smart Contract Generation with Security-Aware Group Relative Policy Optimization
- Towards Understanding Visual Grounding in Visual Language Models
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- TuneAgent: Agentic Operating System Kernel Tuning with Reinforcement Learning
- CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
- Reinforcement Learning for Large Model: A Survey
- Memp: Exploring Agent Procedural Memory
- Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
- Phi-Ground Tech Report: Advancing Perception in GUI Grounding
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
- OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
- GUI-G2: Gaussian Reward Modeling for GUI Grounding
- MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
- UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
- MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
- GTA1: GUI Test-time Scaling Agent
- ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
- Where, What, Why: Towards Explainable Driver Attention Prediction
- Mobile-R1: Towards Interactive Reinforcement Learning for VLM-Based Mobile Agent via Task-Level Rewards
- HiMA-Ecom: Enabling Joint Training of Hierarchical Multi-Agent E-commerce Assistants
- Beyond Syntax: Action Semantics Learning for App Agents
- JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
- GA-S3: Comprehensive Social Network Simulation with Group Agents
- AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
- Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale
- TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence
- Table-R1: Inference-Time Scaling for Table Reasoning
- InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization
- Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
- ProgRM: Build Better GUI Agents with Progress Rewards
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- Qwen-CUA: Native Computer Use for (almost) Everything
- Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
- ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
- NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
- ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
- An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
- Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
- GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
- Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
- GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
- RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning
- ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
- Anticipatory Planning for Multimodal AI Agents
- Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
- A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning
- MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research
- VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
- Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology
- Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
- Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
- BashCoder-R1: Towards Robust and Explainable Bash Code Generation with Robustness-Aware Group Relative Policy Optimization
- InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
Related