SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
2024/01/17 by Kanzhi Cheng, Cheng, Kanzhi, Qiushi Sun +11 · 216 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Interactive and Immersive Displays #Virtual Reality Applications and Impacts
paper · pdf · doi:10.48550/arxiv.2401.10935
openalex publication_date 2024/01/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Graphical User Interface (GUI) agents are designed to automate complex tasks on digital devices, such as smartphones and desktops. Most existing GUI agents interact with the environment through extracted structured data, which can be notably lengthy (e.g., HTML) and occasionally inaccessible (e.g., on desktops). To alleviate this issue, we propose a novel visual GUI agent -- SeeClick, which only relies on screenshots for task automation. In our preliminary study, we have discovered a key challenge in developing visual GUI agents: GUI grounding -- the capacity to accurately locate screen elements based on instructions. To tackle this challenge, we propose to enhance SeeClick with GUI grounding pre-training and devise a method to automate the curation of GUI grounding data. Along with the efforts above, we have also created ScreenSpot, the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments. After pre-training, SeeClick demonstrates significant improvement in ScreenSpot over various baselines. Moreover, comprehensive evaluations on three widely used benchmarks consistently support our finding that advancements in GUI grounding directly correlate with enhanced performance in downstream GUI agent tasks. The model, data and code are available at https://github.com/njucckevin/SeeClick.
Cited by
- LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- Scaling GUI Agents with Visual State Transitions
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
- AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
- Xiaomi MiMo-VL-Miloco Technical Report
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- Step-GUI Technical Report
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
- Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
- Using GUI Agent for Electronic Design Automation
- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
- GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection
- MVP: Multiple View Prediction Improves GUI Grounding
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- HiconAgent: History Context-aware Policy Optimization for GUI Agents
- CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- Qwen3-VL Technical Report
- Fara-7B: An Efficient Agentic Model for Computer Use
- D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
- MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements
- OSGym: Scalable OS Infra for Computer Use Agents
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- Grounding Computer Use Agents on Human Demonstrations
- AUTO-Explorer: Automated Data Collection for GUI Agent
- NVIDIA Nemotron Nano V2 VL
- GUI-360^∘: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
- Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
- The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
- Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
- OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
- MGA: Memory-Driven GUI Agent for Observation-Centric Interaction
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
- Mitigating Coordinate Prediction Bias from Positional Encoding Failures
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
- CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
- AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- PolySkill: Learning Generalizable Skills Through Polymorphic Abstraction
- GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
- Composition-Grounded Data Synthesis for Visual Reasoning
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
- A Survey on Agentic Multimodal Large Language Models
- WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
- Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes
- GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- UI-UG: A Unified MLLM for UI Understanding and Generation
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- Retrieval-augmented GUI Agents with Generative Guidelines
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- GPA: Learning GUI Process Automation from Demonstrations
- Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
- Mano Technical Report
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
- GUI-ReWalk: Massive Data Generation for GUI Agent via Stochastic Exploration and Intent-Aware Reasoning
- DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
- PrivWeb: Unobtrusive and Content-aware Privacy Protection For Web Agents
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- Towards Understanding Visual Grounding in Visual Language Models
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents
- Instruction Agent: Enhancing Agent with Expert Demonstration
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- MobiAgent: A Systematic Framework for Customizable Mobile Agents
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
- You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
- CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- OpenCUA: Open Foundations for Computer-Use Agents
- FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
- Reinforcement Learning for Large Model: A Survey
- Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- CoAct-1: Computer-using Multi-Agent System with Coding Actions
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- NatureGAIA: Pushing the Frontiers of GUI Agents with a Challenging Benchmark and High-Quality Trajectory Dataset
- Phi-Ground Tech Report: Advancing Perception in GUI Grounding
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
- Turbocharging Web Automation: The Impact of Compressed History States
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
- GUI-G2: Gaussian Reward Modeling for GUI Grounding
- MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
- WebGuard: Building a Generalizable Guardrail for Web Agents
- UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
- A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
- LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
- Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
- MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
- VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
- R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
- GTA1: GUI Test-time Scaling Agent
- Hijacking JARVIS: Benchmarking Mobile GUI Agents against Unprivileged Third Parties
- Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
- What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
- ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding
- Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
- GenFlow: Interactive Modular System for Image Generation
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation
- Understanding GUI Agent Localization Biases through Logit Sharpness
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
- Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
- Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
- macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
- MiMo-VL Technical Report
- Surfer-H Meets Holo1: Cost-Efficient Web Agent Powered by Open Weights
- GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
- FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
- AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- Grounded Reinforcement Learning for Visual Reasoning
- Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
- UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
- XBOUND: Exploring Capability Boundaries of Device-Control Agents at the State Level
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Robot Operation of Home Appliances by Reading User Manuals
- Efficient Multi-modal Long Context Learning for Training-free Adaptation
- TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
- FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
- Qwen-CUA: Native Computer Use for (almost) Everything
- Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
- ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
- Building a Stable Planner: An Extended Finite State Machine Based Planning Module for Mobile GUI Agent
- Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
- Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
- MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
- Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
- GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
- Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
- InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction
- RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning
- LLM Agents Can See Code Repositories
- Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
- EcoAgent: An Efficient Device-Cloud Collaborative Multi-Agent Framework for Mobile Automation
- Read More, Think More: Revisiting Observation Reduction for Web Agents
- Visual Test-time Scaling for GUI Agent Grounding
- ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
- Anticipatory Planning for Multimodal AI Agents
- ScaleTrack: Scaling and back-tracking Automated GUI Agents
- AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
- Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
- Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- AndroidGen: Building an Android Language Agent under Data Scarcity
- AppAgent-Claw: CLI Is All You Need for GUI Automation
- OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
- ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
- Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology
- GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
- Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
- Cracking the Code of Action: a Generative Approach to Affordances for Reinforcement Learning
- On the Robustness of GUI Grounding Models Against Image Attacks
- Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
- UFO2: The Desktop AgentOS
- Toward Generation of Test Cases from Task Descriptions via History-aware Planning
- InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
- LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
- StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
- GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
- Kimi-VL Technical Report
- SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
Related