GPT-4V(ision) is a Generalist Web Agent, if Grounded
2024/01/03 by Boyuan Zheng, Boyu Gou, Zheng, Boyuan +7 · 1 voice · 186 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.CV #cs.IR
paper · pdf · doi:10.48550/arxiv.2401.01614
openalex publication_date 2024/01/03 · arxiv published 2024/01/03 · arxiv updated 2024/03/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The recent development on large multimodal models (LMMs), especially GPT-4V(ision) and Gemini, has been quickly expanding the capability boundaries of multimodal models beyond traditional tasks like image captioning and visual question answering. In this work, we explore the potential of LMMs like GPT-4V as a generalist web agent that can follow natural language instructions to complete tasks on any given website. We propose SEEACT, a generalist web agent that harnesses the power of LMMs for integrated visual understanding and acting on the web. We evaluate on the recent MIND2WEB benchmark. In addition to standard offline evaluation on cached websites, we enable a new online evaluation setting by developing a tool that allows running web agents on live websites. We show that GPT-4V presents a great potential for web agents -- it can successfully complete 51.1 of the tasks on live websites if we manually ground its textual plans into actions on the websites. This substantially outperforms text-only LLMs like GPT-4 or smaller models (FLAN-T5 and BLIP-2) specifically fine-tuned for web agents. However, grounding still remains a major challenge. Existing LMM grounding strategies like set-of-mark prompting turns out to be not effective for web agents, and the best grounding strategy we develop in this paper leverages both the HTML structure and visuals. Yet, there is still a substantial gap with oracle grounding, leaving ample room for further improvement. All code, data, and evaluation tools are available at https://github.com/OSU-NLP-Group/SeeAct.
Cited by
- Falsifiable Commitment Planning for Self-Correcting Web Agents
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
- AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
- Reinforcement Learning for Self-Improving Agent with Skill Library
- QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
- Using GUI Agent for Electronic Design Automation
- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
- GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection
- MVP: Multiple View Prediction Improves GUI Grounding
- Reliable agent engineering should integrate machine-compatible organizational principles
- AgentBay: A Hybrid Interaction Sandbox for Seamless Human-AI Intervention in Agentic Systems
- Evaluating Long-Context Reasoning in LLM-Based WebAgents
- Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- LegalWebAgent: Empowering Access to Justice via LLM-Based Web Agents
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering
- Computer-Use Agents as Judges for Generative User Interface
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- Learning from Online Videos at Inference Time for Computer-Use Agents
- GUI-360^∘: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- Promoting Sustainable Web Agents: Benchmarking and Estimating Energy Consumption through Empirical and Theoretical Analysis
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
- What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
- Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
- Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
- CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
- Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- PolySkill: Learning Generalizable Skills Through Polymorphic Abstraction
- Composition-Grounded Data Synthesis for Visual Reasoning
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
- Auto-scaling Continuous Memory for GUI Agent
- Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- Agent Learning via Early Experience
- Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- Watch and Learn: Learning to Use Computers from Online Videos
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- WALT: Web Agents that Learn Tools
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Retrieval-augmented GUI Agents with Generative Guidelines
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
- Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents
- TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
- Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- Query-Centric Diffusion Policy for Generalizable Robotic Assembly
- Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- Detecting Pipeline Failures through Fine-Grained Analysis of Web Agents
- TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
- Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder
- Interaction-Driven Browsing: A Human-in-the-Loop Conceptual Framework Informed by Human Web Browsing for Browser-Using Agents
- Towards Understanding Visual Grounding in Visual Language Models
- VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
- MobiAgent: A Systematic Framework for Customizable Mobile Agents
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- Cybernaut: Towards Reliable Web Automation
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
- Agentic Design Review System
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
- MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents
- Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- WebDS: An End-to-End Benchmark for Web-based Data Science
- Adaptive Content Restriction for Large Language Models via Suffix Optimization
- MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
- Turbocharging Web Automation: The Impact of Compressed History States
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
- Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs
- Natural-Language Agent Harnesses
- WebGuard: Building a Generalizable Guardrail for Web Agents
- Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
- MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
- PyVision: Agentic Vision with Dynamic Tooling
- VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
- Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models
- MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
- R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
- Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion
- Hijacking JARVIS: Benchmarking Mobile GUI Agents against Unprivileged Third Parties
- WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
- Context manipulation attacks : Web agents are susceptible to corrupted memory
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- Exploring the Secondary Risks of Large Language Models
- BIMgent: Towards Autonomous Building Modeling via Computer-use Agents
- MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
- Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
- Gen-n-Val: Agentic Image Data Generation and Validation
- macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
- AI-Native Brand Identity: From Visual Recognition to Cryptographic Verification
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
- DPO Learning with LLMs-Judge Signal for Computer Use Agents
- AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
- MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
- Temac: Multi-Agent Collaboration for Automated Web GUI Testing
- VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making
- A Red Teaming Roadmap Towards System-Level Safety
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
- UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
- XBOUND: Exploring Capability Boundaries of Device-Control Agents at the State Level
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- ALRPHFS: Adversarially Learned Risk Patterns with Hierarchical Fast & Slow Reasoning for Robust Agent Defense
- ProgRM: Build Better GUI Agents with Progress Rewards
- LA-RCS: LLM-Agent-Based Robot Control System
- FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
- GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
- Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
- Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
- EVA: Red-Teaming GUI Agents via Evolving Indirect Prompt Injection
- Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
- Building a Stable Planner: An Extended Finite State Machine Based Planning Module for Mobile GUI Agent
- Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
- Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
- MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
- Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
- A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
- Group-in-Group Policy Optimization for LLM Agent Training
- WebInject: Prompt Injection Attack to Web Agents
- RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning
- From Monoliths to Swarms: A Study of Attack Surface Evolution in the Transition to Multi-Agent Web Systems
- DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
- Visual Test-time Scaling for GUI Agent Grounding
- CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare
- Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
- Anticipatory Planning for Multimodal AI Agents
- ScaleTrack: Scaling and back-tracking Automated GUI Agents
- GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
- AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
- AgentOS: From Application Silos to a Natural Language-Driven Data Ecosystem
- APWA: A Distributed Architecture for Parallelizable Agentic Workflows
- A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning
- Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
- AI Planning Framework for LLM-Based Web Agents
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- AndroidGen: Building an Android Language Agent under Data Scarcity
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
- Okara: Detection and Attribution of TLS Man-in-the-Middle Vulnerabilities in Android Apps with Foundation Models
- From Bug Reports to Browser-Executable Procedures: An LLM-Driven Agent for Web GUI Bug Reproduction
- Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
- Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
- On the Robustness of GUI Grounding Models Against Image Attacks
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- Evaluating the Goal-Directedness of Large Language Models
- StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
- Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents
- The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
- Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
Discussions
Related