WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
2024/01/25 by Hongliang He, Wenlin Yao, He, Hongliang +13 · 1 voice · 141 citations
Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2401.13919
arxiv published 2024/01/25 · arxiv updated 2024/06/06
Abstract
The rapid advancement of large language models (LLMs) has led to a new era marked by the development of autonomous applications in real-world scenarios, which drives innovation in creating advanced web agents. Existing web agents typically only handle one input modality and are evaluated only in simplified web simulators or static web snapshots, greatly limiting their applicability in real-world scenarios. To bridge this gap, we introduce WebVoyager, an innovative Large Multimodal Model (LMM) powered web agent that can complete user instructions end-to-end by interacting with real-world websites. Moreover, we establish a new benchmark by compiling real-world tasks from 15 popular websites and introduce an automatic evaluation protocol leveraging multimodal understanding abilities of GPT-4V to evaluate open-ended web agents. We show that WebVoyager achieves a 59.1% task success rate on our benchmark, significantly surpassing the performance of both GPT-4 (All Tools) and the WebVoyager (text-only) setups, underscoring the exceptional capability of WebVoyager. The proposed automatic evaluation metric achieves 85.3% agreement with human judgment, indicating its effectiveness in providing reliable and accurate assessments of web agents.
Cited by
- LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
- DECEPTICON: How Dark Patterns Manipulate Web Agents
- ClawRec: A Claw-Native Recommender System
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
- Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- Multi-Agent LLM Committees for Autonomous Software Beta Testing
- XAgen: An Explainability Tool for Identifying and Correcting Failures in Multi-Agent Workflows
- WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
- Privacy Practices of Browser Agents
- Evaluating Long-Context Reasoning in LLM-Based WebAgents
- PPTArena: A Benchmark for PowerPoint Editing
- Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- Fara-7B: An Efficient Agentic Model for Computer Use
- Building Browser Agents: Architecture, Security, and Practical Solutions
- UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
- Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- Test-Time Adaptation for LLM Agents via Environment Interaction
- LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- What's the next frontier for Data-centric AI? Data Savvy Agents
- GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
- Affordance Representation and Recognition for Autonomous Agents
- Code Aesthetics with Agentic Reward Feedback
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
- WebRouter: Query-specific Router via Variational Information Bottleneck for Cost-sensitive Web Agent
- SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
- A Survey on Agentic Multimodal Large Language Models
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- Auto-scaling Continuous Memory for GUI Agent
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
- WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
- Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
- WALT: Web Agents that Learn Tools
- WAREX: Web Agent Reliability Evaluation on Existing Benchmarks
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- RISK: A Framework for GUI Agents in E-commerce Risk Management
- Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
- Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
- Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
- GPA: Learning GUI Process Automation from Demonstrations
- WebSight: A Vision-First Architecture for Robust Web Agents
- Towards Understanding Visual Grounding in Visual Language Models
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- IndusGCC: A Data Benchmark and Evaluation Framework for GUI-Based General Computer Control in Industrial Automation
- A Multimodal GUI Architecture for Interfacing with LLM-Based Conversational Assistants
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- Cybernaut: Towards Reliable Web Automation
- A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
- Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
- Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- WebDS: An End-to-End Benchmark for Web-based Data Science
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
- OpenFPL: An open-source forecasting method rivaling state-of-the-art Fantasy Premier League services
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
- DPMT: Dual Process Multi-scale Theory of Mind Framework for Real-time Human-AI Collaboration
- Aime: Towards Fully-Autonomous Multi-Agent Framework
- VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
- A Systematization of Security Vulnerabilities in Computer Use Agents
- Scaling Context Requires Rethinking Attention
- Hijacking JARVIS: Benchmarking Mobile GUI Agents against Unprivileged Third Parties
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- GenFlow: Interactive Modular System for Image Generation
- Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
- Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation
- Context manipulation attacks : Web agents are susceptible to corrupted memory
- Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
- Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills
- Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
- Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
- macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
- Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
- AI-Native Brand Identity: From Visual Recognition to Cryptographic Verification
- NetArena: Dynamic Benchmarks for AI Agents in Network Automation
- GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
- MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
- Self-Challenging Language Model Agents
- WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- Orca: Browsing at Scale Through User-Driven and AI-Facilitated Orchestration Across Malleable Webpages
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model
- Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution
- AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
- lmgame-Bench: How Good are LLMs at Playing Games?
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- The Hidden Dangers of Browsing AI Agents
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
- Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages
- Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
- Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
- MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
- Benchmark Test-Time Scaling of General LLM Agents
- Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
- WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
- NGENT: Next-Generation AI Agents Must Integrate Multi-Domain Abilities to Achieve Artificial General Intelligence
- From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- AppAgent-Claw: CLI Is All You Need for GUI Automation
- ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
- Facilitating Proactive and Reactive Guidance for Decision Making on the Web: A Design Probe with WebSeek
- Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
- WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- AGI Is Coming... Right After AI Learns to Play Wordle
- Toward Generation of Test Cases from Task Descriptions via History-aware Planning
- GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
- Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
- Inducing Programmatic Skills for Agentic Tasks
Discussions
Related