Autonomous Tester Agent Benchmark
2023/07/25 by Zhou, Shuyan, Frank F. Xu, Hao Zhu +19 · 256 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2307.13854
Abstract
Openstreetmap docker files required to self-host the WebArena benchmark, as described here:https://webarena.dev/https://arxiv.org/abs/2307.13854https://github.com/web-arena-x/webarena/tree/main/environmentdocker Copyright to openstreetmaphttps://www.openstreetmap.org/copyright
Cited by
- Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
- It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
- Breaking the illusion: Automated Reasoning of GDPR Consent Violations
- Unbiased Visual Reasoning with Controlled Visual Inputs
- E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
- ClawRec: A Claw-Native Recommender System
- Falsifiable Commitment Planning for Self-Correcting Web Agents
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
- Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
- AgentWatcher: A Rule-based Prompt Injection Monitor
- PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
- TrafficSimAgent: A Hierarchical Agent Framework for Autonomous Traffic Simulation with MCP Control
- Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
- Reinforcement Learning for Self-Improving Agent with Skill Library
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
- PAACE: A Plan-Aware Automated Agent Context Engineering Framework
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment
- ceLLMate: Sandboxing Browser AI Agents
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- Towards a Science of Scaling Agent Systems
- AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
- Trusted AI Agents in the Cloud
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- Evolutionary System 2 Reasoning: An Empirical Proof
- How Well Does Agent Development Reflect Real-World Work?
- Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
- PPTArena: A Benchmark for PowerPoint Editing
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- Real-Time Procedural Learning From Experience for AI Agents
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
- MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)
- Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
- Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
- "Are We Done Yet?": A Vision-Based Judge for Autonomous Task Completion of Computer Use Agents
- CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents
- Fara-7B: An Efficient Agentic Model for Computer Use
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
- ASTRA: Agentic Steerability and Risk Assessment Framework
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
- Why Do Language Model Agents Whistleblow?
- SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
- An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
- Generative Caching for Structurally Similar Prompts and Responses
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- AI Annotation Orchestration: Evaluating LLM verifiers to Improve the Quality of LLM Annotations in Learning Analytics
- OSGym: Super-Scalable Distributed Data Engine for Generalizable Computer Agents
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- Adapting Web Agents with Synthetic Supervision
- AUTO-Explorer: Automated Data Collection for GUI Agent
- Dataforge: Agentic Platform for Autonomous Data Engineering
- Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
- Learning from Online Videos at Inference Time for Computer-Use Agents
- AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent
- Real-Time Reasoning Agents in Evolving Environments
- GUI-360^∘: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- Test-Time Adaptation for LLM Agents via Environment Interaction
- Scaling Agent Learning via Experience Synthesis
- LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
- Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories
- OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
- Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
- Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
- Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasks
- SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
- ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
- What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
- Agents' Last Exam
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- What Limits Agentic Systems Efficiency?
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- WebATLAS: An LLM Agent with Experience-Driven Memory and Action Simulation
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
- Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- Investigating the Impact of Dark Patterns on LLM-Based Web Agents
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- PolySkill: Learning Generalizable Skills Through Polymorphic Abstraction
- WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale
- LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
- ReUseIt: Synthesizing Reusable AI Agent Workflows for Web Automation
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
- TaskAudit: Detecting Functiona11ity Errors in Mobile Apps via Agentic Task Execution
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
- AgentCaster: Reasoning-Guided Tornado Forecasting
- ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
- WebRouter: Query-specific Router via Variational Information Bottleneck for Cost-sensitive Web Agent
- SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
- Scaling Long-Horizon LLM Agent via Context-Folding
- R-WoM: Retrieval-augmented World Model For Computer-use Agents
- Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- ALLOY: Generating Reusable Agent Workflows from User Demonstration
- Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
- Polar: Agentic RL on Any Harness at Scale
- Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
- Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
- Synthetic Sandbox for Training Machine Learning Engineering Agents
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
- How can we assess human-agent interactions? Case studies in software agent design
- Auto-scaling Continuous Memory for GUI Agent
- Fundamentals of Building Autonomous LLM Agents
- Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- Agent Learning via Early Experience
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
- Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
- MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
- A Survey on Agentic Security: Applications, Threats and Defenses
- Toward Systems Foundations for Agentic Exploration
- When Should Users Check? A Decision-Theoretic Model of Confirmation Frequency in Multi-Step AI Agent Tasks
- LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
- JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
- MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
- Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
- VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
- FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents
- Learning Efficient Guardrails for Compliance
- Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
- WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- SCUBA: Salesforce Computer Use Benchmark
- When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets
- A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
- Where LLM Agents Fail and How They can Learn From Failures
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- Scaling Synthetic Task Generation for Agents via Exploration
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- Agentic Services Computing
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- Retrieval-augmented GUI Agents with Generative Guidelines
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
- PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- Estimating the Empowerment of Language Model Agents
- PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning
- CORE: Full-Path Evaluation of LLM Agents Beyond Final State
- Robust, Observable, and Evolvable Agentic Systems Engineering: A Principled Framework Validated via the Fairy GUI Agent
- An LLM-based Agentic Framework for Accessible Network Control
- Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- DeepResearch Agent System
- How Benchmarks Mis-Score Computer-Use Agents
- ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
- AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration
- Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
- What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents
- GPA: Learning GUI Process Automation from Demonstrations
- Agentic AutoSurvey: Let LLMs Survey LLMs
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- An Evaluation-Centric Paradigm for Scientific Visualization Agents
- GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing
- TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
- DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Redefining Website Fingerprinting Attacks With Multiagent LLMs
- Interaction-Driven Browsing: A Human-in-the-Loop Conceptual Framework Informed by Human Web Browsing for Browser-Using Agents
- PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- WebSight: A Vision-First Architecture for Robust Web Agents
- Towards Understanding Visual Grounding in Visual Language Models
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- LLM-Based Agents for Competitive Landscape Mapping in Drug Asset Due Diligence
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
- Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- Throttling Web Agents Using Reasoning Gates
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- Morae: Proactively Pausing UI Agents for User Choices
- AWorld: Orchestrating the Training Recipe for Agentic AI
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- Analyzing Information Sharing and Coordination in Multi-Agent Planning
- Deep Research: A Survey of Autonomous Research Agents
- Agentic Design Review System
- Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
- OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- OpenCUA: Open Foundations for Computer-Use Agents
- FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
- Reinforcement Learning for Large Model: A Survey
- Cognitive Duality for Adaptive Web Agents
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
- HarmonyGuard: Toward Safety and Utility in Web Agents via Adaptive Policy Enhancement and Dual-Objective Optimization
- ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
- VeriGUI: Verifiable Long-Chain GUI Dataset
- ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- Meta-RAG on Large Codebases Using Code Summarization
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- BlockA2A: Towards Secure and Verifiable Agent-to-Agent Interoperability
- WebDS: An End-to-End Benchmark for Web-based Data Science
- SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents
- MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
- A Multi-Agent Generative AI Framework for IC Module-Level Verification Automation
- Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
- Evaluation and Benchmarking of LLM Agents: A Survey
- MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
- Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
Related