OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
2024/04/11 by Xie, Tianbao, Zhang, Danyang, Chen, Jixuan +14 · 176 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2404.07972
Abstract
Autonomous agents that accomplish complex computer tasks with minimal human interventions have the potential to transform human-computer interaction, significantly enhancing accessibility and productivity. However, existing benchmarks either lack an interactive environment or are limited to environments specific to certain applications or domains, failing to reflect the diverse and complex nature of real-world computer use, thereby limiting the scope of tasks and agent scalability. To address this issue, we introduce OSWorld, the first-of-its-kind scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems such as Ubuntu, Windows, and macOS. OSWorld can serve as a unified, integrated computer environment for assessing open-ended computer tasks that involve arbitrary applications. Building upon OSWorld, we create a benchmark of 369 computer tasks involving real web and desktop apps in open domains, OS file I/O, and workflows spanning multiple applications. Each task example is derived from real-world computer use cases and includes a detailed initial state setup configuration and a custom execution-based evaluation script for reliable, reproducible evaluation. Extensive evaluation of state-of-the-art LLM/VLM-based agents on OSWorld reveals significant deficiencies in their ability to serve as computer assistants. While humans can accomplish over 72.36% of the tasks, the best model achieves only 12.24% success, primarily struggling with GUI grounding and operational knowledge. Comprehensive analysis using OSWorld provides valuable insights for developing multimodal generalist agents that were not possible with previous benchmarks. Our code, environment, baseline models, and data are publicly available at https://os-world.github.io.
Cited by
- Scaling GUI Agents with Visual State Transitions
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
- Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation
- ACM: Agentic Context Management for Long Horizon Tasks
- ClawRec: A Claw-Native Recommender System
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
- Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
- When Thinking Before Retrieval Hurts: TraceBound Diagnostics for Adaptive Knowledge-Graph Retrieval
- CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks
- Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- Step-GUI Technical Report
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- Multi-Agent Collaborative Framework for Intelligent IT Operations: An AOI System with Context-Aware Compression and Dynamic Task Scheduling
- From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- Using GUI Agent for Electronic Design Automation
- Benchmarking the Generality of Vision-Language-Action Models
- SoMe: A Realistic Benchmark for LLM-based Social Media Agents
- How Well Does Agent Development Reflect Real-World Work?
- Tipping the Dominos: Topology-Aware Multi-Hop Attacks on LLM-Based Multi-Agent Systems
- DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
- PPTArena: A Benchmark for PowerPoint Editing
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
- MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents
- A Rosetta Stone for AI Benchmarks
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- Qwen3-VL Technical Report
- OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
- Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge
- "Are We Done Yet?": A Vision-Based Judge for Autonomous Task Completion of Computer Use Agents
- AppSelectBench: Application-Level Tool Selection Benchmark
- Fara-7B: An Efficient Agentic Model for Computer Use
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
- WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering
- Building Browser Agents: Architecture, Security, and Practical Solutions
- SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
- ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
- Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- OSGym: Super-Scalable Distributed Data Engine for Generalizable Computer Agents
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- Adapting Web Agents with Synthetic Supervision
- AUTO-Explorer: Automated Data Collection for GUI Agent
- Learning from Online Videos at Inference Time for Computer-Use Agents
- Visual Spatial Tuning
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- Test-Time Adaptation for LLM Agents via Environment Interaction
- Scaling Agent Learning via Experience Synthesis
- LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
- GUI Knowledge Bench: Revealing the Knowledge Gap Behind VLM Failures in GUI Tasks
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
- Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
- SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
- SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
- TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
- Agents' Last Exam
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- MGA: Memory-Driven GUI Agent for Observation-Centric Interaction
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
- GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- SCoPE VLM: Selective Context Processing for Efficient Document Navigation in Vision-Language Models
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
- Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- Experience-Driven Exploration for Efficient API-Free AI Agents
- GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
- LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
- ReUseIt: Synthesizing Reusable AI Agent Workflows for Web Automation
- UniCode: A Framework for Generating High Quality Competitive Coding Problems
- ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
- A Survey on Agentic Multimodal Large Language Models
- Scaling Agents for Computer Use
- R-WoM: Retrieval-augmented World Model For Computer-use Agents
- Don't Just Fine-tune the Agent, Tune the Environment
- Polar: Agentic RL on Any Harness at Scale
- Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
- Auto-scaling Continuous Memory for GUI Agent
- Fundamentals of Building Autonomous LLM Agents
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- Agent Learning via Early Experience
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- Toward Systems Foundations for Agentic Exploration
- When Should Users Check? A Decision-Theoretic Model of Confirmation Frequency in Multi-Step AI Agent Tasks
- A Case for Declarative LLM-friendly Interfaces for Improved Efficiency of Computer-Use Agents
- Watch and Learn: Learning to Use Computers from Online Videos
- BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
- GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
- MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
- WALT: Web Agents that Learn Tools
- Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
- GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- SCUBA: Salesforce Computer Use Benchmark
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- Beyond Manuals and Tasks: Instance-Level Context Learning for LLM Agents
- Retrieval-augmented GUI Agents with Generative Guidelines
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
- Estimating the Empowerment of Language Model Agents
- A Benchmark for Localizing Code and Non-Code Issues in Software Projects
- ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
- Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
- How Benchmarks Mis-Score Computer-Use Agents
- Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale
- Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
- AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration
- Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents
- Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
- GPA: Learning GUI Process Automation from Demonstrations
- Mano Technical Report
- GUI-ReWalk: Massive Data Generation for GUI Agent via Stochastic Exploration and Intent-Aware Reasoning
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
- InfraMind: A Novel Exploration-based GUI Agentic Framework for Mission-critical Industrial Management
- PrivWeb: Unobtrusive and Content-aware Privacy Protection For Web Agents
- Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration
- Towards Understanding Visual Grounding in Visual Language Models
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- How well can LLMs provide planning feedback in grounded environments?
- Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- IndusGCC: A Data Benchmark and Evaluation Framework for GUI-Based General Computer Control in Industrial Automation
- AWorld: Orchestrating the Training Recipe for Agentic AI
- Evaluating Language Model Reasoning about Confidential Information
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
- OpenCUA: Open Foundations for Computer-Use Agents
- Reinforcement Learning for Large Model: A Survey
- SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows
- DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- NatureGAIA: Pushing the Frontiers of GUI Agents with a Challenging Benchmark and High-Quality Trajectory Dataset
- Measuring Harmfulness of Computer-Using Agents
Related