UI-TARS: Pioneering Automated GUI Interaction with Native Agents
2025/01/21 by Yujia Qin, Yining Ye, Qin, Yujia +64 · 144 citations
Computer Science · Business, Management and Accounting · #Multi-Agent Systems and Negotiation #Mobile Agent-Based Network Management #Business Process Modeling and Analysis
paper · pdf · doi:10.48550/arxiv.2501.12326
Abstract
This paper introduces UI-TARS, a native GUI agent model that solely perceives the screenshots as input and performs human-like interactions (e.g., keyboard and mouse operations). Unlike prevailing agent frameworks that depend on heavily wrapped commercial models (e.g., GPT-4o) with expert-crafted prompts and workflows, UI-TARS is an end-to-end model that outperforms these sophisticated frameworks. Experiments demonstrate its superior performance: UI-TARS achieves SOTA performance in 10+ GUI agent benchmarks evaluating perception, grounding, and GUI task execution. Notably, in the OSWorld benchmark, UI-TARS achieves scores of 24.6 with 50 steps and 22.7 with 15 steps, outperforming Claude (22.0 and 14.9 respectively). In AndroidWorld, UI-TARS achieves 46.6, surpassing GPT-4o (34.5). UI-TARS incorporates several key innovations: (1) Enhanced Perception: leveraging a large-scale dataset of GUI screenshots for context-aware understanding of UI elements and precise captioning; (2) Unified Action Modeling, which standardizes actions into a unified space across platforms and achieves precise grounding and interaction through large-scale action traces; (3) System-2 Reasoning, which incorporates deliberate reasoning into multi-step decision making, involving multiple reasoning patterns such as task decomposition, reflection thinking, milestone recognition, etc. (4) Iterative Training with Reflective Online Traces, which addresses the data bottleneck by automatically collecting, filtering, and reflectively refining new interaction traces on hundreds of virtual machines. Through iterative training and reflection tuning, UI-TARS continuously learns from its mistakes and adapts to unforeseen situations with minimal human intervention. We also analyze the evolution path of GUI agents to guide the further development of this domain.
Cited by
- Agentic Entropy-Balanced Policy Optimization
- Scaling GUI Agents with Visual State Transitions
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
- Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction
- MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
- AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
- SpatialTree: How Spatial Abilities Branch Out in MLLMs
- EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- Step-GUI Technical Report
- From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
- Using GUI Agent for Electronic Design Automation
- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection
- MVP: Multiple View Prediction Improves GUI Grounding
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge
- Fara-7B: An Efficient Agentic Model for Computer Use
- UI-CUBE: Enterprise-Grade Computer Use Agent Benchmarking Beyond Task Accuracy to Operational Reliability
- D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
- DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- OSGym: Super-Scalable Distributed Data Engine for Generalizable Computer Agents
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- Grounding Computer Use Agents on Human Demonstrations
- GUI-360^∘: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- Context Engineering 2.0: The Context of Context Engineering
- GUI Knowledge Bench: Revealing the Knowledge Gap Behind VLM Failures in GUI Tasks
- Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
- Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- MGA: Memory-Driven GUI Agent for Observation-Centric Interaction
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- LightAgent: Mobile Agentic Foundation Models
- GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
- ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
- CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
- CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
- AndroidControl-Curated: Revealing the True Potential of GUI Agents through Benchmark Purification
- Unbiased Gradient Low-Rank Projection
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- Experience-Driven Exploration for Efficient API-Free AI Agents
- GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
- ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
- A Survey on Agentic Multimodal Large Language Models
- R-WoM: Retrieval-augmented World Model For Computer-use Agents
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
- Auto-scaling Continuous Memory for GUI Agent
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
- When Should Users Check? A Decision-Theoretic Model of Confirmation Frequency in Multi-Step AI Agent Tasks
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- 3Dify: a Framework for Procedural 3D-CG Generation Assisted by LLMs Using MCP and RAG
- Watch and Learn: Learning to Use Computers from Online Videos
- Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes
- GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
- Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- SCUBA: Salesforce Computer Use Benchmark
- Scaling Synthetic Task Generation for Agents via Exploration
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining Approach
- RISK: A Framework for GUI Agents in E-commerce Risk Management
- ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
- D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- Mano Technical Report
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
- DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- InfraMind: A Novel Exploration-based GUI Agentic Framework for Mission-critical Industrial Management
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
- Interaction-Driven Browsing: A Human-in-the-Loop Conceptual Framework Informed by Human Web Browsing for Browser-Using Agents
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- WebSight: A Vision-First Architecture for Robust Web Agents
- VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
- SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- A Multimodal GUI Architecture for Interfacing with LLM-Based Conversational Assistants
- MobiAgent: A Systematic Framework for Customizable Mobile Agents
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
- PG-Agent: An Agent Powered by Page Graph
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- PerPilot: Personalizing VLM-based Mobile Agents via Memory and Exploration
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
- OpenCUA: Open Foundations for Computer-Use Agents
- MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
- Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
- Reinforcement Learning for Large Model: A Survey
- Memp: Exploring Agent Procedural Memory
- DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- CoAct-1: Computer-using Multi-Agent System with Coding Actions
- NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- NatureGAIA: Pushing the Frontiers of GUI Agents with a Challenging Benchmark and High-Quality Trajectory Dataset
- Measuring Harmfulness of Computer-Using Agents
Related