CogAgent: A Visual Language Model for GUI Agents
2023/12/14 by Wenyi Hong, Weihan Wang, Hong, Wenyi +19 · 249 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2312.08914
openalex publication_date 2023/12/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
People are spending an enormous amount of time on digital devices through graphical user interfaces (GUIs), e.g., computer or smartphone screens. Large language models (LLMs) such as ChatGPT can assist people in tasks like writing emails, but struggle to understand and interact with GUIs, thus limiting their potential to increase automation levels. In this paper, we introduce CogAgent, an 18-billion-parameter visual language model (VLM) specializing in GUI understanding and navigation. By utilizing both low-resolution and high-resolution image encoders, CogAgent supports input at a resolution of 1120*1120, enabling it to recognize tiny page elements and text. As a generalist visual language model, CogAgent achieves the state of the art on five text-rich and four general VQA benchmarks, including VQAv2, OK-VQA, Text-VQA, ST-VQA, ChartQA, infoVQA, DocVQA, MM-Vet, and POPE. CogAgent, using only screenshots as input, outperforms LLM-based methods that consume extracted HTML text on both PC and Android GUI navigation tasks -- Mind2Web and AITW, advancing the state of the art. The model and codes are available at https://github.com/THUDM/CogVLM, with a new version of CogAgent-9B-20241220 available at https://github.com/THUDM/CogAgent.
Cited by
- LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
- One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents
- EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
- Multi-Agent LLM Committees for Autonomous Software Beta Testing
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
- Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
- UniVCD: A New Method for Unsupervised Change Detection in the Open-Vocabulary Era
- Using GUI Agent for Electronic Design Automation
- GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection
- MVP: Multiple View Prediction Improves GUI Grounding
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
- Binary Verification for Zero-Shot Vision
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- OSGym: Scalable OS Infra for Computer Use Agents
- UI2CodeN: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- Grounding Computer Use Agents on Human Demonstrations
- AUTO-Explorer: Automated Data Collection for GUI Agent
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
- Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games
- GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
- Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- Mitigating Coordinate Prediction Bias from Positional Encoding Failures
- GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
- ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- Experience-Driven Exploration for Efficient API-Free AI Agents
- GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
- Composition-Grounded Data Synthesis for Visual Reasoning
- ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Auto-scaling Continuous Memory for GUI Agent
- Agent Learning via Early Experience
- ReInAgent: A Context-Aware GUI Agent Enabling Human-in-the-Loop Mobile Task Navigation
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- Information Seeking for Robust Decision Making under Partial Observability
- WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
- From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- GUI-PRA: Process Reward Agent for GUI Tasks
- MMPB: It's Time for Multi-Modal Personalization
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining Approach
- D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
- Benchmarking MLLM-based Web Understanding: Reasoning, Robustness and Safety
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
- Embodied AI: From LLMs to World Models
- Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models
- GPA: Learning GUI Process Automation from Demonstrations
- Mano Technical Report
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- Towards Understanding Visual Grounding in Visual Language Models
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
- VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
- SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- MobiAgent: A Systematic Framework for Customizable Mobile Agents
- KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- Cybernaut: Towards Reliable Web Automation
- PG-Agent: An Agent Powered by Page Graph
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
- FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
- Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- Cognitive Duality for Adaptive Web Agents
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement
- The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- General Agentic Planning Through Simulative Reasoning with World Models
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
- Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
- GUI-G2: Gaussian Reward Modeling for GUI Grounding
- MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation
- Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs
- MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
- WebGuard: Building a Generalizable Guardrail for Web Agents
- RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
- AD2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
- PyVision: Agentic Vision with Dynamic Tooling
- VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
- MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
- R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
- GTA1: GUI Test-time Scaling Agent
- Vision-Language Models Can't See the Obvious
- Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion
- Hijacking JARVIS: Benchmarking Mobile GUI Agents against Unprivileged Third Parties
- Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Ella: Embodied Social Agents with Lifelong Memory
- What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
- ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding
- Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
- GamerAstra: Supporting 2D Non-Twitch Video Games for Blind and Low-Vision Players through a Multi-Agent Framework
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- Beyond Syntax: Action Semantics Learning for App Agents
- Understanding GUI Agent Localization Biases through Logit Sharpness
- GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
- 3DGS-IEval-15K: A Large-scale Image Quality Evaluation Database for 3D Gaussian-Splatting
- Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation
- Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
- Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
- Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
- macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
- DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models
- Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
- FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
- AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
- Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents
- Temac: Multi-Agent Collaboration for Automated Web GUI Testing
- VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making
- Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
- Grounded Reinforcement Learning for Visual Reasoning
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
- XBOUND: Exploring Capability Boundaries of Device-Control Agents at the State Level
- BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
- AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model
- MSD-LLM: Predicting Ship Detention in Port State Control Inspections with Large Language Model
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World
- TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
- Qwen-CUA: Native Computer Use for (almost) Everything
- Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
- GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
- ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
- ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
- Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
- GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
- Confidence-Regulated Generative Diffusion Models for Reliable AI Agent Migration in Vehicular Metaverses
- MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
- Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
- GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
- Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
- Group-in-Group Policy Optimization for LLM Agent Training
- UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage Designs
- Interpretable Risk Mitigation in LLM Agent Systems
- Visual Instruction Tuning with Chain of Region-of-Interest
- EcoAgent: An Efficient Device-Cloud Collaborative Multi-Agent Framework for Mobile Automation
- Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
- Visual Test-time Scaling for GUI Agent Grounding
- ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
- Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
- UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
- Anticipatory Planning for Multimodal AI Agents
- Video-Based Reward Modeling for Computer-Use Agents
- Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
- HLL: Can Agents Cross Humanity's Last Line of Verification?
- Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
- NGENT: Next-Generation AI Agents Must Integrate Multi-Domain Abilities to Achieve Artificial General Intelligence
- A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning
- Generative Visual Code Mobile World Models
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- AndroidGen: Building an Android Language Agent under Data Scarcity
- AppAgent-Claw: CLI Is All You Need for GUI Automation
- HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
- OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
- Mind the Gap: Action Rebinding Attacks against Android GUI Agents
- Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
- GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
- Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
- Secure and Efficient Access Control for Computer-Use Agents via Context Space
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- "Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents
- On the Robustness of GUI Grounding Models Against Image Attacks
- Building LLM Agents by Incorporating Insights from Computer Systems
- Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
- UFO2: The Desktop AgentOS
- Manipulating Multimodal Agents via Cross-Modal Prompt Injection
- Toward Generation of Test Cases from Task Descriptions via History-aware Planning
- InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
- LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
- StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
- AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
- LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
Related