CogAgent: A Visual Language Model for GUI Agents
2023/12/14 by Wenyi Hong, Weihan Wang, Hong, Wenyi +19 · 124 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2312.08914
openalex publication_date 2023/12/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
People are spending an enormous amount of time on digital devices through graphical user interfaces (GUIs), e.g., computer or smartphone screens. Large language models (LLMs) such as ChatGPT can assist people in tasks like writing emails, but struggle to understand and interact with GUIs, thus limiting their potential to increase automation levels. In this paper, we introduce CogAgent, an 18-billion-parameter visual language model (VLM) specializing in GUI understanding and navigation. By utilizing both low-resolution and high-resolution image encoders, CogAgent supports input at a resolution of 1120*1120, enabling it to recognize tiny page elements and text. As a generalist visual language model, CogAgent achieves the state of the art on five text-rich and four general VQA benchmarks, including VQAv2, OK-VQA, Text-VQA, ST-VQA, ChartQA, infoVQA, DocVQA, MM-Vet, and POPE. CogAgent, using only screenshots as input, outperforms LLM-based methods that consume extracted HTML text on both PC and Android GUI navigation tasks -- Mind2Web and AITW, advancing the state of the art. The model and codes are available at https://github.com/THUDM/CogVLM, with a new version of CogAgent-9B-20241220 available at https://github.com/THUDM/CogAgent.
Cited by
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
- One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents
- EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
- Multi-Agent LLM Committees for Autonomous Software Beta Testing
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
- MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
- From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents
- Beyond Training: Enabling Self-Evolution of Agents with MOBIMEM
- UniVCD: A New Method for Unsupervised Change Detection in the Open-Vocabulary Era
- Using GUI Agent for Electronic Design Automation
- GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection
- MVP: Multiple View Prediction Improves GUI Grounding
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
- Binary Verification for Zero-Shot Vision
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- OSGym: Super-Scalable Distributed Data Engine for Generalizable Computer Agents
- UI2CodeN: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- Grounding Computer Use Agents on Human Demonstrations
- AUTO-Explorer: Automated Data Collection for GUI Agent
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
- Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games
- GUI Knowledge Bench: Revealing the Knowledge Gap Behind VLM Failures in GUI Tasks
- Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- Mitigating Coordinate Prediction Bias from Positional Encoding Failures
- GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
- ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- Experience-Driven Exploration for Efficient API-Free AI Agents
- GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
- Composition-Grounded Instruction Synthesis for Visual Reasoning
- ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Auto-scaling Continuous Memory for GUI Agent
- Agent Learning via Early Experience
- ReInAgent: A Context-Aware GUI Agent Enabling Human-in-the-Loop Mobile Task Navigation
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- Information Seeking for Robust Decision Making under Partial Observability
- WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
- From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- GUI-PRA: Process Reward Agent for GUI Tasks
- MMPB: It's Time for Multi-Modal Personalization
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- Log2Plan: An Adaptive GUI Automation Framework Integrated with Task Mining Approach
- D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
- Benchmarking MLLM-based Web Understanding: Reasoning, Robustness and Safety
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
- Embodied AI: From LLMs to World Models
- Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models
- GPA: Learning GUI Process Automation from Demonstrations
- Mano Technical Report
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- Towards Understanding Visual Grounding in Visual Language Models
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
- VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
- SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- MobiAgent: A Systematic Framework for Customizable Mobile Agents
- KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- Cybernaut: Towards Reliable Web Automation
- PG-Agent: An Agent Powered by Page Graph
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
- FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
- Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- Cognitive Duality for Adaptive Web Agents
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement
- The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding
Related