Cognitive Architectures for Language Agents
2023/09/05 by Theodore R. Sumers, Sumers, Theodore R., Shunyu Yao +5 · 4 voices · 124 citations
Computer Science · Psychology · Social Sciences · #Action (physics) #Artificial intelligence #Chaining #Cognition #Cognitive science #Computer science #Data science #Knowledge management #Language and cultural evolution #Modular design #Natural Language Processing Techniques #Process (computing) #Programming language #Psychology #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2309.02427
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/09/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
Recent efforts have augmented large language models (LLMs) with external resources (e.g., the Internet) or internal control flows (e.g., prompt chaining) for tasks requiring grounding or reasoning, leading to a new class of language agents. While these agents have achieved substantial empirical success, we lack a systematic framework to organize existing agents and plan future developments. In this paper, we draw on the rich history of cognitive science and symbolic artificial intelligence to propose Cognitive Architectures for Language Agents (CoALA). CoALA describes a language agent with modular memory components, a structured action space to interact with internal memory and external environments, and a generalized decision-making process to choose actions. We use CoALA to retrospectively survey and organize a large body of recent work, and prospectively identify actionable directions towards more capable agents. Taken together, CoALA contextualizes today's language agents within the broader history of AI and outlines a path towards language-based general intelligence.
Cited by
- PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents
- Supra Cognitive Modes: A Routed Architecture for Agent Memory
- Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents
- Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
- Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
- Binding Drift in Multi-Step Tool-Augmented Agents
- Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture
- Trajectory-Aware Retrieval Agents for Temporal Decision- Making
- Simulating Human Memory with Language Models
- Is Grep All You Need? How Agent Harnesses Reshape Agentic Search
- Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
- MAPS: Modeling Co-Existing Subjective Perspectives and Shared Meaning in Multi-Agent Cognitive Dialogue
- FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- Position: Modular Memory is the Key to Continual Learning Agents
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
- El Agente: An autonomous agent for quantum chemistry
- From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps
- SymStep: Symbolic Step Verification for Logical Reasoning
- Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
- From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
- El Agente Estructural: An Artificially Intelligent Molecular Editor
- Unifying Dynamic Tool Creation and Cross-Task Experience Sharing through Cognitive Memory Architecture
- Towards a Science of Scaling Agent Systems
- Learning Evolving Latent Strategies for Multi-Agent Language Systems without Model Fine-Tuning
- EWE: An Agentic Framework for Extreme Weather Analysis
- Episodic Memory in Agentic Frameworks: Suggesting Next Tasks
- Smarter Together: Creating Agentic Communities of Practice through Shared Experiential Learning
- AUTO-Explorer: Automated Data Collection for GUI Agent
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval
- The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
- Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype?
- Agentic Meta-Orchestrator for Multi-task Copilots
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems
- Code-enabled language models can outperform reasoning models on diverse tasks
- Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
- From Craft to Constitution: A Governance-First Paradigm for Principled Agent Engineering
- Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
- How can we assess human-agent interactions? Case studies in software agent design
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- NL2GenSym: Natural Language to Generative Symbolic Rules for SOAR Cognitive Architecture via Large Language Models
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- Agent Learning via Early Experience
- Opponent Shaping in LLM Agents
- A Qualitative Comparative Evaluation of Cognitive and Generative Theories
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
- LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
- AutoMem: Automated Learning of Memory as a Cognitive Skill
- An LLM-based multi-agent framework for agile effort estimation
- Towards Fully Automated Molecular Simulations: Multi-Agent Framework for Simulation Setup and Force Field Extraction
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
- Guidelines for Empirical Studies in Software Engineering involving Large Language Models
- Cognitive Agents Powered by Large Language Models for Agile Software Project Management
- Virtual Community: An Open World for Humans, Robots, and Society
- Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- Cognitive Duality for Adaptive Web Agents
- Sari Sandbox: A Virtual Retail Store Environment for Embodied AI Agents
- PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
- GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts
- Games Agents Play: Towards Transactional Analysis in LLM-based Multi-Agent Systems
- MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
- AgentFly: Extensible and Scalable Reinforcement Learning for LM Agents
- Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
- Agentic Large Language Models for Conceptual Systems Engineering and Design
- Agent Safety Alignment via Reinforcement Learning
- Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
- ROS Help Desk: GenAI Powered, User-Centric Framework for ROS Error Diagnosis and Debugging
- RALLY: Role-Adaptive LLM-Driven Yoked Navigation for Agentic UAV Swarms
- GAIus: Combining Genai with Legal Clauses Retrieval for Knowledge-based Assistant
- Ella: Embodied Social Agents with Lifelong Memory
- TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
- HiBerNAC: Hierarchical Brain-emulated Robotic Neural Agent Collective for Disentangling Complex Manipulation
- KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
- Context manipulation attacks : Web agents are susceptible to corrupted memory
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
- Querying Large Automotive Software Models: Agentic vs. Direct LLM Approaches
- Tiered Agentic Oversight: A Hierarchical Multi-Agent System for Healthcare Safety
- Can Theoretical Physics Research Benefit from Language Agents?
- Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
- Build Agent Advocates, Not Platform Agents
- AI-Native Brand Identity: From Visual Recognition to Cryptographic Verification
- The Future of Continual Learning in the Era of Foundation Models: Three Key Directions
- MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
- ARIA: Training Language Agents with Intention-Driven Reward Aggregation
- Cross-Task Experiential Learning on LLM-based Multi-Agent Collaboration
- Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
- AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
- VECSR: Virtually Embodied Common Sense Reasoning System
- Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
- SEM: Reinforcement Learning for Search-Efficient Large Language Models
- Applying Cognitive Design Patterns to General LLM Agents
- MARK: Memory Augmented Refinement of Knowledge
- Graph-based Agent Memory: Taxonomy, Techniques, and Applications
- Quo Vadis, World Modeling?
- Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
- Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
- The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
- Characterizing AI Agents for Alignment and Governance
- Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
- A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology
- Systematic Bias in Large Language Models: Discrepant Response Patterns in Binary vs. Continuous Judgment Tasks
- Mesh Memory Protocol: Semantic Infrastructure for Multi-Agent LLM Systems
- Structured LLM Reasoning for Zero-Shot Human--Robot Coordination Under Hidden Goals
- CogniFold: Always-On Proactive Memory via Cognitive Folding
- IMPersona: Evaluating Individual Level LM Impersonation
- Toward Generation of Test Cases from Task Descriptions via History-aware Planning
- GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
- Review of Case-Based Reasoning for LLM Agents: Theoretical Foundations, Architectural Components, and Cognitive Integration
- A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
- Inherent and emergent liability issues in LLM-based agentic systems: a principal-agent perspective
Discussions
Related