From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms
2026/01/01 by Jinghao Luo, Yuchen Tian, Chuxue Cao +6 · 1 voice
Computer Science · Engineering · #Action (physics) #Data collection #Encoding (memory) #Mechanism (biology) #Mobile Agent-Based Network Management #Modular Robots and Swarm Intelligence #Multi-Agent Systems and Negotiation #Troubleshooting #cs.AI #cs.CL
paper · pdf · doi:10.18653/v1/2026.findings-acl.2069
openalex publication_date 2026/01/01 · arxiv published 2026/05/07 · arxiv updated 2026/05/07 · openalex created_date 2026/07/02 · openalex updated_date 2026/07/29
Abstract
Large Language Model (LLM)-based agents have fundamentally reshaped artificial intelligence by integrating external tools and planning capabilities.While memory mechanisms have emerged as the architectural cornerstone of these systems, current research remains fragmented, oscillating between operating system engineering and cognitive science.This theoretical divide prevents a unified view of technological synthesis and a coherent evolutionary perspective.To bridge this gap, this survey proposes a novel evolutionary framework for LLM agent memory mechanisms, formalizing the development process into three stages: Storage (trajectory preservation), Reflection (trajectory refinement), and Experience (trajectory abstraction).We first formally define these three stages before analyzing the three core drivers of this evolution: the necessity for long-range consistency, the challenges in dynamic environments, and the ultimate goal of continual learning.Furthermore, we specifically explore two transformative mechanisms in the frontier Experience stage: active exploration and cross-trajectory abstraction.By synthesizing these disparate views, this work offers robust design principles and a clear roadmap for the development of next-generation LLM agents.
Citations
- Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory
- Experiential Reflective Learning for Self-Improving LLM Agents
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- LoongFlow: Directed Evolutionary Search via a Cognitive Plan-Execute-Summarize Paradigm
- Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
- End-to-End Test-Time Training for Long Context
- Context as a Tool: Context Management for Long-Horizon SWE-Agents
- Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
- MemR3: Memory Retrieval via Reflective Reasoning for LLM Agents
- Memory as Resonance: A Biomimetic Architecture for Infinite Context Memory on Ergodic Phonetic Manifolds
- Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
- MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs
- Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
- GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
- MemEvolve: Meta-Evolution of Agent Memory Systems
- From Personalization to Prejudice: Bias and Discrimination in Memory-Enhanced AI Agents for Recruitment
- Reinforcement Learning for Self-Improving Agent with Skill Library
- MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
- Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution
- Guided Self-Evolving LLMs with Minimal Human Supervision
- WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
- AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
- O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents
- AgentEvolver: Towards Efficient Self-Evolving Agent System
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- Beyond A Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
- AgentFold: Long-Horizon Web Agents with Proactive Context Management
- WebATLAS: An LLM Agent with Experience-Driven Memory and Action Simulation
- DeepAgent: A General Reasoning Agent with Scalable Toolsets
- Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
- PolySkill: Learning Generalizable Skills Through Polymorphic Abstraction
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- Internalizing World Models via Self-Play Finetuning for Agentic RL
- Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
- Scaling Long-Horizon LLM Agent via Context-Folding
- Preference-Aware Memory Update for Long-Term LLM Agents
- Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory and User Profiles
- ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
- CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension
- LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
- Constructing coherent spatial memory in LLM agents through graph rectification
- EvolProver: Advancing Automated Theorem Proving by Evolving Formalized Problems via Symmetry and Difficulty
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
- Where LLM Agents Fail and How They can Learn From Failures
- Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration
- MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service
- SEDM: Scalable Self-Evolving Distributed Memory for Agents
- REMI: A Novel Causal Schema Memory Architecture for Personalized Lifestyle Recommendation Agents
- Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- HiPlan: Hierarchical Planning for LLM-Based Agents with Adaptive Global-Local Guidance
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- Coarse-to-Fine Grounded Memory for LLM Agent Planning
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information
- GraphCogent: Mitigating LLMs' Working Memory Constraints via Multi-Agent Collaboration in Complex Graph Understanding
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- Narrative Memory in Machines: Multi-Agent Arc Extraction in Serialized TV
- RCR-Router: Efficient Role-Aware Context Routing for Multi-Agent LLM Systems with Structured Memory
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- SWE-Exp: Experience-Driven Software Issue Resolution
- CoEx -- Co-evolving World-model and Exploration
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
- AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- Contextual Experience Replay for Self-Improvement of Language Agents
- Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
- WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
- Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
- LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
- Group-in-Group Policy Optimization for LLM Agent Training
- MemEngine: A Unified and Modular Library for Developing Advanced Memory of LLM-based Agents
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- Monte Carlo Planning with Large Language Model for Text-Based Game Agents
- From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs
- Evaluating the Goal-Directedness of Large Language Models
- HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
- Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
- Inducing Programmatic Skills for Agentic Tasks
- MemInsight: Autonomous Memory Augmentation for LLM Agents
- Toward Multi-Session Personalized Conversation: A Large-Scale Dataset and Hierarchical Tree Framework for Implicit Reasoning
- A-MEM: Agentic Memory for LLM Agents
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
- Minerva: A Programmable Memory Test Benchmark for Language Models
- Memento No More: Coaching AI Agents to Master Multiple Tasks via Hints Internalization
- Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos
- Curiosity-Driven Reinforcement Learning from Human Feedback
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs
- AgentRefine: Enhancing Agent Generalization through Refinement Tuning
- Titans: Learning to Memorize at Test Time
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- COLD: Causal reasOning in cLosed Daily activities
- WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
- Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
- SHARE: Shared Memory-Aware Open-Domain Long-Term Dialogue Dataset Constructed from Movie Script
- MADial-Bench: Towards Real-world Evaluation of Memory-Augmented Dialogue Generation
- Symbolic Working Memory Enhances Language Models for Complex Rule Application
- IDEA: Enhancing the Rule Learning Ability of Large Language Model Agent through Induction, Deduction, and Abduction
- Enhancing Agent Learning through World Dynamics Modeling
- AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
- BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
- StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
- GrowOVER: How Can LLMs Adapt to Growing Real-World Knowledge?
- SelfGoal: Your Language Agents Already Know How to Achieve High-level Goals
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
- A Survey on the Memory Mechanism of Large Language Model based Agents
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
- Larimar: Large Language Models with Episodic Memory Control
- VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- PerLTQA: A Personal Long-Term Memory Dataset for Memory Classification, Retrieval, and Synthesis in Question Answering
- On the Multi-turn Instruction Following for Conversational Web Agents
- AgentScope: A Flexible yet Robust Multi-Agent Platform
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment
- InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
- Understanding the planning of LLM agents: A survey
- LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
- Personalized Large Language Model Assistant with Evolving Conditional Memory
- Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
- CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization
- MemGPT: Towards LLMs as Operating Systems
- Large Language Models Cannot Self-Correct Reasoning Yet
- Efficient Streaming Language Models with Attention Sinks
- The Rise and Potential of Large Language Model Based Agents: A Survey
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation
- Counterfactually Auditable Lifecycle Certification for Autonomous Agents
- ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory
- SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks
- How Language Model Hallucinations Can Snowball
- Reflexion: Language Agents with Verbal Reinforcement Learning
- LLaMA: Open and Efficient Foundation Language Models
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- RealTime QA: What's the Answer Right Now?
- TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Mind the Gap: Assessing Temporal Generalization in Neural Language Models
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- Beyond Context Limits: Subconscious Threads for Long-Horizon Reasoning
- Large Language Models and Causal Inference in Collaboration: A Survey
- Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
- DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
Discussions
Related