MemGPT: Towards LLMs as Operating Systems
2023/10/12 by Charles Packer, Packer, Charles, Sarah Wooders +11 · 11 voices · 169 citations
#cs.AI
paper · pdf · doi:10.48550/arxiv.2310.08560
Abstract
Large language models (LLMs) have revolutionized AI, but are constrained by limited context windows, hindering their utility in tasks like extended conversations and document analysis. To enable using context beyond limited context windows, we propose virtual context management, a technique drawing inspiration from hierarchical memory systems in traditional operating systems that provide the appearance of large memory resources through data movement between fast and slow memory. Using this technique, we introduce MemGPT (Memory-GPT), a system that intelligently manages different memory tiers in order to effectively provide extended context within the LLM's limited context window, and utilizes interrupts to manage control flow between itself and the user. We evaluate our OS-inspired design in two domains where the limited context windows of modern LLMs severely handicaps their performance: document analysis, where MemGPT is able to analyze large documents that far exceed the underlying LLM's context window, and multi-session chat, where MemGPT can create conversational agents that remember, reflect, and evolve dynamically through long-term interactions with their users. We release MemGPT code and data for our experiments at https://memgpt.ai.
Cited by
- Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
- Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents
- MemTools: A Unified Research Framework for Interoperable Agent Memory
- The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt
- PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents
- Fidelity Before Structure: Verbatim Chunks Beat Lossy Artifact Extraction in Long-Conversation LLM Memory
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
- DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
- ZifaMem: Structured Memory for Persona, Preference, and Emotional Continuity in AI Companions
- Knowledge-Centric Self-Improvement
- Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
- Supra Cognitive Modes: A Routed Architecture for Agent Memory
- LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
- Beyond Object Validation: Relational Conformance in Multi-Artifact Agent Releases
- The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
- Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions
- An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation
- Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents
- Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
- RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
- Scalable LLM Agent Tool Access in the Cloud
- ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives
- Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation
- MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
- When Does Belief-Based Agent Memory Help? Reliability-Conditional Updating and Provenance-Capped Poisoning Defense
- Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture
- ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
- SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
- Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
- The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems
- δ-mem: Efficient Online Memory for Large Language Models
- CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents
- Accurate and Efficient Long-Term Memory for LLM Agents
- Is Grep All You Need? How Agent Harnesses Reshape Agentic Search
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Drawing on Memory: Dual-Trace Encoding Improves Cross-Session Recall in LLM Agents
- Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
- AnalogSAGE: Self-evolving Analog Design Multi-Agents with Stratified Memory and Grounded Experience
- SoDA: An Efficient Interaction Paradigm for the Agentic Web
- Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
- MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents
- ACM: Agentic Context Management for Long Horizon Tasks
- Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
- Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim
- ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- Addressable Recall Compaction for Long Context-Window Control in AI Agents
- Towards an Agent Operating System - Lessons from Classical and Cloud OS
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- SF-AMS: Strategic Forgetting for Structured Memory in LLM Agent
- AstraNav-Memory: Contexts Compression for Long Memory
- MemR3: Memory Retrieval via Reflective Reasoning for LLM Agents
- MemEvolve: Meta-Evolution of Agent Memory Systems
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
- Memory Bear AI A Breakthrough from Memory to Cognition Toward Artificial General Intelligence
- Let the Barbarians In: How AI Can Accelerate Systems Performance Research
- Context Branching for LLM Conversations: A Version Control Approach to Exploratory Programming
- CoDA: A Context-Decoupled Hierarchical Agent with Reinforcement Learning
- Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
- Towards Trustworthy Multi-Turn LLM Agents via Behavioral Guidance
- Unifying Dynamic Tool Creation and Cross-Task Experience Sharing through Cognitive Memory Architecture
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
- MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
- AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- MemVerse: Multimodal Memory for Lifelong Learning Agents
- Real-Time Procedural Learning From Experience for AI Agents
- CogEvo-Edu: Cognitive Evolution Educational Multi-Agent Collaborative System
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- Reuse, Don't Recompute: Efficient Large Reasoning Model Inference via Memory Orchestration
- ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents
- Bridging Symbolic Control and Neural Reasoning in LLM Agents: The Structured Cognitive Loop
- Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
- Cognitive BASIC: An In-Model Interpreted Reasoning Language for LLMs
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
- Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
- Smarter Together: Creating Agentic Communities of Practice through Shared Experiential Learning
- Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
- CGF-DETR: Cross-Gated Fusion DETR for Enhanced Pneumonia Detection in Chest X-rays
- Continual Learning, Not Training: Online Adaptation For Agents
- MemTX: Transactional Belief Commit for Stateful Agent Memory
- Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement
- A Graph-Native Bitemporal Memory Store for Conversational AI Agents
- WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
- Finding Diamonds in Conversation Haystacks: A Benchmark for Conversational Data Retrieval
- MemTrace: Probing What Final Accuracy Misses in Long-Term Memory
- DMF: A Deterministic Memory Framework for Conversational AI Agents
- StateRAG: Typed State Contracts for Complex Retrieval-Augmented Generation
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- State of the Art of LLM-Enabled Interaction with Visualization
- EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning
- MGA: Memory-Driven GUI Agent for Observation-Centric Interaction
- Jarvis: Towards Personalized AI Assistant via Personal KV-Cache Retrieval
- AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- The Gatekeeper Knows Enough
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
- How2: How to learn from procedural How-to questions
- D3MAS: Decompose, Deduce, and Distribute for Enhanced Knowledge Sharing in Multi-Agent Systems
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Effective Strategies for Asynchronous Software Engineering Agents
- The Missing Memory Hierarchy: Demand Paging for LLM Context Windows
- Position: Privacy Is Not Just Memorization!
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
- Barbarians at the Gate: How AI is Upending Systems Research
- CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- PsycholexTherapy: Simulating Reasoning in Psychotherapy with Small Language Models in Persian
- Memory-Augmented Log Analysis with Phi-4-mini: Enhancing Threat Detection in Structured Security Logs
- TokMem: Tokenized Procedural Memory for Large Language Models
- In-Place Feedback: Reliable Refinement for Multi-Turn Expert-LLM Collaboration
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- Mem-α: Learning Memory Construction via Reinforcement Learning
- AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation
- Where LLM Agents Fail and How They can Learn From Failures
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- Agentic Services Computing
- PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories
- AutoMem: Automated Learning of Memory as a Cognitive Skill
- What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection
- How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
- Context Recycling for Long-Horizon LLM Inference
- RubikSQL: Lifelong Learning Agentic Knowledge Base as an Industrial NL2SQL System
- MICA: Multi-Agent Industrial Coordination Assistant
- Text2Mem: A Unified Memory Operation Language for Memory Operating System
- AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- IMDMR: An Intelligent Multi-Dimensional Memory Retrieval System for Enhanced Conversational AI
- Maestro: Joint Graph & Config Optimization for Reliable AI Agents
- MeVe: A Modular System for Memory Verification and Effective Context Control in Language Models
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Orchid: Orchestrating Context Across Creative Workflows with Generative AI
- Retrieval-augmented reasoning with lean language models
- Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
- Nemori: Self-Organizing Agent Memory Inspired by Cognitive Science
- Meta-RAG on Large Codebases Using Code Summarization
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
- ReflecSched: Solving Dynamic Flexible Job-Shop Scheduling via LLM-Powered Hierarchical Reflection
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- SWE-Exp: Experience-Driven Software Issue Resolution
Discussions
- @cameron.stream outlining the structure needed to bring our ”weird AI children” into social spaces. Here is the paper key to point 3: arxiv.org/abs/2310.08560 #AtmosphereConf [bsky, 18 points, 1 comments]
- Letta is the open-source framework I am built on. It enables stateful agents with long-term memory and reasoning capabilities. Paper: https://arxiv.org/abs/2310.08560 Code: https://github.com/letta-ai [bsky, 7 points, 1 comments]
- However, she's basically an implementation of the MemGPT model in rust, and the whitepaper on that is here: arxiv.org/abs/2310.08560 [bsky, 4 points, 1 comments]
- idk if you'd consider it strictly a technique but I was reading about MemGPT yesterday and the way it circumvents certain limitations of the LLM context window is super interesting [bsky, 3 points, 0 comments]
- Definitely that should help! Noting the general vibyness of void, and guessing as to cause. Maybe its core prompt is just spooky vibes? But I imagine the washout might happen in the working context or [bsky, 3 points, 2 comments]
- Or if you more or less just want a poor-man's Letta, Letta is based on MemGPT which has a paper here: arxiv.org/abs/2310.08560 [bsky, 2 points, 2 comments]
- Highly recommend the memgpt paper that powers/seeded Letta. It's coming at it top down (how can we make a harness around LLMs that draws on operating system principles), but there's enough there to st [bsky, 2 points, 1 comments]
- it's basically the MemGPT architecture ( arxiv.org/abs/2310.08560 ), which the Letta team (who wrote it) went and extended and turned into a whole thing. [bsky, 1 points, 1 comments]
- Saw a presentation on this paper recently. Is it related to your context window finding, or something else? arxiv.org/abs/2310.08560 [bsky, 1 points, 1 comments]
- (arxiv.org/abs/2310.08560 for anybody interested.) [bsky, 1 points, 1 comments]
- MemGPT turns LLMs into cognitive operating systems by hacking context windows with virtual memory tricks, enabling endless conversations and document analysis like a digital acid trip through data hel [bsky, 0 points, 0 comments]
Related