Gradient Episodic Memory for Continual Learning
2017/06/26 by David López-Paz, Lopez-Paz, David, Marc’Aurelio Ranzato +1 · 275 citations
Computer Science · #Artificial Intelligence (cs.AI) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1706.08840
openalex publication_date 2017/06/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
One major obstacle towards AI is the poor ability of models to solve new problems quicker, and without forgetting previously acquired knowledge. To better understand this issue, we study the problem of continual learning, where the model observes, once and one by one, examples concerning a sequence of tasks. First, we propose a set of metrics to evaluate models learning over a continuum of data. These metrics characterize models not only by their test accuracy, but also in terms of their ability to transfer knowledge across tasks. Second, we propose a model for continual learning, called Gradient Episodic Memory (GEM) that alleviates forgetting, while allowing beneficial transfer of knowledge to previous tasks. Our experiments on variants of the MNIST and CIFAR-100 datasets demonstrate the strong performance of GEM when compared to the state-of-the-art.
Cited by
- Merge before Forget: A Single LoRA Continual Learning via Continual Merging
- SLAM: Structured and Localized Analytic Manifold Adaptation for Forgetting-Immune and Domain-Robust Lifelong VPR
- Dynamic Feedback Engines: Layer-Wise Control for Self-Regulating Continual Learning
- Mixture-of-Experts with Gradient Conflict-Driven Subspace Topology Pruning for Emergent Modularity
- Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios
- InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
- Demonstration-Guided Continual Reinforcement Learning in Dynamic Environments
- AL-GNN: Privacy-Preserving and Replay-Free Continual Graph Learning via Analytic Learning
- Sophia: A Persistent Agent Framework of Artificial Life
- Sequencing to Mitigate Catastrophic Forgetting in Continual Learning
- Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed Memory
- Out-of-Distribution Detection for Continual Learning: Design Principles and Benchmarking
- PPSEBM: An Energy-Based Model with Progressive Parameter Selection for Continual Learning
- Feature Aggregation for Efficient Continual Learning of Complex Facial Expressions
- On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
- Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
- Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
- Multi-Generator Continual Learning for Robust Delay Prediction in 6G
- Heads collapse, features stay: Why Replay needs big buffers
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Bring Your Dreams to Life: Continual Text-to-Video Customization
- Parameter-Efficient Augment Plugin for Class-Incremental Learning
- Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
- MoB: Mixture of Bidders
- Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent
- Stable-Drift: A Patient-Aware Latent Drift Replay Method for Stabilizing Representations in Continual Learning
- Bandit Guided Submodular Curriculum for Adaptive Subset Selection
- Probabilistic Hash Embeddings for Online Learning of Categorical Features
- Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
- Prototype-Guided Non-Exemplar Continual Learning for Cross-subject EEG Decoding
- OpenCML: End-to-End Framework of Open-world Machine Learning to Learn Unknown Classes Incrementally
- Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
- Continual Alignment for SAM: Rethinking Foundation Models for Medical Image Segmentation in Continual Learning
- MedPEFT-CL: Dual-Phase Parameter-Efficient Continual Learning with Medical Semantic Adapter and Bidirectional Memory Consolidation
- Continual Reinforcement Learning for Cyber-Physical Systems: Lessons Learned and Open Challenges
- Parameter Importance-Driven Continual Learning for Foundation Models
- Neuro-Logic Lifelong Learning
- ConSurv: Multimodal Continual Learning for Survival Analysis
- Compact Memory for Continual Logistic Regression
- Learning with Preserving for Continual Multitask Learning
- Mitigating Negative Flips via Margin Preserving Training
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective
- Solving bilevel optimization via sequential minimax optimization
- Sharing the Learned Knowledge-base to Estimate Convolutional Filter Parameters for Continual Image Restoration
- The brain as a blueprint: a survey of brain-inspired approaches to learning in artificial intelligence
- Clustering-Based Weight Orthogonalization for Stabilizing Deep Reinforcement Learning
- Model Inversion with Layer-Specific Modeling and Alignment for Data-Free Continual Learning
- The Art of Not Forgetting A Local Learning Architecture for Continual Learning
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- On the importance of cross-task features for class-incremental learning
- Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
- Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?
- Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
- Knowledge-guided Continual Learning for Behavioral Analytics Systems
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Continual Knowledge Consolidation LORA for Domain Incremental Learning
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- Online Time Series Forecasting with Theoretical Guarantees
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- Mapping Post-Training Forgetting in Language Models at Scale
- CaMiT: A Time-Aware Car Model Dataset for Classification and Generation
- The Three Regimes of Offline-to-Online Reinforcement Learning
- Towards Error Centric Intelligence I, Beyond Observational Learning
- BI-MAML: Balanced Incremental Approach for Meta Learning
- FedGTEA: Federated Class-Incremental Learning with Gaussian Task Embedding and Alignment
- REPAIR: Robust Editing via Progressive Adaptive Intervention and Reintegration
- VAO: Validation-Aligned Optimization for Cross-Task Generative Auto-Bidding
- Source-Free Cross-Domain Continual Learning
- End-to-End Test-Time Training for Long Context
- The Unreasonable Effectiveness of Randomized Representations in Online Continual Graph Learning
- An Attention-based Feature Memory Design for Energy-Efficient Continual Learning
- Rehearsal-free and Task-free Online Continual Learning With Contrastive Prompt
- Overcoming Catastrophic Forgetting by Soft Parameter Pruning
- A Computational Perspective on NeuroAI and Synthetic Biological Intelligence
- Temporal Generalization: A Reality Check
- The Lie of the Average: How Class Incremental Learning Evaluation Deceives You?
- Blockwise Hadamard high-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
- LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
- EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense
- DATS: Distance-Aware Temperature Scaling for Calibrated Class-Incremental Learning
- Adaptive Model Ensemble for Continual Learning
- C2Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning
- Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
- DevFD: Developmental Face Forgery Detection by Learning Shared and Orthogonal LoRA Subspaces
- AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
- SPICED: A Synaptic Homeostasis-Inspired Framework for Unsupervised Continual EEG Decoding
- VidCLearn: A Continual Learning Approach for Text-to-Video Generation
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
- Incremental Multistep Forecasting of Battery Degradation Using Pseudo Targets
- AFT: An Exemplar-Free Class Incremental Learning Method for Environmental Sound Classification
- Global Pre-fixing, Local Adjusting: A Simple yet Effective Contrastive Strategy for Continual Learning
- CL2GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction
- Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
- Unbiased Online Curvature Approximation for Regularized Graph Continual Learning
- Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function
- Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities
- Bilevel Continual Learning
- Flattening Sharpness for Dynamic Gradient Projection Memory Benefits\n Continual Learning
- Online time series prediction using feature adjustment
- Holographic Knowledge Manifolds: A Novel Pipeline for Continual Learning Without Catastrophic Forgetting in Large Language Models
- Selfless Sequential Learning
- A Metaverse: Taxonomy, Components, Applications, and Open Challenges
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Unsupervised Video Continual Learning via Non-Parametric Deep Embedded Clustering
- Adapting to Change: A Comparison of Continual and Transfer Learning for Modeling Building Thermal Dynamics under Concept Drifts
- Revisiting Deepfake Detection: Chronological Continual Learning and the Limits of Generalization
- Class Incremental Continual Learning with Self-Organizing Maps and Variational Autoencoders Using Synthetic Replay
- High-dimensional Asymptotics of Generalization Performance in Continual Ridge Regression
- C-Flat++: Towards a More Efficient and Powerful Framework for Continual Learning
- Shift Detection and Adaptation for Network Intrusion Detection
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Calibrated Semantic Diffusion: A p-Laplacian Synthesis with Learnable Dissipation, Quantified Constants, and Graph-Aware Calibration
- Representative Task Self-selection for Flexible Clustered Lifelong Learning
- Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
- Learning Wisdom from Errors: Promoting LLM's Continual Relation Learning through Exploiting Error Cases
- SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph Embedding
- Index-Aligned Query Distillation for Transformer-based Incremental Object Detection
- Meta Continual Learning
- Can Synthetic Images Conquer Forgetting? Beyond Unexplored Doubts in Few-Shot Class-Incremental Learning
- Dynamic Mixture-of-Experts for Incremental Graph Learning
- CitySeg: A 3D Open Vocabulary Semantic Segmentation Foundation Model in City-scale Scenarios
- GaussianUpdate: Continual 3D Gaussian Splatting Update for Changing Environments
- OpenHAIV: A Framework Towards Practical Open-World Learning
- LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning
- Statistical Theory of Multi-stage Newton Iteration Algorithm for Online Continual Learning
- Lifelong Learner: Discovering Versatile Neural Solvers for Vehicle Routing Problems
- Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
- A Study on Regularization-Based Continual Learning Methods for Indic ASR
- Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
- CRAM: Large-scale Video Continual Learning with Bootstrapped Compression
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- Measuring the stability and plasticity of recommender systems
- Prediction-Oriented Subsampling from Data Streams
- Online Continual Graph Learning
- Exploring Stability-Plasticity Trade-offs for Continual Named Entity Recognition
- Learning to Evolve: Bayesian-Guided Continual Knowledge Graph Embedding
- Towards Field-Ready AI-based Malaria Diagnosis: A Continual Learning Approach
- Continual Learning with Support Boundary Experience Blending
- Forgetting of task-specific knowledge in model merging-based continual learning
- RainbowPrompt: Diversity-Enhanced Prompt-Evolving for Continual Learning
- Modular Delta Merging with Orthogonal Constraints: A Scalable Framework for Continual and Reversible Model Composition
- GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
- Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
- GRID: Scalable Task-Agnostic Prompt-Based Continual Learning for Language Models
- Noradrenergic-inspired gain modulation attenuates the stability gap in joint training
- Look the Other Way: Designing 'Positive' Molecules with Negative Data via Task Arithmetic
- Adapt, But Don't Forget: Fine-Tuning and Contrastive Routing for Lane Detection under Distribution Shift
- Optimal Empirical Risk Minimization under Temporal Distribution Shifts
- RegCL: Continual Adaptation of Segment Anything Model via Model Merging
- A gradient of complementary learning systems emerges through meta-learning
- Auto-Compressing Networks
- Using Continual Learning for Real-Time Detection of Vulnerable Road Users in Complex Traffic Scenarios
- CLA: Latent Alignment for Online Continual Self-Supervised Learning
- Confounder-Free Continual Learning via Recursive Feature Normalization
- Catastrophic Forgetting Mitigation Through Plateau Phase Activity Profiling
- The Bayesian Approach to Continual Learning: An Overview
- Rethinking Query-based Transformer for Continual Image Segmentation
- Information Must Flow: Recursive Bootstrapping for Information Bottleneck in Optimal Transport
- Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
- Meta-Learning Transformers to Improve In-Context Generalization
- Exploring Kolmogorov-Arnold Network Expansions in Vision Transformers for Mitigating Catastrophic Forgetting in Continual Learning
- ACE: Adapting to Changing Environments for Semantic Segmentation
- Continual Multiple Instance Learning with Enhanced Localization for Histopathological Whole Slide Image Analysis
- SHIELD: Secure Hypernetworks for Incremental Expansion Learning Defense
- ETT-CKGE: Efficient Task-driven Tokens for Continual Knowledge Graph Embedding
- A Survey of Continual Reinforcement Learning
- Federated Learning with Fair Averaging
- Little by Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory Experts
- Dealing with the Evil Twins: Improving Random Augmentation by Addressing Catastrophic Forgetting of Diverse Augmentations
- Catastrophic Forgetting Mitigation via Discrepancy-Weighted Experience Replay
- Pathway-based Progressive Inference (PaPI) for Energy-Efficient Continual Learning
- Continual Learning with Columnar Spiking Neural Networks
- Weight Factorization and Centralization for Continual Learning in Speech Recognition
- The Condition Number as a Scale-Invariant Proxy for Information Encoding in Neural Units
- Task-Agnostic Experts Composition for Continual Learning
- SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention
- Knowledge Adaptation as Posterior Correction
- Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
- Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence
- DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models
- CARoL: Context-aware Adaptation for Robot Learning
- Revisiting Clustering of Neural Bandits: Selective Reinitialization for Mitigating Loss of Plasticity
- Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
- SWE-Bench-CL: Continual Learning for Coding Agents
- Continual Hyperbolic Learning of Instances and Classes
- Class-Incremental Learning for Honey Botanical Origin Classification with Hyperspectral Images: A Study with Continual Backpropagation
- Dynamic Mixture of Progressive Parameter-Efficient Expert Library for Lifelong Robot Learning
- SupportNet: solving catastrophic forgetting in class incremental learning with support data
- Adversarial Incremental Learning
- How neural networks find generalizable solutions: Self-tuned annealing in deep learning
- Replay Can Provably Increase Forgetting
- The Future of Continual Learning in the Era of Foundation Models: Three Key Directions
- Class Incremental Learning for Algorithm Selection
- Continual Speech Learning with Fused Speech Features
- Deep Generative Dual Memory Network for Continual Learning
- iDPA: Instance Decoupled Prompt Attention for Incremental Medical Object Detection
- Flashbacks to Harmonize Stability and Plasticity in Continual Learning
- CL-LoRA: Continual Low-Rank Adaptation for Rehearsal-Free Class-Incremental Learning
- Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
- Rethinking Continual Learning with Progressive Neural Collapse
- Unlocking the Power of Rehearsal in Continual Learning: A Theoretical Perspective
- Multi-task Learning for Heterogeneous Multi-source Block-Wise Missing Data
- Continual Learning in Vision-Language Models via Aligned Model Merging
- Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
- LADA: Scalable Label-Specific CLIP Adapter for Continual Learning
- Shared and Private VAEs with Generative Replay for Continual Learning
- Adaptive Budget Allocation for Orthogonal-Subspace Adapter Tuning in LLMs Continual Learning
- LoKI: Low-damage Knowledge Implanting of Large Language Models
- Efficient Continual Learning in Keyword Spotting using Binary Neural Networks
- Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors
- Continuous Learning for Children's ASR: Overcoming Catastrophic Forgetting with Elastic Weight Consolidation and Synaptic Intelligence
- Dynamic Dual Buffer with Divide-and-Conquer Strategy for Online Continual Learning
- What is the role of memorization in Continual Learning?
- HERO: Heterogeneous Continual Graph Learning via Meta-Knowledge Distillation
- Improving Generalization in Heterogeneous Federated Continual Learning via Spatio-Temporal Gradient Matching with Prototypical Coreset
- Gated Integration of Low-Rank Adaptation for Continual Learning of Large Language Models
- Exploiting Age of Information in Network Digital Twins for AI-driven Real-Time Link Blockage Detection
- AdaHAT: Adaptive Hard Attention to the Task in Task-Incremental Learning
- A Unified Gradient-based Framework for Task-agnostic Continual Learning-Unlearning
- Toward Embodied AGI: A Review of Embodied AI and the Road Ahead
- Contrastive Consolidation of Top-Down Modulations Achieves Sparsely Supervised Continual Learning
- CCD: Continual Consistency Diffusion for Lifelong Generative Modeling
- Continuous Subspace Optimization for Continual Learning
- AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning for Small-footprint Keyword Spotting
- Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning
- SEAL: Searching Expandable Architectures for Incremental Learning
- Advancing Multiple Instance Learning with Continual Learning for Whole Slide Imaging
- Task-Core Memory Management and Consolidation for Long-term Continual Learning
- Reinforced Interactive Continual Learning via Real-time Noisy Human Feedback
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
- PrePrompt: Predictive prompting for class incremental learning
- GradMix: Gradient-based Selective Mixup for Robust Data Augmentation in Class-Incremental Learning
- Attention-based Generative Latent Replay: A Continual Learning Approach for WSI Analysis
- TaskVAE: Task-Specific Variational Autoencoders for Exemplar Generation in Continual Learning for Human Activity Recognition
- Fine-Tuning Without Forgetting: Adaptation of YOLOv8 Preserves COCO Performance
- AHC: Meta-Learned Adaptive Compression for Continual Object Detection on Memory-Constrained Microcontrollers
- In-Context Collapse in Vision-Language Models and How to Mitigate it?
- MedCL-Bench: Benchmarking stability-efficiency trade-offs and scaling in biomedical continual learning
- The Gentle Collapse: Distributional Metrics for Continual Learning
- Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
- Toward Interpretable and Generalizable AI in Regulatory Genomics
- Panini: Continual Learning in Token Space via Structured Memory
- Balanced Online Class-Incremental Learning via Dual Classifiers
- Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning
- TailLoR: Protecting Principal Components in Parameter-Efficient Continual Learning
- MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
- Anchor-Aided Multi-User Semantic Communication with Adaptive Decoders
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- Closed-Form Optimal Stepsizes for Cyclic Continual Learning: The Complete d = 2 Theory and the Algebraic Boundary at K = 3
- An active-learning framework for real-time depth perception from monocular vision streams
- Optimal Training-Time Scaling in Gradual Adaptation
- Adaptive Intrusion Detection System using Transformer-Based Neural Networks and Continual Learning Approach with Adversarial Investigation
- CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental Segmentation
- Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute
- Noise-Tolerant Coreset-Based Class Incremental Continual Learning
- SPECI: Skill Prompts based Hierarchical Continual Imitation Learning for Robot Manipulation
- Evaluating Temporal Plasticity in Foundation Time Series Models for Incremental Fine-tuning
- SwitchMT: An Adaptive Context Switching Methodology for Scalable Multi-Task Learning in Intelligent Autonomous Agents
- Continual Learning Strategies for 3D Engineering Regression Problems: A Benchmarking Study
- CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
- Offline Learning and Forgetting for Reasoning with Large Language Models
- Self-Controlled Dynamic Expansion Model for Continual Learning
- Continual learning for rotating machinery fault diagnosis with cross-domain environmental and operational variations
- GPS: Distilling Compact Memories via Grid-based Patch Sampling for Efficient Online Class-Incremental Learning
- Task-conditioned Ensemble of Expert Models for Continuous Learning
- Proxy-Anchor and EVT-Driven Continual Learning Method for Generalized Category Discovery
- Explainability and Continual Learning meet Federated Learning at the Network Edge
- Boosting-inspired online learning with transfer for railway maintenance
- CMIP-CIL: A Cross-Modal Benchmark for Image-Point Class Incremental Learning
- Boosting the Class-Incremental Learning in 3D Point Clouds via Zero-Collection-Cost Basic Shape Pre-Training
- Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
- LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
- SEE: Continual Fine-tuning with Sequential Ensemble of Experts
Related