Gradient Episodic Memory for Continual Learning
2017/06/26 by David López-Paz, Lopez-Paz, David, Marc’Aurelio Ranzato +1 · 139 citations
Computer Science · #Artificial Intelligence (cs.AI) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1706.08840
openalex publication_date 2017/06/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
One major obstacle towards AI is the poor ability of models to solve new problems quicker, and without forgetting previously acquired knowledge. To better understand this issue, we study the problem of continual learning, where the model observes, once and one by one, examples concerning a sequence of tasks. First, we propose a set of metrics to evaluate models learning over a continuum of data. These metrics characterize models not only by their test accuracy, but also in terms of their ability to transfer knowledge across tasks. Second, we propose a model for continual learning, called Gradient Episodic Memory (GEM) that alleviates forgetting, while allowing beneficial transfer of knowledge to previous tasks. Our experiments on variants of the MNIST and CIFAR-100 datasets demonstrate the strong performance of GEM when compared to the state-of-the-art.
Cited by
- Merge before Forget: A Single LoRA Continual Learning via Continual Merging
- SLAM: Structured and Localized Analytic Manifold Adaptation for Forgetting-Immune and Domain-Robust Lifelong VPR
- Dynamic Feedback Engines: Layer-Wise Control for Self-Regulating Continual Learning
- Mixture-of-Experts with Gradient Conflict-Driven Subspace Topology Pruning for Emergent Modularity
- Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios
- InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
- Demonstration-Guided Continual Reinforcement Learning in Dynamic Environments
- AL-GNN: Privacy-Preserving and Replay-Free Continual Graph Learning via Analytic Learning
- Sophia: A Persistent Agent Framework of Artificial Life
- Sequencing to Mitigate Catastrophic Forgetting in Continual Learning
- Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed Memory
- Out-of-Distribution Detection for Continual Learning: Design Principles and Benchmarking
- PPSEBM: An Energy-Based Model with Progressive Parameter Selection for Continual Learning
- Feature Aggregation for Efficient Continual Learning of Complex Facial Expressions
- On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
- Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
- Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
- Multi-Generator Continual Learning for Robust Delay Prediction in 6G
- Heads collapse, features stay: Why Replay needs big buffers
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Bring Your Dreams to Life: Continual Text-to-Video Customization
- Parameter-Efficient Augment Plugin for Class-Incremental Learning
- Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
- MoB: Mixture of Bidders
- Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent
- Stable-Drift: A Patient-Aware Latent Drift Replay Method for Stabilizing Representations in Continual Learning
- Bandit Guided Submodular Curriculum for Adaptive Subset Selection
- Probabilistic Hash Embeddings for Online Learning of Categorical Features
- Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
- Prototype-Guided Non-Exemplar Continual Learning for Cross-subject EEG Decoding
- OpenCML: End-to-End Framework of Open-world Machine Learning to Learn Unknown Classes Incrementally
- Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
- Continual Alignment for SAM: Rethinking Foundation Models for Medical Image Segmentation in Continual Learning
- MedPEFT-CL: Dual-Phase Parameter-Efficient Continual Learning with Medical Semantic Adapter and Bidirectional Memory Consolidation
- Continual Reinforcement Learning for Cyber-Physical Systems: Lessons Learned and Open Challenges
- Parameter Importance-Driven Continual Learning for Foundation Models
- Neuro-Logic Lifelong Learning
- ConSurv: Multimodal Continual Learning for Survival Analysis
- Compact Memory for Continual Logistic Regression
- Learning with Preserving for Continual Multitask Learning
- Mitigating Negative Flips via Margin Preserving Training
- Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective
- Solving bilevel optimization via sequential minimax optimization
- Sharing the Learned Knowledge-base to Estimate Convolutional Filter Parameters for Continual Image Restoration
- The brain as a blueprint: a survey of brain-inspired approaches to learning in artificial intelligence
- Clustering-Based Weight Orthogonalization for Stabilizing Deep Reinforcement Learning
- Model Inversion with Layer-Specific Modeling and Alignment for Data-Free Continual Learning
- The Art of Not Forgetting A Local Learning Architecture for Continual Learning
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- On the importance of cross-task features for class-incremental learning
- Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
- Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?
- Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
- Knowledge-guided Continual Learning for Behavioral Analytics Systems
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Continual Knowledge Consolidation LORA for Domain Incremental Learning
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- Online Time Series Forecasting with Theoretical Guarantees
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- Mapping Post-Training Forgetting in Language Models at Scale
- CaMiT: A Time-Aware Car Model Dataset for Classification and Generation
- The Three Regimes of Offline-to-Online Reinforcement Learning
- Towards Error Centric Intelligence I, Beyond Observational Learning
- BI-MAML: Balanced Incremental Approach for Meta Learning
- FedGTEA: Federated Class-Incremental Learning with Gaussian Task Embedding and Alignment
- REPAIR: Robust Editing via Progressive Adaptive Intervention and Reintegration
- VAO: Validation-Aligned Optimization for Cross-Task Generative Auto-Bidding
- Source-Free Cross-Domain Continual Learning
- End-to-End Test-Time Training for Long Context
- The Unreasonable Effectiveness of Randomized Representations in Online Continual Graph Learning
- IMLP: An Energy-Efficient Continual Learning Method for Tabular Data Streams
- Rehearsal-free and Task-free Online Continual Learning With Contrastive Prompt
- Overcoming Catastrophic Forgetting by Soft Parameter Pruning
- A Computational Perspective on NeuroAI and Synthetic Biological Intelligence
- Temporal Generalization: A Reality Check
- The Lie of the Average: How Class Incremental Learning Evaluation Deceives You?
- Blockwise Hadamard high-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
- LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
- EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email Defense
- DATS: Distance-Aware Temperature Scaling for Calibrated Class-Incremental Learning
- Adaptive Model Ensemble for Continual Learning
- C2Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning
- Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
- DevFD: Developmental Face Forgery Detection by Learning Shared and Orthogonal LoRA Subspaces
- AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
- SPICED: A Synaptic Homeostasis-Inspired Framework for Unsupervised Continual EEG Decoding
- VidCLearn: A Continual Learning Approach for Text-to-Video Generation
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
- Incremental Multistep Forecasting of Battery Degradation Using Pseudo Targets
- AFT: An Exemplar-Free Class Incremental Learning Method for Environmental Sound Classification
- Global Pre-fixing, Local Adjusting: A Simple yet Effective Contrastive Strategy for Continual Learning
- CL2GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction
- Forget What's Sensitive, Remember What Matters: Token-Level Differential Privacy in Memory Sculpting for Continual Learning
- Unbiased Online Curvature Approximation for Regularized Graph Continual Learning
- Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function
- Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities
- Bilevel Continual Learning
- Flattening Sharpness for Dynamic Gradient Projection Memory Benefits\n Continual Learning
- Online time series prediction using feature adjustment
- Holographic Knowledge Manifolds: A Novel Pipeline for Continual Learning Without Catastrophic Forgetting in Large Language Models
- Selfless Sequential Learning
- A Metaverse: Taxonomy, Components, Applications, and Open Challenges
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Unsupervised Video Continual Learning via Non-Parametric Deep Embedded Clustering
- Adapting to Change: A Comparison of Continual and Transfer Learning for Modeling Building Thermal Dynamics under Concept Drifts
- Revisiting Deepfake Detection: Chronological Continual Learning and the Limits of Generalization
- Class Incremental Continual Learning with Self-Organizing Maps and Variational Autoencoders Using Synthetic Replay
- High-dimensional Asymptotics of Generalization Performance in Continual Ridge Regression
- C-Flat++: Towards a More Efficient and Powerful Framework for Continual Learning
- Shift Detection and Adaptation for Network Intrusion Detection
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Calibrated Semantic Diffusion: A p-Laplacian Synthesis with Learnable Dissipation, Quantified Constants, and Graph-Aware Calibration
- Representative Task Self-selection for Flexible Clustered Lifelong Learning
- Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
- Learning Wisdom from Errors: Promoting LLM's Continual Relation Learning through Exploiting Error Cases
- SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph Embedding
- Index-Aligned Query Distillation for Transformer-based Incremental Object Detection
- Meta Continual Learning
- Dynamic Mixture-of-Experts for Incremental Graph Learning
- CitySeg: A 3D Open Vocabulary Semantic Segmentation Foundation Model in City-scale Scenarios
- GaussianUpdate: Continual 3D Gaussian Splatting Update for Changing Environments
- OpenHAIV: A Framework Towards Practical Open-World Learning
- Statistical Theory of Multi-stage Newton Iteration Algorithm for Online Continual Learning
- Lifelong Learner: Discovering Versatile Neural Solvers for Vehicle Routing Problems
- Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
- A Study on Regularization-Based Continual Learning Methods for Indic ASR
- Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
- CRAM: Large-scale Video Continual Learning with Bootstrapped Compression
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- Measuring the stability and plasticity of recommender systems
- Prediction-Oriented Subsampling from Data Streams
- Online Continual Graph Learning
- Exploring Stability-Plasticity Trade-offs for Continual Named Entity Recognition
- Learning to Evolve: Bayesian-Guided Continual Knowledge Graph Embedding
- Towards Field-Ready AI-based Malaria Diagnosis: A Continual Learning Approach
- Continual Learning with Synthetic Boundary Experience Blending
- Forgetting of task-specific knowledge in model merging-based continual learning
- RainbowPrompt: Diversity-Enhanced Prompt-Evolving for Continual Learning
- Modular Delta Merging with Orthogonal Constraints: A Scalable Framework for Continual and Reversible Model Composition
Related