vix.ing · top · new · best · stats

Generative replay with feedback connections as a general strategy for continual learning

2018/09/27 by Gido M. van de Ven, Andreas S. Tolias, van de Ven, Gido M. +1 · 184 citations
Computer Science · Engineering · Mathematics · #Artificial Intelligence (cs.AI) #Artificial intelligence #Artificial neural network #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Domain Adaptation and Few-Shot Learning #Engineering #FOS: Computer and information sciences #Forgetting #Generative grammar #Generative model #Human Pose and Action Recognition #MNIST database #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Multimodal Machine Learning Applications #Scalability #Task (project management) #cs.AI #cs.CV #cs.LG #stat.ML

paper · pdf · open access · doi:10.48550/arxiv.1809.10635

published in Lirias · 17 pages, 8 figures, 4 tables

openalex publication_date 2018/09/27 · openalex created_date 2018/10/05 · arxiv created 2019/04/17 · arxiv updated 2019/04/18 · openalex updated_date 2026/07/28

Abstract

A major obstacle to developing artificial intelligence applications capable of true lifelong learning is that artificial neural networks quickly or catastrophically forget previously learned tasks when trained on a new one. Numerous methods for alleviating catastrophic forgetting are currently being proposed, but differences in evaluation protocols make it difficult to directly compare their performance. To enable more meaningful comparisons, here we identified three distinct scenarios for continual learning based on whether task identity is known and, if it is not, whether it needs to be inferred. Performing the split and permuted MNIST task protocols according to each of these scenarios, we found that regularization-based approaches (e.g., elastic weight consolidation) failed when task identity needed to be inferred. In contrast, generative replay combined with distillation (i.e., using class probabilities as "soft targets") achieved superior performance in all three scenarios. Addressing the issue of efficiency, we reduced the computational cost of generative replay by integrating the generative model into the main model by equipping it with generative feedback or backward connections. This Replay-through-Feedback approach substantially shortened training time with no or negligible loss in performance. We believe this to be an important first step towards making the powerful technique of generative replay scalable to real-world continual learning applications.

Citations

Cited by

Related