2021/10/09 by Danruo Deng, Deng, Danruo, Guangyong Chen +7 · 2 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2110.04593
openalex publication_date 2021/10/09 · openalex created_date 2021/11/22 · openalex updated_date 2026/07/28
The backpropagation networks are notably susceptible to catastrophic\nforgetting, where networks tend to forget previously learned skills upon\nlearning new ones. To address such the 'sensitivity-stability' dilemma, most\nprevious efforts have been contributed to minimizing the empirical risk with\ndifferent parameter regularization terms and episodic memory, but rarely\nexploring the usages of the weight loss landscape. In this paper, we\ninvestigate the relationship between the weight loss landscape and\nsensitivity-stability in the continual learning scenario, based on which, we\npropose a novel method, Flattening Sharpness for Dynamic Gradient Projection\nMemory (FS-DGPM). In particular, we introduce a soft weight to represent the\nimportance of each basis representing past tasks in GPM, which can be\nadaptively learned during the learning process, so that less important bases\ncan be dynamically released to improve the sensitivity of new skill learning.\nWe further introduce Flattening Sharpness (FS) to reduce the generalization gap\nby explicitly regulating the flatness of the weight loss landscape of all seen\ntasks. As demonstrated empirically, our proposed method consistently\noutperforms baselines with the superior ability to learn new skills while\nalleviating forgetting effectively.\n