2022/02/01 by Guanglei Yang, Enrico Fini, Yang, Guanglei +13 · 1 citation
Computer Science · Medicine · #Artificial intelligence #COVID-19 diagnosis using AI #Cognitive psychology #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Deep learning #Deep neural networks #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Forgetting #Image (mathematics) #Interpretability #Machine learning #Multimodal Machine Learning Applications #Pascal (unit) #Popularity #Segmentation #Semantic gap #cs.CV
paper · pdf · open access · doi:10.48550/arxiv.2202.00432
published in arXiv (Cornell University) (Cornell University)
arxiv created 2022/02/01 · openalex publication_date 2022/02/01 · arxiv updated 2022/02/02 · openalex created_date 2022/09/06 · openalex updated_date 2026/07/28
Over the past years, semantic segmentation, as many other tasks in computer vision, benefited from the progress in deep neural networks, resulting in significantly improved performance. However, deep architectures trained with gradient-based techniques suffer from catastrophic forgetting, which is the tendency to forget previously learned knowledge while learning new tasks. Aiming at devising strategies to counteract this effect, incremental learning approaches have gained popularity over the past years. However, the first incremental learning methods for semantic segmentation appeared only recently. While effective, these approaches do not account for a crucial aspect in pixel-level dense prediction problems, i.e. the role of attention mechanisms. To fill this gap, in this paper we introduce a novel attentive feature distillation approach to mitigate catastrophic forgetting while accounting for semantic spatial- and channel-level dependencies. Furthermore, we propose a continual attentive fusion structure, which takes advantage of the attention learned from the new and the old tasks while learning features for the new task. Finally, we also introduce a novel strategy to account for the background class in the distillation loss, thus preventing biased predictions. We demonstrate the effectiveness of our approach with an extensive evaluation on Pascal-VOC 2012 and ADE20K, setting a new state of the art.