vix.ing · top · new · best · stats

Multi-Branch Learning for Weakly-Labeled Sound Event Detection

2020/02/22 by Yuxin Huang, Xiangdong Wang, Huang, Yuxin +7
Computer Science · Engineering · #Artificial intelligence #Artificial neural network #Audio and Speech Processing (eess.AS) #Computer science #Event (particle physics) #FOS: Electrical engineering #Feature (linguistics) #Feature learning #Machine learning #Multi-task learning #Music and Audio Processing #Overfitting #Pattern recognition (psychology) #Pooling #Speech and Audio Processing #Speech recognition #Task (project management) #Water Systems and Optimization #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2002.09661

published in arXiv (Cornell University) (Cornell University) · Accepted by ICASSP 2020

arxiv created 2020/02/22 · openalex publication_date 2020/02/22 · arxiv updated 2020/02/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

There are two sub-tasks implied in the weakly-supervised SED: audio tagging and event boundary detection. Current methods which combine multi-task learning with SED requires annotations both for these two sub-tasks. Since there are only annotations for audio tagging available in weakly-supervised SED, we design multiple branches with different learning purposes instead of pursuing multiple tasks. Similar to multiple tasks, multiple different learning purposes can also prevent the common feature which the multiple branches share from overfitting to any one of the learning purposes. We design these multiple different learning purposes based on combinations of different MIL strategies and different pooling methods. Experiments on the DCASE 2018 Task 4 dataset and the URBAN-SED dataset both show that our method achieves competitive performance.

Related