2020/07/10 by Konstantinos Drossos, Drossos, Konstantinos, Stylianos Ioannis Mimilakis +3
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2007.05183
openalex publication_date 2020/07/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Sound event detection (SED) is the task of identifying sound events along with their onset and offset times. A recent, convolutional neural networks based SED method, proposed the usage of depthwise separable (DWS) and time-dilated convolutions. DWS and time-dilated convolutions yielded state-of-the-art results for SED, with considerable small amount of parameters. In this work we propose the expansion of the time-dilated convolutions, by conditioning them with jointly learned embeddings of the SED predictions by the SED classifier. We present a novel algorithm for the conditioning of the time-dilated convolutions which functions similarly to language modelling, and enhances the performance of the these convolutions. We employ the freely available TUT-SED Synthetic dataset, and we assess the performance of our method using the average per-frame F1 score and average per-frame error rate, over the 10 experiments. We achieve an increase of 2% (from 0.63 to 0.65) at the average F1 score (the higher the better) and a decrease of 3% (from 0.50 to 0.47) at the error rate (the lower the better).