2018/01/29 by Sharath Adavanne, Adavanne, Sharath, Archontis Politis +3
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1801.09522
openalex publication_date 2018/01/29 · openalex created_date 2021/11/08 · openalex updated_date 2026/07/28
In this paper, we propose a stacked convolutional and recurrent neural\nnetwork (CRNN) with a 3D convolutional neural network (CNN) in the first layer\nfor the multichannel sound event detection (SED) task. The 3D CNN enables the\nnetwork to simultaneously learn the inter- and intra-channel features from the\ninput multichannel audio. In order to evaluate the proposed method,\nmultichannel audio datasets with different number of overlapping sound sources\nare synthesized. Each of this dataset has a four-channel first-order Ambisonic,\nbinaural, and single-channel versions, on which the performance of SED using\nthe proposed method are compared to study the potential of SED using\nmultichannel audio. A similar study is also done with the binaural and\nsingle-channel versions of the real-life recording TUT-SED 2017 development\ndataset. The proposed method learns to recognize overlapping sound events from\nmultichannel features faster and performs better SED with a fewer number of\ntraining epochs. The results show that on using multichannel Ambisonic audio in\nplace of single-channel audio we improve the overall F-score by 7.5%, overall\nerror rate by 10% and recognize 15.6% more sound events in time frames with\nfour overlapping sound sources.\n