2019/04/04 by Chaitanya Narisetty, Narisetty, Chaitanya, Tatsuya Komatsu +3
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1904.02852
openalex publication_date 2019/04/04 · openalex created_date 2022/07/29 · openalex updated_date 2026/07/28
This paper proposes an effective modelling of sound event spectra with a\nhidden data-size-imbalance, for improved Acoustic Event Detection (AED). The\nproposed method models each event as an aggregated representation of a few\nlatent factors, while conventional approaches try to find acoustic elements\ndirectly from the event spectra. In the method, all the latent factors across\nall events are assigned comparable importance and complexity to overcome the\nhidden imbalance of data-sizes in event spectra. To extract latent factors in\neach event, the proposed method employs clustering and performs non-negative\nmatrix factorization to each latent factor, and learns its acoustic elements as\na sub-dictionary. Separate sub-dictionary learning effectively models the\nacoustic elements with limited data-sizes and avoids over-fitting due to hidden\nimbalances in training data. For the task of polyphonic sound event detection\nfrom DCASE 2013 challenge, an AED based on the proposed modelling achieves a\ndetection F-measure of 46.5%, a significant improvement of more than 19% as\ncompared to the existing state-of-the-art methods.\n