2016/07/13 by Yong Xu, Qiang Huang, Xu, Yong +5
Computer Science · #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Pattern recognition (psychology) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech recognition #cs.CV #cs.LG #cs.SD
paper · pdf · doi:10.48550/arxiv.1607.03682
5 pages, DCASE 2016 challenge workshop paper, poster
openalex publication_date 2016/07/13 · arxiv created 2016/08/13 · arxiv updated 2016/08/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we present a deep neural network (DNN)-based acoustic scene classification framework. Two hierarchical learning methods are proposed to improve the DNN baseline performance by incorporating the hierarchical taxonomy information of environmental sounds. Firstly, the parameters of the DNN are initialized by the proposed hierarchical pre-training. Multi-level objective function is then adopted to add more constraint on the cross-entropy based loss function. A series of experiments were conducted on the Task1 of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2016 challenge. The final DNN-based system achieved a 22.9% relative improvement on average scene classification error as compared with the Gaussian Mixture Model (GMM)-based benchmark system across four standard folds.