2018/12/04 by Seongkyu Mun, Mun, Seongkyu, Suwon Shon +1
Arts and Humanities · Computer Science · Engineering · Mathematics · #Artificial intelligence #Audio and Speech Processing (eess.AS) #Autoencoder #Channel (broadcasting) #Computer science #Computer vision #Data mining #Deep learning #Diverse Musicological Studies #Domain (mathematical analysis) #Domain adaptation #Engineering #Event (particle physics) #FOS: Computer and information sciences #FOS: Electrical engineering #Field (mathematics) #Mathematics #Music and Audio Processing #Pattern recognition (psychology) #Process (computing) #Sound (cs.SD) #Speech and Audio Processing #Speech recognition #Task (project management) #Telecommunications #Time domain #cs.SD #eess.AS #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1812.01731
published in arXiv (Cornell University) (Cornell University)
arxiv created 2018/12/04 · openalex publication_date 2018/12/04 · arxiv updated 2018/12/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In a recent acoustic scene classification (ASC) research field, training and test device channel mismatch have become an issue for the real world implementation. To address the issue, this paper proposes a channel domain conversion using factorized hierarchical variational autoencoder. Proposed method adapts both the source and target domain to a pre-defined specific domain. Unlike the conventional approach, the relationship between the target and source domain and information of each domain are not required in the adaptation process. Based on the experimental results using the IEEE detection and classification of acoustic scenes and event 2018 task 1-B dataset and the baseline system, it is shown that the proposed approach can mitigate the channel mismatching issue of different recording devices.