2018/10/30 by Hao Zhou, Ke Chen, Zhou, Hao +1
Computer Science · Mathematics · Psychology · #Emotion and Mood Recognition #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1810.12782
5 pages, 3 figures, accepted to ICASSP 2019
openalex publication_date 2018/10/30 · openalex created_date 2018/11/09 · arxiv created 2019/02/14 · arxiv updated 2019/02/15 · openalex updated_date 2026/07/28
Speech emotion recognition plays an important role in building more intelligent and human-like agents. Due to the difficulty of collecting speech emotional data, an increasingly popular solution is leveraging a related and rich source corpus to help address the target corpus. However, domain shift between the corpora poses a serious challenge, making domain shift adaptation difficult to function even on the recognition of positive/negative emotions. In this work, we propose class-wise adversarial domain adaptation to address this challenge by reducing the shift for all classes between different corpora. Experiments on the well-known corpora EMODB and Aibo demonstrate that our method is effective even when only a very limited number of target labeled examples are provided.