2020/05/18 by Siddique Latif, Latif, Siddique, Rajib Rana +7
Computer Science · #Anomaly Detection Techniques and Applications #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2005.08453
openalex publication_date 2020/05/18 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Speech emotion recognition systems (SER) can achieve high accuracy when the\ntraining and test data are identically distributed, but this assumption is\nfrequently violated in practice and the performance of SER systems plummet\nagainst unforeseen data shifts. The design of robust models for accurate SER is\nchallenging, which limits its use in practical applications. In this paper we\npropose a deeper neural network architecture wherein we fuse DenseNet, LSTM and\nHighway Network to learn powerful discriminative features which are robust to\nnoise. We also propose data augmentation with our network architecture to\nfurther improve the robustness. We comprehensively evaluate the architecture\ncoupled with data augmentation against (1) noise, (2) adversarial attacks and\n(3) cross-corpus settings. Our evaluations on the widely used IEMOCAP and\nMSP-IMPROV datasets show promising results when compared with existing studies\nand state-of-the-art models.\n