2017/06/02 by Michael Neumann, Ngoc Thang Vu, Neumann, Michael +1 · 1 citation
Computer Science · Psychology · #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and Audio Processing
paper · pdf · doi:10.48550/arxiv.1706.00612
openalex publication_date 2017/06/02 · openalex created_date 2022/10/05 · openalex updated_date 2026/07/28
Speech emotion recognition is an important and challenging task in the realm\nof human-computer interaction. Prior work proposed a variety of models and\nfeature sets for training a system. In this work, we conduct extensive\nexperiments using an attentive convolutional neural network with multi-view\nlearning objective function. We compare system performance using different\nlengths of the input signal, different types of acoustic features and different\ntypes of emotion speech (improvised/scripted). Our experimental results on the\nInteractive Emotional Motion Capture (IEMOCAP) database reveal that the\nrecognition performance strongly depends on the type of speech data independent\nof the choice of input features. Furthermore, we achieved state-of-the-art\nresults on the improvised speech data of IEMOCAP.\n