vix.ing · top · new · best · stats · spec

Speech Emotion Recognition Using Quaternion Convolutional Neural\n Networks

2021/10/31 by Aneesh Muppidi, Muppidi, Aneesh, Martin Radfar +1 · 1 citation
Computer Science · Psychology · #Speech and Audio Processing #Emotion and Mood Recognition #Hand Gesture Recognition Systems

paper · pdf · doi:10.48550/arxiv.2111.00404

Abstract

Although speech recognition has become a widespread technology, inferring\nemotion from speech signals still remains a challenge. To address this problem,\nthis paper proposes a quaternion convolutional neural network (QCNN) based\nspeech emotion recognition (SER) model in which Mel-spectrogram features of\nspeech signals are encoded in an RGB quaternion domain. We show that our QCNN\nbased SER model outperforms other real-valued methods in the Ryerson\nAudio-Visual Database of Emotional Speech and Song (RAVDESS, 8-classes)\ndataset, achieving, to the best of our knowledge, state-of-the-art results. The\nQCNN also achieves comparable results with the state-of-the-art methods in the\nInteractive Emotional Dyadic Motion Capture (IEMOCAP 4-classes) and Berlin\nEMO-DB (7-classes) datasets. Specifically, the model achieves an accuracy of\n77.87 %, 70.46 %, and 88.78 % for the RAVDESS, IEMOCAP, and EMO-DB datasets,\nrespectively. In addition, our results show that the quaternion unit structure\nis better able to encode internal dependencies to reduce its model size\nsignificantly compared to other methods.\n

Cited by

Related