vix.ing · top · new · best · stats · spec

On the Robustness of Speech Emotion Recognition for Human-Robot\n Interaction with Deep Neural Networks

2018/04/06 by Egor Lakomkin, Mohammad Ali Zamani, Lakomkin, Egor +7 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Robotics (cs.RO) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1804.02173

openalex publication_date 2018/04/06 · openalex created_date 2022/09/25 · openalex updated_date 2026/07/28

Abstract

Speech emotion recognition (SER) is an important aspect of effective\nhuman-robot collaboration and received a lot of attention from the research\ncommunity. For example, many neural network-based architectures were proposed\nrecently and pushed the performance to a new level. However, the applicability\nof such neural SER models trained only on in-domain data to noisy conditions is\ncurrently under-researched. In this work, we evaluate the robustness of\nstate-of-the-art neural acoustic emotion recognition models in human-robot\ninteraction scenarios. We hypothesize that a robot's ego noise, room\nconditions, and various acoustic events that can occur in a home environment\ncan significantly affect the performance of a model. We conduct several\nexperiments on the iCub robot platform and propose several novel ways to reduce\nthe gap between the model's performance during training and testing in\nreal-world conditions. Furthermore, we observe large improvements in the model\nperformance on the robot and demonstrate the necessity of introducing several\ndata augmentation techniques like overlaying background noise and loudness\nvariations to improve the robustness of the neural approaches.\n

Cited by

Related