vix.ing · top · new · best · stats · spec

SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition

2024/08/14 by Mohamed Nasrun Osman, Osman, Mohamed, Daniel Z. Kaplan +3 · 2 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and Audio Processing

paper · pdf · doi:10.48550/arxiv.2408.07851

openalex publication_date 2024/08/14 · openalex created_date 2025/01/03 · openalex updated_date 2026/07/28

Abstract

Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models. However, the generalization of these models to diverse languages and emotional expressions remains a challenge. We propose a large-scale benchmark to evaluate the robustness and adaptability of state-of-the-art SER models in both in-domain and out-of-domain settings. Our benchmark includes a diverse set of multilingual datasets, focusing on less commonly used corpora to assess generalization to new data. We employ logit adjustment to account for varying class distributions and establish a single dataset cluster for systematic evaluation. Surprisingly, we find that the Whisper model, primarily designed for automatic speech recognition, outperforms dedicated SSL models in cross-lingual SER. Our results highlight the need for more robust and generalizable SER models, and our benchmark serves as a valuable resource to drive future research in this direction.

Cited by

Related