vix.ing · top · new · best · stats

EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection

2025/06/11 by Christoph Schuhmann, Robert Kaczmarczyk, Schuhmann, Christoph +15
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Benchmark (surveying) #Computation and Language (cs.CL) #Emotion and Mood Recognition #Emotion classification #Emotion detection #Emotion recognition #FOS: Computer and information sciences #Face (sociological concept) #Mental Health via Writing #Resource (disambiguation) #Sadness #Sentiment Analysis and Opinion Mining

paper · pdf · doi:10.48550/arxiv.2506.09827

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/06/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Speech emotion recognition (SER) systems are constrained by existing datasets that typically cover only 6-10 basic emotions, lack scale and diversity, and face ethical challenges when collecting sensitive emotional states. We introduce EMONET-VOICE, a comprehensive resource addressing these limitations through two components: (1) EmoNet-Voice Big, a 5,000-hour multilingual pre-training dataset spanning 40 fine-grained emotion categories across 11 voices and 4 languages, and (2) EmoNet-Voice Bench, a rigorously validated benchmark of 4,7k samples with unanimous expert consensus on emotion presence and intensity levels. Using state-of-the-art synthetic voice generation, our privacy-preserving approach enables ethical inclusion of sensitive emotions (e.g., pain, shame) while maintaining controlled experimental conditions. Each sample underwent validation by three psychology experts. We demonstrate that our Empathic Insight models trained on our synthetic data achieve strong real-world dataset generalization, as tested on EmoDB and RAVDESS. Furthermore, our comprehensive evaluation reveals that while high-arousal emotions (e.g., anger: 95% accuracy) are readily detected, the benchmark successfully exposes the difficulty of distinguishing perceptually similar emotions (e.g., sadness vs. distress: 63% discrimination), providing quantifiable metrics for advancing nuanced emotion AI. EMONET-VOICE establishes a new paradigm for large-scale, ethically-sourced, fine-grained SER research.

Citations

Related