vix.ing · top · new · best · stats · spec

MMER: Multimodal Multi-task Learning for Speech Emotion Recognition

2022/03/31 by Sreyan Ghosh, Ghosh, Sreyan, Tyagi, Utkarsh +4 · 3 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2203.16794

openalex publication_date 2022/03/31 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28

Abstract

In this paper, we propose MMER, a novel Multimodal Multi-task learning approach for Speech Emotion Recognition. MMER leverages a novel multimodal network based on early-fusion and cross-modal self-attention between text and acoustic modalities and solves three novel auxiliary tasks for learning emotion recognition from spoken utterances. In practice, MMER outperforms all our baselines and achieves state-of-the-art performance on the IEMOCAP benchmark. Additionally, we conduct extensive ablation studies and results analysis to prove the effectiveness of our proposed approach.

Cited by

Related