vix.ing · top · new · best · stats · spec

Sentiment-Aware Automatic Speech Recognition pre-training for enhanced Speech Emotion Recognition

2022/01/27 by Ayoub Ghriss, Ghriss, Ayoub, Bo Yang +7
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Sentiment Analysis and Opinion Mining #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2201.11826

openalex publication_date 2022/01/27 · openalex created_date 2022/07/23 · openalex updated_date 2026/07/28

Abstract

We propose a novel multi-task pre-training method for Speech Emotion Recognition (SER). We pre-train SER model simultaneously on Automatic Speech Recognition (ASR) and sentiment classification tasks to make the acoustic ASR model more ``emotion aware''. We generate targets for the sentiment classification using text-to-sentiment model trained on publicly available data. Finally, we fine-tune the acoustic ASR on emotion annotated speech data. We evaluated the proposed approach on the MSP-Podcast dataset, where we achieved the best reported concordance correlation coefficient (CCC) of 0.41 for valence prediction.

Related