vix.ing · top · new · best · stats · spec

Early Joint Learning of Emotion Information Makes MultiModal Model Understand You Better

2024/09/12 by Mengying Ge, Ge, Mengying, Mingyang Li +17 · 2 citations
Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Innovative Teaching and Learning Methods #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2409.18971

openalex publication_date 2024/09/12 · openalex created_date 2024/10/28 · openalex updated_date 2026/07/28

Abstract

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy based on a large language model, where joint training of audio and text is conducted initially. And the joint Audio-Text modal feature will be late-fused with other unimodal features. In order to solve the problems of data insufficiency and class imbalance, We use multiple turns of multi-model voting for data mining. Moreover, to enhance the quality of audio features, we employ speech source separation to preprocess audios. Our model ranks 2nd in both MER2024-SEMI and MER2024-NOISE, validating our method's effectiveness.

Cited by

Related