vix.ing · top · new · best · stats · spec

Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model

2025/09/10 by Mana Ihori, Taiga Yamane, Ihori, Mana +13
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Electrical engineering #Sentiment Analysis and Opinion Mining #Speech Recognition and Synthesis #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2509.08344

openalex publication_date 2025/09/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper proposes a personalization method for speech emotion recognition (SER) through in-context learning (ICL). Since the expression of emotions varies from person to person, speaker-specific adaptation is crucial for improving the SER performance. Conventional SER methods have been personalized using emotional utterances of a target speaker, but it is often difficult to prepare utterances corresponding to all emotion labels in advance. Our idea to overcome this difficulty is to obtain speaker characteristics by conditioning a few emotional utterances of the target speaker in ICL-based inference. ICL is a method to perform unseen tasks by conditioning a few input-output examples through inference in large language models (LLMs). We meta-train a speech-language model extended from the LLM to learn how to perform personalized SER via ICL. Experimental results using our newly collected SER dataset demonstrate that the proposed method outperforms conventional methods.

Citations

Related