vix.ing · top · new · best · stats · spec

EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast

2025/05/29 by Shreeram Suresh Chandra, Lucas Goncalves, Chandra, Shreeram Suresh +7 · 3 citations
Computer Science · Psychology · #Emotion and Mood Recognition #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Sentiment Analysis and Opinion Mining

paper · pdf · doi:10.48550/arxiv.2505.23732

openalex publication_date 2025/05/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Current emotion-based contrastive language-audio pretraining (CLAP) methods typically learn by naïvely aligning audio samples with corresponding text prompts. Consequently, this approach fails to capture the ordinal nature of emotions, hindering inter-emotion understanding and often resulting in a wide modality gap between the audio and text embeddings due to insufficient alignment. To handle these drawbacks, we introduce EmotionRankCLAP, a supervised contrastive learning approach that uses dimensional attributes of emotional speech and natural language prompts to jointly capture fine-grained emotion variations and improve cross-modal alignment. Our approach utilizes a Rank-N-Contrast objective to learn ordered relationships by contrasting samples based on their rankings in the valence-arousal space. EmotionRankCLAP outperforms existing emotion-CLAP methods in modeling emotion ordinality across modalities, measured via a cross-modal retrieval task.

Cited by

Related