Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC
2025/08/06 by Hu, Guanyu, Kollias, Dimitrios, Yang, Xinyu · 2 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2508.06564
Abstract
Multimodal Emotion Recognition in Conversations remains a challenging task due to the complex interplay of textual, acoustic and visual signals. While recent models have improved performance via advanced fusion strategies, they often lack psychologically meaningful priors to guide multimodal alignment. In this paper, we revisit the use of CLIP and propose a novel Visual Emotion Guided Anchoring (VEGA) mechanism that introduces class-level visual semantics into the fusion and classification process. Distinct from prior work that primarily utilizes CLIP's textual encoder, our approach leverages its image encoder to construct emotion-specific visual anchors based on facial exemplars. These anchors guide unimodal and multimodal features toward a perceptually grounded and psychologically aligned representation space, drawing inspiration from cognitive theories (prototypical emotion categories and multisensory integration). A stochastic anchor sampling strategy further enhances robustness by balancing semantic stability and intra-class diversity. Integrated into a dual-branch architecture with self-distillation, our VEGA-augmented model achieves sota performance on IEMOCAP and MELD. Code is available at: https://github.com/dkollias/VEGA.
Citations
- Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
- Rethinking Affect Analysis: A Protocol for Ensuring Fairness and Consistency
- Robust Facial Reactions Generation: An Emotion-Aware Framework with Modality Compensation
- 7th ABAW Competition: Multi-Task Learning and Compound Expression Recognition
- Ensuring UAV Safety: A Vision-only and Real-time Framework for Collision Avoidance Through Object Detection, Tracking, and Distance Estimation
- CUE-Net: Violence Detection Video Analytics with Spatial Cropping, Enhanced UniformerV2 and Modified Efficient Additive Attention
- COVID-19 Computer-aided Diagnosis through AI-assisted CT Imaging Analysis: Deploying a Medical AI System
- The 6th Affective Behavior Analysis in-the-wild (ABAW) Competition
- Distribution Matching for Multi-Task Learning of Classification Tasks: a Large-Scale Study on Faces & Beyond
- BTDNet: a Multi-Modal Approach for Brain Tumor Radiogenomic Classification
- Multimodal Continuous Emotion Recognition: A Technical Report for ABAW5
- ABAW: Valence-Arousal Estimation, Expression Recognition, Action Unit Detection & Emotional Reaction Intensity Estimation Challenges
- MMA-MRNNet: Harnessing Multiple Models of Affect and Dynamic Masked RNN for Precise Facial Expression Intensity Estimation
- A Deep Neural Architecture for Harmonizing 3-D Input Data Analysis and Decision Making in Medical Imaging
- Deep Emotion Recognition in Textual Conversations: A Survey
- MM-DFN: Multimodal Dynamic Fusion Network for Emotion Recognition in Conversations
- ABAW: Valence-Arousal Estimation, Expression Recognition, Action Unit Detection & Multi-Task Learning Challenges
- Analysing Affective Behavior in the second ABAW2 Competition
- MIA-COV19D: COVID-19 Detection through 3-D Chest CT Image Analysis
- Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study
- Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework
- Learning Transferable Visual Models From Natural Language Supervision
- DialogueTRM: Exploring the Intra- and Inter-Modal Emotional Behaviors in the Conversation
- Deep Transparent Prediction through Latent Representation Analysis
- Analysing Affective Behavior in the First ABAW 2020 Competition
- Face Behavior a la carte: Expressions, Affect and Action Units in a Single Network
- Exploiting multi-CNN features in CNN-RNN based Dimensional Emotion\n Recognition on the OMG in-the-wild Dataset
- Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace
- DialogueGCN: A Graph Convolutional Neural Network for Emotion\n Recognition in Conversation
- Emotion Recognition in Conversation: Research Challenges, Datasets, and Recent Advances
- A Multi-Task Learning & Generation Framework: Valence-Arousal, Action Units & Primary Expressions
- Aff-Wild2: Extending the Aff-Wild Database for Affect Recognition
- DialogueRNN: An Attentive RNN for Emotion Detection in Conversations
- Multimodal Speech Emotion Recognition Using Audio and Text
- MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in\n Conversations
- Photorealistic Facial Synthesis in the Dimensional Affect Space
- Training Deep Neural Networks with Different Datasets In-the-wild: The Emotion Recognition Paradigm
- A Multi-component CNN-RNN Approach for Dimensional Emotion Recognition in-the-wild
- Memory Fusion Network for Multi-view Sequential Learning
- An argument for basic emotions
- Grounded Cognition
Cited by
Related