Multimodal Video Emotion Recognition with Reliable Reasoning Priors
2025/07/29 by Wang, Zhepeng, Zhu, Yingjian, Dong, Guanghao +4
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2508.03722
Abstract
This study investigates the integration of trustworthy prior reasoning knowledge from MLLMs into multimodal emotion recognition. We employ Gemini to generate fine-grained, modality-separable reasoning traces, which are injected as priors during the fusion stage to enrich cross-modal interactions. To mitigate the pronounced class-imbalance in multimodal emotion recognition, we introduce Balanced Dual-Contrastive Learning, a loss formulation that jointly balances inter-class and intra-class distributions. Applied to the MER2024 benchmark, our prior-enhanced framework yields substantial performance gains, demonstrating that the reliability of MLLM-derived reasoning can be synergistically combined with the domain adaptability of lightweight fusion networks for robust, scalable emotion recognition.
Citations
- TACFN: Transformer-based Adaptive Cross-modal Fusion Network for Multimodal Emotion Recognition
- Visual-RFT: Visual Reinforcement Fine-Tuning
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
- Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
- Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Improving Visual Commonsense in Language Models via Multiple Image Generation
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
- MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning
- Visual Programming: Compositional visual reasoning without training
- Balanced Contrastive Learning for Long-Tailed Visual Recognition
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Representation Learning with Contrastive Predictive Coding
Related