OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
2026/08/06 by Jiahao Huang, Zheng Lian, Jingyi Zhang +3
Computer Science · #cs.HC
paper · pdf
arxiv created 2026/08/06 · arxiv updated 2026/08/07
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.
Citations
- Cosmos 3: Omnimodal World Models for Physical AI
- Qwen3-VL Technical Report
- Emotional Artificial Intelligence in Education: A Systematic Review and Meta-Analysis
- Facial-R1: Aligning Reasoning and Recognition for Facial Emotion Analysis
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
- Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning
- AffectGPT-R1: Leveraging Reinforcement Learning for Open-Vocabulary Multimodal Emotion Recognition
- Beyond Empathy: Integrating Diagnostic and Therapeutic Reasoning with Large Language Models for Mental Health Counseling
- MER 2025: When Affective Computing Meets Large Language Models
- R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
- Qwen2.5-VL Technical Report
- Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
- Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
- OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
- ECR-Chain: Advancing Generative Language Models to Better Emotion-Cause Reasoners through Reasoning Chains
- MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
- MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- MERBench: A Unified Evaluation Benchmark for Multimodal Emotion Recognition
- MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning
- Make Acoustic and Visual Cues Matter: CH-SIMS v2.0 Dataset and AV-Mixup Consistent Module
- DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the Wild
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Towards Multimodal Sarcasm Detection (An Obviously_ Perfect Paper)
- UR-FUNNY: A Multimodal Language Dataset for Understanding Humor
- MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
- MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos
- An argument for basic emotions
- Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding