Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
2025/06/01 by Youngmin Kim, Jiwan Chung, Kim, Youngmin +13 · 2 citations
Arts and Humanities · Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Body language #Bridging (networking) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Facial expression #Gesture #Language, Discourse, Communication Strategies #Language, Metaphor, and Cognition #Limiting #Nonverbal communication #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2506.00958
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Nonverbal communication is integral to human interaction, with gestures, facial expressions, and body language conveying critical aspects of intent and emotion. However, existing large language models (LLMs) fail to effectively incorporate these nonverbal elements, limiting their capacity to create fully immersive conversational experiences. We introduce MARS, a multimodal language model designed to understand and generate nonverbal cues alongside text, bridging this gap in conversational AI. Our key innovation is VENUS, a large-scale dataset comprising annotated videos with time-aligned text, facial expressions, and body language. Leveraging VENUS, we train MARS with a next-token prediction objective, combining text with vector-quantized nonverbal representations to achieve multimodal understanding and generation within a unified framework. Based on various analyses of the VENUS datasets, we validate its substantial scale and high effectiveness. Our quantitative and qualitative results demonstrate that MARS successfully generates text and nonverbal languages, corresponding to conversational input.
Citations
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
- TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- MoMask: Generative Masked Modeling of 3D Human Motions
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- HumanTOMATO: Text-aligned Whole-body Motion Generation
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Can Language Models Learn to Listen?
- Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
- Audio-Driven 3D Facial Animation from In-the-Wild Videos
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- Learning Emotion Representations from Verbal and Nonverbal Communication
- One-Stage 3D Whole-Body Mesh Recovery with Component Aware Transformer
- CHAMPAGNE: Learning Real-world Conversation from Large-Scale Web Videos
- GPT-4 Technical Report
- A Light Weight Model for Active Speaker Detection
- WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
- T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations
- Generating Holistic 3D Human Motion from Speech
- TweetNLP: Cutting-Edge Natural Language Processing for Social Media
- EMOCA: Emotion Driven Monocular Face Capture and Animation
- Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion
- MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound
- MERLOT: Multimodal Neural Script Knowledge Models
- pyannote.audio: neural building blocks for speaker diarization
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- BERTScore: Evaluating Text Generation with BERT
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
- MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
- Neural Discrete Representation Learning
- Attention Is All You Need
Cited by
Related