InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
2025/10/15 by Wenwen Tong, Tong, Wenwen, Dongchuan Ran +46 · 4 citations
Arts and Humanities · Computer Science · Psychology · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Language, Metaphor, and Cognition #Speech and dialogue systems #Subtitles and Audiovisual Media
paper · pdf · doi:10.48550/arxiv.2510.13747
openalex publication_date 2025/10/15 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/28
Abstract
We introduce InteractiveOmni, a unified and open-source omni-modal large language model for audio-visual multi-turn interaction, ranging from 4B to 8B parameters, designed to lead the field of lightweight models by offering comprehensive omni-modal understanding and speech generation capabilities. To achieve this, we integrate the vision encoder, audio encoder, large language model, and speech decoder into a unified model for understanding and generation tasks. We design a multi-stage training strategy to ensure robust cross-modal capabilities, including pre-training for omni-modal understanding, followed by post-training with speech conversation and audio-visual interaction. To enable human-like long-term conversational ability, we meticulously curate a multi-turn training dataset that enhances the model's ability to handle complex and multi-turn interactions. To effectively evaluate the multi-turn memory and speech interaction capabilities, we construct the multi-modal multi-turn memory benchmark and the multi-turn speech interaction benchmark. Experiments demonstrate that InteractiveOmni significantly outperforms leading open-source models and provides a more intelligent multi-turn audio-visual experience, particularly in its long-term memory capabilities. Notably, InteractiveOmni-4B is comparable to the much larger model like Qwen2.5-Omni-7B on general benchmarks, and it can retain 97% of the performance of the InteractiveOmni-8B while utilizing only 50% of the model size. Achieving state-of-the-art results against similarly sized models across image, audio, video understanding, and speech generation tasks, InteractiveOmni is an accessible, open-source foundation for next-generation intelligent interactive systems.
Citations
- VibeVoice Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- MiDashengLM: Efficient Audio Understanding with General Audio Captions
- Marco-Voice Technical Report
- Step-Audio 2 Technical Report
- NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
- Model Merging in Pre-training of Large Language Models
- Qwen3 Technical Report
- Seed1.5-VL Technical Report
- LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
- Kimi-Audio Technical Report
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- Kimi-VL Technical Report
- Qwen2.5-Omni Technical Report
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
- Qwen2.5-VL Technical Report
- Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
- MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
- Baichuan-Omni-1.5 Technical Report
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
- STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
- Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
- Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
- MIO: A Foundation Model on Multimodal Tokens
- E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
- Moshi: a speech-text foundation model for real-time dialogue
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- VITA: Towards Open-Source Interactive Omni Multimodal LLM
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- Qwen2-Audio Technical Report
- MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
- Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
- VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
- ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- MathWriting+
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Mustango: Toward Controllable Text-to-Music Generation
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Joint Audio and Speech Understanding
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
- Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
- Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
- Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and Captioning
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
- MMBench: Is Your Multi-modal Model an All-around Player?
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
- Perception Test: A Diagnostic Benchmark for Multimodal Video Models
- Textually Pretrained Speech Language Models
- Better speech synthesis through scaling
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- Hierarchical Video-Moment Retrieval and Step-Captioning
- GPT-4 Technical Report
- Epic-Sounds: A Large-scale Dataset of Actions That Sound
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Robust Speech Recognition via Large-Scale Weak Supervision
- UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
- MapQA: A Dataset for Question Answering on Choropleth Maps
- CochlScene: Acquisition of acoustic scene data using crowdsourcing
- EgoTaskQA: Understanding Human Tasks in Egocentric Videos
- Audio Retrieval with WavText5K and CLAP Training
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- SPot-the-Difference Self-Supervised Pre-training for Anomaly Detection and Segmentation
- A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
- FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
- Visual Spatial Reasoning
- Flamingo: a Visual Language Model for Few-Shot Learning
- Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
- GigaST: A 10,000-hour Pseudo Speech Translation Corpus
- IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
- WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
- Conditional Variational Autoencoder with Adversarial Learning for\n End-to-End Text-to-Speech
- GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
- Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
- InfographicVQA
- SPGISpeech: 5,000 hours of transcribed financial audio for fully\n formatted end-to-end speech recognition
- Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
- Learning Transferable Visual Models From Natural Language Supervision
- SLURP: A Spoken Language Understanding Resource Package
- A Curated Dataset of Urban Scenes for Audio-Visual Scene Analysis
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines
- CoVoST 2 and Massively Multilingual Speech-to-Text Translation
- DocVQA: A Dataset for VQA on Document Images
- Language Models are Few-Shot Learners
- VGGSound: A Large-scale Audio-Visual Dataset
- CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus
- Common Voice: A Massively-Multilingual Speech Corpus
- Clotho: An Audio Captioning Dataset
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- ICDAR 2019 Competition on Large-scale Street View Text with Partial Labeling -- RRC-LSVT
- ICDAR2019 Robust Reading Challenge on Arbitrary-Shaped Text (RRC-ArT)
- ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard
- Scene Text Visual Question Answering
- OK-VQA: A Visual Question Answering Benchmark Requiring External\n Knowledge
- Towards VQA Models That Can Read
- LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
- TallyQA: Answering Complex Counting Questions
- MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in\n Conversations
- TVQA: Localized, Compositional Video Question Answering
- AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
- General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline
- FigureQA: An Annotated Figure Dataset for Visual Reasoning
- AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
- ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for\n Reading Comprehension
- FMA: A Dataset For Music Analysis
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- Visual Storytelling
- A Diagram Is Worth A Dozen Images
- COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images
- Visual7W: Grounded Question Answering in Images
- A Dataset for Movie Description
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Cited by
Related