NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
2025/10/15 by Run Luo, Luo, Run, Xiaobo Xia +13 · 3 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2510.13721
Abstract
Next-generation multimodal foundation models capable of any-to-any cross-modal generation and multi-turn interaction will serve as core components of artificial general intelligence systems, playing a pivotal role in human-machine interaction. However, most existing multimodal models remain constrained by autoregressive architectures, whose inherent limitations prevent a balanced integration of understanding and generation capabilities. Although hybrid and decoupling strategies have been explored to address these tasks within unified frameworks separately, their redundant, non-integrated designs limit their applicability to broader scenarios, such as cross-modal retrieval. In this work, we introduce NExT-OMNI, an open-source omnimodal foundation model that achieves unified modeling through discrete flow paradigms. By leveraging metric-induced probability paths and kinetic optimal velocities, NExT-OMNI natively supports any-to-any understanding and generation with enhanced response efficiency, while enabling broader application scenarios through concise unified representations rather than task-decoupled designs. Trained on large-scale interleaved text, image, video, and audio data, NExT-OMNI delivers competitive performance on multimodal generation and understanding benchmarks, while outperforming prior unified models in multi-turn multimodal interaction and cross-modal retrieval, highlighting its architectural advantages as a next-generation multimodal foundation model. To advance further research, we release training details, data protocols, and open-source both the code and model checkpoints.
Citations
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Dream 7B: Diffusion Large Language Models
- InterSyn: Interleaved Learning for Dynamic Motion Synthesis in the Wild
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
- Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
- MMaDA: Multimodal Large Diffusion Language Models
- Emerging Properties in Unified Multimodal Pretraining
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
- MMGen: Unified Multi-modal Image Generation and Understanding in One Go
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
- Qwen2.5-VL Technical Report
- Large Language Diffusion Models
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- Dual Diffusion for Unified Image Generation and Understanding
- Qwen2.5 Technical Report
- MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
- MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
- Flow Matching with General Discrete Paths: A Kinetic-Optimal Perspective
- AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
- OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
- TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
- MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
- Scaling Diffusion Language Models via Adaptation from Autoregressive Models
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- Emu3: Next-Token Prediction is All You Need
- EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
- Moshi: a speech-text foundation model for real-time dialogue
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models
- MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- VITA: Towards Open-Source Interactive Omni Multimodal LLM
- LLaVA-OneVision: Easy Visual Task Transfer
- ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
- video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
- What If We Recaption Billions of Web Images with LLaMA-3?
- OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
- DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Better & Faster Large Language Models via Multi-token Prediction
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- Generative Multimodal Models are In-Context Learners
- OneLLM: One Framework to Align All Modalities with Language
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- UniIR: Training and Benchmarking Universal Multimodal Information Retrievers
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- Improved Baselines with Visual Instruction Tuning
- Making LLaMA SEE and Draw with SEED Tokenizer
- DreamLLM: Synergistic Multimodal Comprehension and Creation
- NExT-GPT: Any-to-Any Multimodal LLM
- Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- Planting a SEED of Vision in Large Language Model
- MMBench: Is Your Multi-modal Model an All-around Player?
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
- Evaluating Object Hallucination in Large Vision-Language Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- VideoChat: Chat-Centric Video Understanding
- WizardLM: Empowering large pre-trained language models to follow complex instructions
- Visual Instruction Tuning
- Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text
- LLaMA: Open and Efficient Foundation Language Models
- Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
- Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities
- Robust Speech Recognition via Large-Scale Weak Supervision
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Flow Matching for Generative Modeling
- Classifier-Free Diffusion Guidance
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
- WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Style-Based Generator Architecture for Generative Adversarial Networks
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Neural Discrete Representation Learning
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- TGIF: A New Dataset and Benchmark on Animated GIF Description
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- Liquid: Language Models are Scalable and Unified Multi-modal Generators
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
Cited by
Related