Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
2025/11/16 by Li, Yunxin, Chen, Xinyu, Jiang, Shenyuan +9 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2511.12609
Abstract
We present Uni-MoE 2.0 from the Lychee family. As a fully open-source omnimodal large model (OLM), it substantially advances Lychee's Uni-MoE series in language-centric multimodal understanding, reasoning, and generating. Based on the dense LLM, we build Uni-MoE-2.0-Omni from scratch through three core contributions: dynamic-capacity Mixture-of-Experts (MoE) design, a progressive training strategy enhanced with an iterative reinforcement strategy, and a carefully curated multimodal data matching technique. It is capable of omnimodal understanding, as well as generating images, text, and speech. Architecturally, our new MoE framework balances computational efficiency and capability for 10 cross-modal inputs using shared, routed, and null experts, while our Omni-Modality 3D RoPE ensures spatio-temporal cross-modality alignment in the self-attention layer. For training, following cross-modal pretraining, we use a progressive supervised fine-tuning strategy that activates modality-specific experts and is enhanced by balanced data composition and an iterative GSPO-DPO method to stabilise RL training and improve reasoning. Data-wise, the base model, trained on approximately 75B tokens of open-source multimodal data, is equipped with special speech and image generation tokens, allowing it to learn these generative tasks by conditioning its outputs on linguistic cues. Extensive evaluation across 85 benchmarks demonstrates that our model achieves SOTA or highly competitive performance against leading OLMs, surpassing Qwen2.5-Omni (trained with 1.2T tokens) on over 50 of 76 benchmarks. Key strengths include video understanding (+7% avg. of 8), omnimodallity understanding (+7% avg. of 4), and audiovisual reasoning (+4%). It also advances long-form speech processing (reducing WER by 4.2%) and leads in low-level image processing and controllable generation across 5 metrics.
Citations
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining
- ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
- OmniGen2: Exploration to Advanced Multimodal Generation
- Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
- Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
- VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization
- Emerging Properties in Unified Multimodal Pretraining
- A Survey on Large Language Models in Multimodal Recommender Systems
- Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
- Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
- Step1X-Edit: A Practical Framework for General Image Editing
- VideoVista-CulturalLingo: 360^∘ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
- SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
- Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
- Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering
- Qwen2.5-Omni Technical Report
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
- MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Baichuan-Omni-1.5 Technical Report
- Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Neptune: The Long Orbit to Benchmarking Long Video Understanding
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
- GPT-4o System Card
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
- WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- LLaVA-OneVision: Easy Visual Task Transfer
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
- UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Long Story Short: Story-level Video Understanding from 20K Short Films
- From Pixels to Prose: A Large Dataset of Dense Image Captions
- VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
- ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
- CinePile: A Long Video Question Answering Dataset and Benchmark
- SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
- HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
- ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- A Survey on Multimodal Large Language Models for Autonomous Driving
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Emu Edit: Precise Image Editing via Recognition and Generation Tasks
- Mustango: Toward Controllable Text-to-Music Generation
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- LP-MusicCaps: LLM-Based Pseudo Music Captioning
- Meta-Transformer: A Unified Framework for Multimodal Learning
- MMBench: Is Your Multi-modal Model an All-around Player?
- JourneyDB: A Benchmark for Generative Image Understanding
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- FunQA: Towards Surprising Video Comprehension
- MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
- Valley: Video Assistant with Large Language model Enhanced abilitY
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
- Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Visual Instruction Tuning
- Sigmoid Loss for Language Image Pre-Training
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Robust Speech Recognition via Large-Scale Weak Supervision
- InstructPix2Pix: Learning to Follow Image Editing Instructions
- EgoTaskQA: Understanding Human Tasks in Egocentric Videos
- A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
- Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
- Learning to Answer Questions in Dynamic Audio-Visual Scenarios
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage
- Emotional Voice Conversion: Theory, Databases and ESD
- Emotional voice conversion: Theory, databases and ESD
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize\n Long-Tail Visual Concepts
- WDNet: Watermark-Decomposition Network for Visible Watermark Removal
- AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines
- DocVQA: A Dataset for VQA on Document Images
- NH-HAZE: An Image Dehazing Benchmark with Non-Homogeneous Hazy and Haze-Free Images
- Defocus Deblurring Using Dual-Pixel Data
- Common Voice: A Massively-Multilingual Speech Corpus
- Clotho: An Audio Captioning Dataset
- Heavy Rain Image Restoration: Integrating Physics Model and Conditional Adversarial Learning
- Dense Haze: A benchmark for image dehazing with dense-haze and haze-free\n images
- LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
- Toward Real-World Single Image Super-Resolution: A New Benchmark and A New Model
- MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in\n Conversations
- The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems
- Benchmarking Single Image Dehazing and Beyond
- Attentive Generative Adversarial Network for Raindrop Removal from a Single Image
- AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
- TALL: Temporal Activity Localization via Language Query
- RACE: Large-scale ReAding Comprehension Dataset From Examinations
- Deep Multi-scale Convolutional Neural Network for Dynamic Scene\n Deblurring
- Deep Joint Rain Detection and Removal from a Single Image
- A Diagram Is Worth A Dozen Images
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for\n Richer Image-to-Sentence Models
- VQA: Visual Question Answering
- Microsoft COCO: Common Objects in Context
- ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Cited by
Related