Growing Visual Generative Capacity for Pre-Trained MLLMs
2025/10/02 by Hanyu Wang, Jiaming Han, Wang, Hanyu +14
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Semantic Web and Ontologies
paper · pdf · doi:10.48550/arxiv.2510.01546
openalex publication_date 2025/10/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Multimodal large language models (MLLMs) extend the success of language models to visual understanding, and recent efforts have sought to build unified MLLMs that support both understanding and generation. However, constructing such models remains challenging: hybrid approaches combine continuous embeddings with diffusion or flow-based objectives, producing high-quality images but breaking the autoregressive paradigm, while pure autoregressive approaches unify text and image prediction over discrete visual tokens but often face trade-offs between semantic alignment and pixel-level fidelity. In this work, we present Bridge, a pure autoregressive unified MLLM that augments pre-trained visual understanding models with generative ability through a Mixture-of-Transformers architecture, enabling both image understanding and generation within a single next-token prediction framework. To further improve visual generation fidelity, we propose a semantic-to-pixel discrete representation that integrates compact semantic tokens with fine-grained pixel tokens, achieving strong language alignment and precise description of visual details with only a 7.9% increase in sequence length. Extensive experiments across diverse multimodal benchmarks demonstrate that Bridge achieves competitive or superior results in both understanding and generation benchmarks, while requiring less training data and reduced training time compared to prior unified MLLMs.
Citations
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- Qwen-Image Technical Report
- GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- Show-o2: Improved Native Unified Multimodal Models
- VINCIE: Unlocking In-context Image Editing from Video
- Thinking with Generated Images
- ImgEdit: A Unified Image Editing Dataset and Benchmark
- Emerging Properties in Unified Multimodal Pretraining
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- Step1X-Edit: A Practical Framework for General Image Editing
- Seedream 3.0 Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Transfer between Modalities with MetaQueries
- Long Context Tuning for Video Generation
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
- UniTok: A Unified Tokenizer for Visual Generation and Understanding
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
- LMFusion: Adapting Pretrained Language Models for Multimodal Generation
- MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
- ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
- Liquid: Language Models are Scalable and Unified Multi-modal Generators
- Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
- AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
- JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
- OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
- Randomized Autoregressive Visual Generation
- LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
- Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
- HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
- Emu3: Next-Token Prediction is All You Need
- OmniGen: Unified Image Generation
- VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- LLaVA-OneVision: Easy Visual Task Transfer
- UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
- OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- Evaluating Text-to-Visual Generation with Image-to-Text Generation
- ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
- World Model on Million-Length Video And Language With Blockwise RingAttention
- Learned representation-guided diffusion models for large-image generation
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- CapsFusion: Rethinking Image-Text Data at Scale
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- Improved Baselines with Visual Instruction Tuning
- Qwen Technical Report
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MMBench: Is Your Multi-modal Model an All-around Player?
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- JourneyDB: A Benchmark for Generative Image Understanding
- Conditional Text Image Generation with Diffusion Models
- Evaluating Object Hallucination in Large Vision-Language Models
- Visual Instruction Tuning
- LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
- Sigmoid Loss for Language Image Pre-Training
- Unified Multi-Modal Latent Diffusion for Joint Subject and Text Conditional Image Generation
- LLaMA: Open and Efficient Foundation Language Models
- Scalable Diffusion Models with Transformers
- InstructPix2Pix: Learning to Follow Image Editing Instructions
- LAION-5B: An open large-scale dataset for training next generation image-text models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Conditional Image Generation with Score-Based Diffusion Models
- Cascaded Diffusion Models for High Fidelity Image Generation
- Learning Transferable Visual Models From Natural Language Supervision
- Taming Transformers for High-Resolution Image Synthesis
- Language Models are Few-Shot Learners
- Neural Discrete Representation Learning
Related