BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
2025/05/14 by Chen, Jiuhai, Xu, Zhiyang, Pan, Xichen +10 · 123 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2505.09568
Abstract
Unifying image understanding and generation has gained growing attention in recent research on multimodal models. Although design choices for image understanding have been extensively studied, the optimal model architecture and training recipe for a unified framework with image generation remain underexplored. Motivated by the strong potential of autoregressive and diffusion models for high-quality generation and scalability, we conduct a comprehensive study of their use in unified multimodal settings, with emphasis on image representations, modeling objectives, and training strategies. Grounded in these investigations, we introduce a novel approach that employs a diffusion transformer to generate semantically rich CLIP image features, in contrast to conventional VAE-based representations. This design yields both higher training efficiency and improved generative quality. Furthermore, we demonstrate that a sequential pretraining strategy for unified models-first training on image understanding and subsequently on image generation-offers practical advantages by preserving image understanding capability while developing strong image generation ability. Finally, we carefully curate a high-quality instruction-tuning dataset BLIP3o-60k for image generation by prompting GPT-4o with a diverse set of captions covering various scenes, objects, human gestures, and more. Building on our innovative model design, training recipe, and datasets, we develop BLIP3-o, a suite of state-of-the-art unified multimodal models. BLIP3-o achieves superior performance across most of the popular benchmarks spanning both image understanding and generation tasks. To facilitate future research, we fully open-source our models, including code, model weights, training scripts, and pretraining and instruction tuning datasets.
Cited by
- ThinkGen: Generalized Thinking for Visual Generation
- PixelGen: Improving Pixel Diffusion with Perceptual Supervision
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- Loom: Diffusion-Transformer for Interleaved Generation
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- Do-Undo: Generating and Reversing Physical Actions in Vision-Language Models
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- What Happens Next? Next Scene Prediction with a Unified Video Model
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- Mull-Tokens: Modality-Agnostic Latent Thinking
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval
- Accelerating Inference of Masked Image Generators via Reinforcement Learning
- RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
- Visual Generation Tuning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Ovis-Image Technical Report
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
- Plan-X: Instruct Video Generation via Semantic Planning
- UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- Moving Pictures of Thought: Extracting Visual Knowledge in Charles S. Peirce's Manuscripts with Vision-Language Models
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- Emu3.5: Native Multimodal Models are World Learners
- PairUni: Pairwise Training for Unified Multimodal Language Models
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering
- Amortized Moment Matching for Visual Generation
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- SafeEditor: Unified MLLM for Efficient Post-hoc T2I Safety Editing
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- Data-Centric Lessons To Improve Speech-Language Pretraining
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Generative Universal Verifier as Multimodal Meta-Reasoner
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Diffusion Transformers with Representation Autoencoders
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- InstructX: Towards Unified Visual Editing with MLLM Guidance
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- VUGEN: Visual Understanding priors for GENeration
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
- Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack Framework for Unified Multimodal Understanding and Generation Models
- OmniDFA: A Unified Framework for Open Set Synthesis Image Detection and Few-Shot Attribution
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- UniVid: The Open-Source Unified Video Model
- Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
- Planning with Unified Multimodal Models
- X-Streamer: Unified Human World Modeling with Audiovisual Interaction
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- UniECG: Understanding and Generating ECG in One Unified Model
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- GenExam: A Multidisciplinary Text-to-Image Exam
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Unified Multimodal Model as Auto-Encoder
- Reconstruction Alignment Improves Unified Multimodal Models
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- Interleaving Reasoning for Better Text-to-Image Generation
- The Telephone Game: Evaluating Semantic Drift in Unified Models
- Transition Models: Rethinking the Generative Learning Objective
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control
- Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- Elastic Diffusion Transformer
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
- Qwen-Image Technical Report
- PixNerd: Pixel Neural Field Diffusion
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
Related