Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
2024/08/22 by Jinheng Xie, Weijia Mao, Xie, Jinheng +17 · 256 citations
Computer Science · #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2408.12528
Abstract
We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibly supports a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation. Across various benchmarks, it demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation. This significantly highlights its potential as a next-generation foundation model. Code and models are released at https://github.com/showlab/Show-o.
Cited by
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- LongCat-Image Technical Report
- Scaling GUI Agents with Visual State Transitions
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- Bridging Your Imagination with Audio-Video Generation via a Unified Director
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
- VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- Loom: Diffusion-Transformer for Interleaved Generation
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
- ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
- Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- TV2TV: A Unified Framework for Interleaved Language and Video Generation
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- PhotoFramer: Multi-modal Image Composition Instruction
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
- FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- DINO-Tok: Adapting DINO for Visual Tokenizers
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
- Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
- Plan-X: Instruct Video Generation via Semantic Planning
- Diversity Has Always Been There in Your Visual Autoregressive Models
- UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- Emu3.5: Native Multimodal Models are World Learners
- PairUni: Pairwise Training for Unified Multimodal Language Models
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- Uniform Discrete Diffusion with Metric Path for Video Generation
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- End-to-End Multi-Modal Diffusion Mamba
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- InstructX: Towards Unified Visual Editing with MLLM Guidance
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Reinforcing Diffusion Models by Direct Group Preference Optimization
- Resolving the Identity Crisis in Text-to-Image Generation
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- VUGEN: Visual Understanding priors for GENeration
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- Pulp Motion: Framing-aware multimodal camera and human motion generation
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- LieHMR: Autoregressive Human Mesh Recovery with SO(3) Diffusion
- Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- IRIS: Intrinsic Reward Image Synthesis
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- UniVid: The Open-Source Unified Video Model
- Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
- HunyuanImage 3.0 Technical Report
- Planning with Unified Multimodal Models
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
- Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- AToken: A Unified Tokenizer for Vision
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- Unified Multimodal Model as Auto-Encoder
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Reconstruction Alignment Improves Unified Multimodal Models
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- Interleaving Reasoning for Better Text-to-Image Generation
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
- DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- AudioStory: Generating Long-Form Narrative Audio with Large Language Models
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
- Thyme: Think Beyond Images
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- Multimodal learning with next-token prediction for large multimodal models
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- Qwen-Image Technical Report
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
- PixNerd: Pixel Neural Field Diffusion
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
- Image Generators are Generalist Vision Learners
- FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models
- CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
- Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
- SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping
- SynCoGen: Synthesizable 3D Molecule Generation via Joint Reaction and Coordinate Modeling
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- IGD: Instructional Graphic Design with Multimodal Layer Generation
- Theory-Informed Improvements to Classifier-Free Guidance for Discrete Diffusion Models
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
- AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
- MotionGPT3: Human Motion as a Second Modality
- Dreamland: Controllable World Creation with Simulator and Generative Models
- EAR: Erasing Concepts from Unified Autoregressive Models
- MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- OmniGen2: Exploration to Advanced Multimodal Generation
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation
- UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
- FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
- Watermarking Autoregressive Image Generation
- Show-o2: Improved Native Unified Multimodal Models
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
- AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
- SpectralAR: Spectral Autoregressive Visual Generation
- Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
- STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- How Far Are We from Generating Missing Modalities with Foundation Models?
- HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
- UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
- IMAGHarmony: Controllable Image Editing with Consistent Object Quantity and Layout
- Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
- TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- D-AR: Diffusion via Autoregressive Models
- Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
- Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling
- MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
- OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
- R-Genie: Reasoning-Guided Generative Image Editing
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
- RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
- MMaDA: Multimodal Large Diffusion Language Models
- IA-T2I: Internet-Augmented Text-to-Image Generation
- Emerging Properties in Unified Multimodal Pretraining
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- Context-Aware Autoregressive Models for Multi-Conditional Image Generation
Related