Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
2024/08/20 by Chunting Zhou, Zhou, Chunting, Lili Yu +17 · 4 voices · 196 citations
Neuroscience · #Artificial intelligence #Brain Tumor Detection and Classification #Computer science #Computer security #Materials science #Modal #Security token
paper · pdf · doi:10.48550/arxiv.2408.11039
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/08/20 · openalex created_date 2024/10/01 · openalex updated_date 2026/07/28
Abstract
We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over mixed-modality sequences. We pretrain multiple Transfusion models up to 7B parameters from scratch on a mixture of text and image data, establishing scaling laws with respect to a variety of uni- and cross-modal benchmarks. Our experiments show that Transfusion scales significantly better than quantizing images and training a language model over discrete image tokens. By introducing modality-specific encoding and decoding layers, we can further improve the performance of Transfusion models, and even compress each image to just 16 patches. We further demonstrate that scaling our Transfusion recipe to 7B parameters and 2T multi-modal tokens produces a model that can generate images and text on a par with similar scale diffusion models and language models, reaping the benefits of both worlds.
Cited by
- Twins: Learn to Predict Unified Representations with Focal Loss
- RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
- Pixel-Space Diffusion Transformers
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
- CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation
- StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
- T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
- STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
- Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- WorldVLA: Towards Autoregressive Action World Model
- Transfer between Modalities with MetaQueries
- Vision as LoRA
- Scaling GUI Agents with Visual State Transitions
- Self-Evaluation Unlocks Any-Step Text-to-Image Generation
- Bridging Your Imagination with Audio-Video Generation via a Unified Director
- Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- TV2TV: A Unified Framework for Interleaved Language and Video Generation
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- First Frame Is the Place to Go for Video Content Customization
- EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- TransactionGPT
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- PairUni: Pairwise Training for Unified Multimodal Language Models
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- InstructX: Towards Unified Visual Editing with MLLM Guidance
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- VUGEN: Visual Understanding priors for GENeration
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- IRIS: Intrinsic Reward Image Synthesis
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- UniVid: The Open-Source Unified Video Model
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- HunyuanImage 3.0 Technical Report
- Planning with Unified Multimodal Models
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- CAD-Tokenizer: Towards Text-based CAD Prototyping via Modality-Specific Tokenization
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- VIG-RL: Learning to Search and Insert for Verified Image Grounding
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- AToken: A Unified Tokenizer for Vision
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- Unified Multimodal Model as Auto-Encoder
- Reconstruction Alignment Improves Unified Multimodal Models
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- Interleaving Reasoning for Better Text-to-Image Generation
- The Telephone Game: Evaluating Semantic Drift in Unified Models
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- Multimodal learning with next-token prediction for large multimodal models
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
- PixNerd: Pixel Neural Field Diffusion
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
- Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
- Generative Distribution Distillation
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
- Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
- From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization
- Token Communication in the Era of Large Models: An Information Bottleneck-Based Approach
- MotionGPT3: Human Motion as a Second Modality
- Transition Matching: Scalable and Flexible Generative Modeling
- Dreamland: Controllable World Creation with Simulator and Generative Models
- SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- Highly Compressed Tokenizer Can Generate Without Training
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
- Watermarking Autoregressive Image Generation
- Aligning Text, Images, and 3D Structure Token-by-Token
- Show-o2: Improved Native Unified Multimodal Models
- LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
- LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
- Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
- STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- X-Driver: Explainable Autonomous Driving with Vision-Language Models
- SeedEdit 3.0: Fast and High-Quality Generative Image Editing
- HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
- Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack
- D-AR: Diffusion via Autoregressive Models
- Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
- Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling
- OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
- R-Genie: Reasoning-Guided Generative Image Editing
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- MMaDA: Multimodal Large Diffusion Language Models
- Emerging Properties in Unified Multimodal Pretraining
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- Context-Aware Autoregressive Models for Multi-Conditional Image Generation
- Video-GPT via Next Clip Diffusion
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
- A Survey of Interactive Generative Video
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- YoChameleon: Personalized Vision and Language Generation
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- DiMeR: Disentangled Mesh Reconstruction Model
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- PixelFlow: Pixel-Space Generative Models with Flow
- Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability
- MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
- Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
- Diffusion model [wikipedia]
Discussions
Related