Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
2024/08/20 by Chunting Zhou, Zhou, Chunting, Lili Yu +17 · 4 voices · 116 citations
Neuroscience · #Brain Tumor Detection and Classification
paper · pdf · doi:10.48550/arxiv.2408.11039
Abstract
We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over mixed-modality sequences. We pretrain multiple Transfusion models up to 7B parameters from scratch on a mixture of text and image data, establishing scaling laws with respect to a variety of uni- and cross-modal benchmarks. Our experiments show that Transfusion scales significantly better than quantizing images and training a language model over discrete image tokens. By introducing modality-specific encoding and decoding layers, we can further improve the performance of Transfusion models, and even compress each image to just 16 patches. We further demonstrate that scaling our Transfusion recipe to 7B parameters and 2T multi-modal tokens produces a model that can generate images and text on a par with similar scale diffusion models and language models, reaping the benefits of both worlds.
Cited by
- Twins: Learn to Predict Unified Representations with Focal Loss
- RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
- Pixel-Space Diffusion Transformers
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
- CreatiPoster: Towards Editable and Controllable Multi-Layer Graphic Design Generation
- StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
- T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
- STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
- Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- WorldVLA: Towards Autoregressive Action World Model
- Transfer between Modalities with MetaQueries
- Vision as LoRA
- Scaling GUI Agents with Visual State Transitions
- Self-Evaluation Unlocks Any-Step Text-to-Image Generation
- Bridging Your Imagination with Audio-Video Generation via a Unified Director
- Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- TV2TV: A Unified Framework for Interleaved Language and Video Generation
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- First Frame Is the Place to Go for Video Content Customization
- EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- TransactionGPT
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- PairUni: Pairwise Training for Unified Multimodal Language Models
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- InstructX: Towards Unified Visual Editing with MLLM Guidance
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- VUGEN: Visual Understanding priors for GENeration
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- IRIS: Intrinsic Reward Image Synthesis
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- UniVid: The Open-Source Unified Video Model
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- HunyuanImage 3.0 Technical Report
- Planning with Unified Multimodal Models
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- CAD-Tokenizer: Towards Text-based CAD Prototyping via Modality-Specific Tokenization
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- VIG-RL: Learning to Search and Insert for Verified Image Grounding
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- AToken: A Unified Tokenizer for Vision
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- Unified Multimodal Model as Auto-Encoder
- Reconstruction Alignment Improves Unified Multimodal Models
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- Interleaving Reasoning for Better Text-to-Image Generation
- The Telephone Game: Evaluating Semantic Drift in Unified Models
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- Multimodal learning with next-token prediction for large multimodal models
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
- PixNerd: Pixel Neural Field Diffusion
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
- Diffusion model [wikipedia]
Discussions
Related