Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
2024/06/10 by Peize Sun, Yi Jiang, Sun, Peize +11 · 166 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing
paper · pdf · doi:10.48550/arxiv.2406.06525
openalex publication_date 2024/06/10 · openalex created_date 2024/06/12 · openalex updated_date 2026/07/28
Abstract
We introduce LlamaGen, a new family of image generation models that apply original ``next-token prediction'' paradigm of large language models to visual generation domain. It is an affirmative answer to whether vanilla autoregressive models, e.g., Llama, without inductive biases on visual signals can achieve state-of-the-art image generation performance if scaling properly. We reexamine design spaces of image tokenizers, scalability properties of image generation models, and their training data quality. The outcome of this exploration consists of: (1) An image tokenizer with downsample ratio of 16, reconstruction quality of 0.94 rFID and codebook usage of 97% on ImageNet benchmark. (2) A series of class-conditional image generation models ranging from 111M to 3.1B parameters, achieving 2.18 FID on ImageNet 256x256 benchmarks, outperforming the popular diffusion models such as LDM, DiT. (3) A text-conditional image generation model with 775M parameters, from two-stage training on LAION-COCO and high aesthetics quality images, demonstrating competitive performance of visual quality and text alignment. (4) We verify the effectiveness of LLM serving frameworks in optimizing the inference speed of image generation models and achieve 326% - 414% speedup. We release all models and codes to facilitate open-source community of visual generation and multimodal foundation models.
Cited by
- Distribution Matching Variational AutoEncoder
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- Normalizing Trajectory Models
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- Next-Embedding Prediction Makes Strong Vision Learners
- SFTok: Bridging the Performance Gap in Discrete Tokenizers
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- Spherical Leech Quantization for Visual Tokenization and Generation
- ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
- Grounding Everything in Tokens for Multimodal Large Language Models
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- Training-Free Vector Quantization via Gaussian VAEs
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
- DeRA: Decoupled Representation Alignment for Video Tokenization
- SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
- FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
- Visual Generation Tuning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Guiding Visual Autoregressive Models through Spectrum Weakening
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- DINO-Tok: Adapting DINO for Visual Tokenizers
- PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
- Diversity Has Always Been There in Your Visual Autoregressive Models
- Spanning Tree Autoregressive Visual Generation
- Decoupling Complexity from Scale in Latent Diffusion Model
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- MixAR: Mixture Autoregressive Image Generation
- Improved Masked Image Generation with Knowledge-Augmented Token Representations
- ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
- Emu3.5: Native Multimodal Models are World Learners
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- Amortized Moment Matching for Visual Generation
- FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- Uniform Discrete Diffusion with Metric Path for Video Generation
- TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- FARMER: Flow AutoRegressive Transformer over Pixels
- Autoregressive Styled Text Image Generation, but Make it Reliable
- Nested AutoRegressive Models
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- EchoDistill: Bidirectional Concept Distillation for One-Step Diffusion Personalization
- EditInfinity: Image Editing with Binary-Quantized Generative Models
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- Latent Diffusion Model without Variational Autoencoder
- ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- End-to-End Multi-Modal Diffusion Mamba
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- BIGFix: Bidirectional Image Generation with Token Fixing
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Heptapod: Language Modeling on Visual Signals
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Visual Self-Refinement for Autoregressive Models
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Autoregressive Video Generation beyond Next Frames Prediction
- Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Stochastic Interpolants via Conditional Dependent Coupling
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Scale-Wise VAR is Secretly Discrete Diffusion
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- Autoregressive End-to-End Planning with Time-Invariant Spatial Alignment and Multi-Objective Policy Refinement
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- GenExam: A Multidisciplinary Text-to-Image Exam
- Image Tokenizer Needs Post-Training
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- RewardDance: Reward Scaling in Visual Generation
- Exploring Autoregressive Vision Foundation Models for Image Compression
- Missing Fine Details in Images: Last Seen in High Frequencies
- Transition Models: Rethinking the Generative Learning Objective
- MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- 2D Gaussians Meet Visual Tokenizer
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
- Next Visual Granularity Generation
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- Grouped Speculative Decoding for Autoregressive Image Generation
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Towards Robust Red-Green Watermarking for Autoregressive Image Generators
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- StarPose: 3D Human Pose Estimation via Spatial-Temporal Autoregressive Diffusion
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
Related