Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
2024/04/03 by Keyu Tian, Yi Jiang, Tian, Keyu +7 · 6 voices · 240 citations
#cs.CV #cs.AI
paper · pdf · doi:10.48550/arxiv.2404.02905
Abstract
We present Visual AutoRegressive modeling (VAR), a new generation paradigm that redefines the autoregressive learning on images as coarse-to-fine "next-scale prediction" or "next-resolution prediction", diverging from the standard raster-scan "next-token prediction". This simple, intuitive methodology allows autoregressive (AR) transformers to learn visual distributions fast and generalize well: VAR, for the first time, makes GPT-like AR models surpass diffusion transformers in image generation. On ImageNet 256x256 benchmark, VAR significantly improve AR baseline by improving Frechet inception distance (FID) from 18.65 to 1.73, inception score (IS) from 80.4 to 350.2, with around 20x faster inference speed. It is also empirically verified that VAR outperforms the Diffusion Transformer (DiT) in multiple dimensions including image quality, inference speed, data efficiency, and scalability. Scaling up VAR models exhibits clear power-law scaling laws similar to those observed in LLMs, with linear correlation coefficients near -0.998 as solid evidence. VAR further showcases zero-shot generalization ability in downstream tasks including image in-painting, out-painting, and editing. These results suggest VAR has initially emulated the two important properties of LLMs: Scaling Laws and zero-shot task generalization. We have released all models and codes to promote the exploration of AR/VAR models for visual generation and unified learning.
Cited by
- NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Normalizing Trajectory Models
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
- GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
- Next-Scale Prediction: A Self-Supervised Approach for Real-World Image Denoising
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
- Next-Embedding Prediction Makes Strong Vision Learners
- SFTok: Bridging the Performance Gap in Discrete Tokenizers
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- Spherical Leech Quantization for Visual Tokenization and Generation
- From Feature Interaction to Feature Generation: A Generative Paradigm of CTR Prediction Models
- RecTok: Reconstruction Distillation along Rectified Flow
- Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- Bidirectional Normalizing Flow: From Data to Noise and Back
- Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- Training-Free Vector Quantization via Gaussian VAEs
- M-STAR: Multi-Scale Spatiotemporal Autoregression for Human Mobility Modeling
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- An Efficient Test-Time Scaling Approach for Image Generation
- StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow Generation
- FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
- SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- On the Design of One-step Diffusion via Shortcutting Flow Paths
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- Visual Generation Tuning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- REVEAL: Reasoning-enhanced Forensic Evidence Analysis for Explainable AI-generated Image Detection
- Guiding Visual Autoregressive Models through Spectrum Weakening
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- CoFiRec: Coarse-to-Fine Tokenization for Generative Recommendation
- Adversarial Flow Models
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
- The Collapse of Patches
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- MeanFlow Transformers with Representation Autoencoders
- Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- MFM-point: Multi-scale Flow Matching for Point Cloud Generation
- From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- DINO-Tok: Adapting DINO for Visual Tokenizers
- PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- Understanding, Accelerating, and Improving MeanFlow Training
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
- Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
- Diversity Has Always Been There in Your Visual Autoregressive Models
- Spanning Tree Autoregressive Visual Generation
- Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content
- Progressive Supernet Training for Efficient Visual Autoregressive Modeling
- AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
- CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- Self-Supervised Compression and Artifact Correction for Streaming Underwater Imaging Sonar
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- MixAR: Mixture Autoregressive Image Generation
- ReCast: Reliability-aware Codebook Assisted Lightweight Time Series Forecasting
- Improved Masked Image Generation with Knowledge-Augmented Token Representations
- ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
- Image Aesthetic Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- Medical Referring Image Segmentation via Next-Token Mask Prediction
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- Principled Coarse-Grained Acceptance for Speculative Decoding in Speech
- THD-BAR: Topology Hierarchical Derived Brain Autoregressive Modeling for EEG Generic Representations
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
- Emu3.5: Native Multimodal Models are World Learners
- Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- Amortized Moment Matching for Visual Generation
- Multi-scale Autoregressive Models are Laplacian, Discrete, and Latent Diffusion Models in Disguise
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry
- Approaching Low-Cost Cardiac Intelligence with Semi-Supervised Knowledge Distillation
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
- FARMER: Flow AutoRegressive Transformer over Pixels
- Nested AutoRegressive Models
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
- TEMPO: Temporal Multi-scale Autoregressive Generation of Protein Conformational Ensembles
- Improved Training Technique for Shortcut Models
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- EchoDistill: Bidirectional Concept Distillation for One-Step Diffusion Personalization
- UniSE: A Unified Framework for Decoder-only Autoregressive LM-based Speech Enhancement
- EditInfinity: Image Editing with Binary-Quantized Generative Models
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
- ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- Group-Wise Optimization for Self-Extensible Codebooks in Vector Quantized Models
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- LayerSync: Self-aligning Intermediate Layers
- BIGFix: Bidirectional Image Generation with Token Fixing
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- Diffusion Transformers with Representation Autoencoders
- Improving Text-to-Image Generation with Input-Side Inference-Time Scaling
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- Joint Discriminative-Generative Modeling via Dual Adversarial Training
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Next Semantic Scale Prediction via Hierarchical Diffusion Language Models
- Dynamic Mixture-of-Experts for Visual Autoregressive Model
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Heptapod: Language Modeling on Visual Signals
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- BlockGPT: Spatio-Temporal Modelling of Rainfall via Frame-Level Autoregression
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- GenAR: Next-Scale Autoregressive Generation for Spatial Gene Expression Prediction
- MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
- Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Visual Self-Refinement for Autoregressive Models
- UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- MARS: Audio Generation via Multi-Channel Autoregression on Spectrograms
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling View
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Scalable GANs with Transformers
- Autoregressive Video Generation beyond Next Frames Prediction
- Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
- Texture Vector-Quantization and Reconstruction Aware Prediction for Generative Super-Resolution
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Stochastic Interpolants via Conditional Dependent Coupling
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Scale-Wise VAR is Secretly Discrete Diffusion
- Un-Doubling Diffusion: LLM-guided Disambiguation of Homonym Duplication
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- ARMesh: Autoregressive Mesh Generation via Next-Level-of-Detail Prediction
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- GenExam: A Multidisciplinary Text-to-Image Exam
- FoundDiff: Foundational Diffusion Model for Generalizable Low-Dose CT Denoising
- Image Tokenizer Needs Post-Training
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- Medverse: A Universal Model for Full-Resolution 3D Medical Image Segmentation, Transformation and Enhancement
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- Exploring Autoregressive Vision Foundation Models for Image Compression
- Missing Fine Details in Images: Last Seen in High Frequencies
- Transition Models: Rethinking the Generative Learning Objective
- OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction
- Human Motion Video Generation: A Survey
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Deep learning-enabled virtual multiplexed immunostaining of label-free tissue for vascular invasion assessment
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- Distillation of a tractable model from the VQ-VAE
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- Localizing and Mitigating Memorization in Image Autoregressive Models
- Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- Fractal Flow: Hierarchical and Interpretable Normalizing Flow via Topic Modeling and Recursive Strategy
- Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model
- Generative AI in Map-Making: A Technical Exploration and Its Implications for Cartographers
- Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- SATURN: Autoregressive Image Generation Guided by Scene Graphs
- InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
- Next Visual Granularity Generation
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- Images Speak Louder Than Scores: Failure Mode Escape for Enhancing Generative Quality
- Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation
- Scaling Learned Image Compression Models up to 1 Billion
- LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
- Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Towards Robust Red-Green Watermarking for Autoregressive Image Generators
- M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
- Multitask Learning with Stochastic Interpolants
- HPSv3: Towards Wide-Spectrum Human Preference Score
- Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
- CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- StarPose: 3D Human Pose Estimation via Spatial-Temporal Autoregressive Diffusion
- StrandDesigner: Towards Practical Strand Generation with Sketch Guidance
- Graph Lineages and Skeletal Graph Products
- GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
- HRVVS: A High-resolution Video Vasculature Segmentation Network via Hierarchical Autoregressive Residual Priors
- Neural Autoregressive Modeling of Brain Aging
Discussions
- 📖For this week's @milanlp.bsky.social reading group, Yujie Ren presented "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction" by Keyu Tian et al.
Paper: arxiv.org/pd [bsky, 7 points, 0 comments]
- there’s been some interesting work lately on multiscale autoregressive image modeling arxiv.org/abs/2404.029... [bsky, 6 points, 1 comments]
- GPT 4o's new image capabilities seem to be liked. The insinuation from OpenAI seems to be that it is not based on diffusion. I wonder how their work relates to the infamous NeurIPS paper "Visual Autor [bsky, 3 points, 1 comments]
- 4/5 The debate of auto-regressive vs. diffusion continues. I thought for images diffusion had won, but the NeurIPS best paper (below) and the new @GroqInc image model are auto-regressive. Diffusion fo [bsky, 1 points, 1 comments]
- Visual Autoregressive Modeling: Image Generation via Next-Resolution Prediction [hn, 1 points, 1 comments]
- I may be late to the party, but these image generation models are SUPERB! Visual autoregressive modeling:
arxiv.org/pdf/2404.02905
arxiv.org/pdf/2412.04431
1- Use a VQ auto encoder to encode patches o [bsky, 0 points, 0 comments]
Related