Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
2024/06/10 by Peize Sun, Yi Jiang, Sun, Peize +11 · 1 voice · 310 citations
Computer Science · Mathematics · #Advanced Image and Video Retrieval Techniques #Artificial intelligence #Autoregressive model #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #Computer graphics (images) #Computer science #Computer vision #Database #Diffusion #Econometrics #FOS: Computer and information sciences #Image (mathematics) #Mathematics #Music and Audio Processing #Physics #Scalability #cs.CV
paper · pdf · doi:10.48550/arxiv.2406.06525
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/06/10 · arxiv published 2024/06/10 · arxiv updated 2024/06/10 · openalex created_date 2024/06/12 · openalex updated_date 2026/07/28
Abstract
We introduce LlamaGen, a new family of image generation models that apply original ``next-token prediction'' paradigm of large language models to visual generation domain. It is an affirmative answer to whether vanilla autoregressive models, e.g., Llama, without inductive biases on visual signals can achieve state-of-the-art image generation performance if scaling properly. We reexamine design spaces of image tokenizers, scalability properties of image generation models, and their training data quality. The outcome of this exploration consists of: (1) An image tokenizer with downsample ratio of 16, reconstruction quality of 0.94 rFID and codebook usage of 97% on ImageNet benchmark. (2) A series of class-conditional image generation models ranging from 111M to 3.1B parameters, achieving 2.18 FID on ImageNet 256x256 benchmarks, outperforming the popular diffusion models such as LDM, DiT. (3) A text-conditional image generation model with 775M parameters, from two-stage training on LAION-COCO and high aesthetics quality images, demonstrating competitive performance of visual quality and text alignment. (4) We verify the effectiveness of LLM serving frameworks in optimizing the inference speed of image generation models and achieve 326% - 414% speedup. We release all models and codes to facilitate open-source community of visual generation and multimodal foundation models.
Cited by
- Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression
- Distribution Matching Variational AutoEncoder
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
- Normalizing Trajectory Models
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- Next-Embedding Prediction Makes Strong Vision Learners
- SFTok: Bridging the Performance Gap in Discrete Tokenizers
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- Spherical Leech Quantization for Visual Tokenization and Generation
- ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
- Grounding Everything in Tokens for Multimodal Large Language Models
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- Training-Free Vector Quantization via Gaussian VAEs
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- Rethinking Training Dynamics in Scale-wise Autoregressive Generation
- Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
- DeRA: Decoupled Representation Alignment for Video Tokenization
- SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
- FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
- Visual Generation Tuning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Guiding Visual Autoregressive Models through Spectrum Weakening
- Markovian Scale Prediction: A New Era of Visual Autoregressive Generation
- Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment
- DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- DINO-Tok: Adapting DINO for Visual Tokenizers
- PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
- Diversity Has Always Been There in Your Visual Autoregressive Models
- Spanning Tree Autoregressive Visual Generation
- Decoupling Complexity from Scale in Latent Diffusion Model
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation
- VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
- Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
- MixAR: Mixture Autoregressive Image Generation
- Improved Masked Image Generation with Knowledge-Augmented Token Representations
- ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
- Emu3.5: Native Multimodal Models are World Learners
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- Amortized Moment Matching for Visual Generation
- FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- Uniform Discrete Diffusion with Metric Path for Video Generation
- TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
- MC-SJD : Maximal Coupling Speculative Jacobi Decoding for Autoregressive Visual Generation Acceleration
- FARMER: Flow AutoRegressive Transformer over Pixels
- Autoregressive Styled Text Image Generation, but Make it Reliable
- Nested AutoRegressive Models
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- Adversarial Concept Distillation for One-Step Diffusion Personalization
- EditInfinity: Image Editing with Binary-Quantized Generative Models
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Generation then Reconstruction: Accelerating Masked Autoregressive Models via Two-Stage Sampling
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- Latent Diffusion Model without Variational Autoencoder
- ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- End-to-End Multi-Modal Diffusion Mamba
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- BIGFix: Bidirectional Image Generation with Token Fixing
- Your VAR Model is Secretly an Efficient and Explainable Generative Classifier
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Heptapod: Language Modeling on Visual Signals
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- \bfD3QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
- Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Visual Self-Refinement for Autoregressive Models
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- Autoregressive Video Generation beyond Next Frames Prediction
- Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Stochastic Interpolants via Conditional Dependent Coupling
- ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Scale-Wise VAR is Secretly Discrete Diffusion
- Group Critical-token Policy Optimization for Autoregressive Image Generation
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- Autoregressive End-to-End Planning with Time-Invariant Spatial Alignment and Multi-Objective Policy Refinement
- The Unanticipated Asymmetry Between Perceptual Optimization and Assessment
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- GenExam: A Multidisciplinary Text-to-Image Exam
- Image Tokenizer Needs Post-Training
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- RewardDance: Reward Scaling in Visual Generation
- Exploring Autoregressive Vision Foundation Models for Image Compression
- Missing Fine Details in Images: Last Seen in High Frequencies
- Transition Models: Rethinking the Generative Learning Objective
- MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- 2D Gaussians Meet Visual Tokenizer
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
- Next Visual Granularity Generation
- Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- Grouped Speculative Decoding for Autoregressive Image Generation
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Towards Robust Red-Green Watermarking for Autoregressive Image Generators
- NEP: Autoregressive Image Editing via Next Editing Token Prediction
- WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- StarPose: 3D Human Pose Estimation via Spatial-Temporal Autoregressive Diffusion
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- Imbalance in Balance: Online Concept Balancing in Generation Models
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- A Practical Investigation of Spatially-Controlled Image Generation with Transformers
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
- Latent Denoising Makes Good Tokenizers
- SCALAR: Scale-wise Controllable Visual Autoregressive Learning
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
- CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
- SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping
- Quantize-then-Rectify: Efficient VQ-VAE Training
- MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
- ICAS: Detecting Training Data from Autoregressive Image Generative Models
- AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- Cautious Next Token Prediction
- Progressive Checkerboards for Autoregressive Multiscale Image Generation
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
- Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
- Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
- Masks make discriminative models great again!
- Transition Matching: Scalable and Flexible Generative Modeling
- CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation
- Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
- Exploring Image Generation via Mutually Exclusive Probability Spaces and Local Correlation Hypothesis
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
- Auto-Regressively Generating Multi-View Consistent Images
- Highly Compressed Tokenizer Can Generate Without Training
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
- Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
- RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought
- Watermarking Autoregressive Image Generation
- Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study
- Show-o2: Improved Native Unified Multimodal Models
- VideoMAR: Autoregressive Video Generatio with Continuous Tokens
- Discrete Diffusion in Large Language and Multimodal Models: A Survey
- Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
- SpectralAR: Spectral Autoregressive Visual Generation
- Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
- MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
- DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
- STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
- Aligning Latent Spaces with Flow Priors
- Multi-scale Image Super Resolution with a Single Auto-Regressive Model
- Video World Models with Long-term Spatial Memory
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- Beyond Token-Level Cross-Entropy: Fréchet Distributional Post-Training for Autoregressive Image Generation
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
- ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
- Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts
- HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
- Native-Resolution Image Synthesis
- Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
- TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
- D-AR: Diffusion via Autoregressive Models
- Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
- Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization
- Training Free Stylized Abstraction
- Thinking with Generated Images
- PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models
- Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
- DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction
- Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
- Long-Context State-Space Video World Models
- ReDDiT: Rehashing Noise for Discrete Visual Generation
- Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
- LlamaSeg: Image Segmentation via Autoregressive Mask Generation
- StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- ImgEdit: A Unified Image Editing Dataset and Benchmark
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation
- MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
- GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
- Conditional Panoramic Image Generation via Masked Autoregressive Modeling
- From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- Learning to Tessellate: Point Cloud Generation via Recursive Spectral Partitioning
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
- ChemMLLM: Chemical Multimodal Large Language Model
- MMaDA: Multimodal Large Diffusion Language Models
- IA-T2I: Internet-Augmented Text-to-Image Generation
- Intentional Gesture: Deliver Your Intentions with Gestures for Speech
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
- Context-Aware Autoregressive Models for Multi-Conditional Image Generation
- Video-GPT via Next Clip Diffusion
- FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
- Multi-Token Prediction Needs Registers
- Continuous Visual Autoregressive Generation via Score Maximization
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- Adaptive Protein Tokenization
- Geometric Autoencoder for Diffusion Models
- End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
- AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
- Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
- GarmentX: Autoregressive Parametric Representations for High-Fidelity 3D Garment Generation
- Learning Streaming Video Representation via Multitask Training
- Signal or Noise? Understanding Generative Models for Real-World Sensor Time Series
- PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- SMART: When is it Actually Worth Expanding a Speculative Tree?
- InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- Distilling Specialized Orders for Visual Generation
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
- Personalized Text-to-Image Generation with Auto-Regressive Models
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing
- OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape Generation
- D2iT: Dynamic Diffusion Transformer for Accurate Image Generation
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
- Position: Beyond Euclidean -- Foundation Models Should Embrace Non-Euclidean Geometries
- PixelFlow: Pixel-Space Generative Models with Flow
- VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
Discussions
Related