Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
2023/10/09 by Lijun Yu, José Lezama, Yu, Lijun +28 · 203 citations
Computer Science · Medicine · #Multimodal Machine Learning Applications #Topic Modeling #Artificial Intelligence in Healthcare and Education
paper · pdf · doi:10.48550/arxiv.2310.05737
Abstract
While Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effectively use LLMs for visual generation, one crucial component is the visual tokenizer that maps pixel-space inputs to discrete tokens appropriate for LLM learning. In this paper, we introduce MAGVIT-v2, a video tokenizer designed to generate concise and expressive tokens for both videos and images using a common token vocabulary. Equipped with this new tokenizer, we show that LLMs outperform diffusion models on standard image and video generation benchmarks including ImageNet and Kinetics. In addition, we demonstrate that our tokenizer surpasses the previously top-performing video tokenizer on two more tasks: (1) video compression comparable to the next-generation video codec (VCC) according to human evaluations, and (2) learning effective representations for action recognition tasks.
Cited by
- Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- TopoGR: Revealing and Preserving Latent Structure of Semantic ID in Generative Recommendation
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
- SFTok: Bridging the Performance Gap in Discrete Tokenizers
- ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens
- Spherical Leech Quantization for Visual Tokenization and Generation
- ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
- RecTok: Reconstruction Distillation along Rectified Flow
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models
- Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- Training-Free Vector Quantization via Gaussian VAEs
- Efficient Generative Transformer Operators For Million-Point PDEs
- DeRA: Decoupled Representation Alignment for Video Tokenization
- SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL
- Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models
- MM-ACT: Learn from Multimodal Parallel Generation to Act
- What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
- DINO-Tok: Adapting DINO for Visual Tokenizers
- STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
- One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control
- FVAR: Visual Autoregressive Modeling via Next Focus Prediction
- GloTok: Global Perspective Tokenizer for Image Reconstruction and Generation
- Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
- Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
- MixAR: Mixture Autoregressive Image Generation
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
- InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
- Towards One-step Causal Video Generation via Adversarial Self-Distillation
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
- Nested AutoRegressive Models
- Semantic Communications with World Models
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
- CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
- Group-Wise Optimization for Self-Extensible Codebooks in Vector Quantized Models
- BIGFix: Bidirectional Image Generation with Token Fixing
- Diffusion Transformers with Representation Autoencoders
- Test-Time Anchoring for Discrete Diffusion Posterior Sampling
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Heptapod: Language Modeling on Visual Signals
- Control-Augmented Autoregressive Diffusion for Data Assimilation
- BlockGPT: Spatio-Temporal Modelling of Rainfall via Frame-Level Autoregression
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- NeRV-Diffusion: Diffuse Implicit Neural Representations for Video Synthesis
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- Stochastic Interpolants via Conditional Dependent Coupling
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
- X-Streamer: Unified Human World Modeling with Audiovisual Interaction
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- AToken: A Unified Tokenizer for Vision
- Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
- Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting
- Missing Fine Details in Images: Last Seen in High Frequencies
- Transition Models: Rethinking the Generative Learning Objective
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- Robust Residual Finite Scalar Quantization for Neural Compression
- Pixels to Play: A Foundation Model for 3D Gameplay
- EgoTwin: Dreaming Body and View in First Person
- Matrix-game 2.0: An open-source real-time and streaming interactive world model
- Next Visual Granularity Generation
- Representing Speech Through Autoregressive Prediction of Cochlear Tokens
- Versatile Video Representation via Feed-Forward 2D Gaussian Splatting Tokenization
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
- Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
- LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
- Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
- Kronos: A Foundation Model for the Language of Financial Markets
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
- A Practical Investigation of Spatially-Controlled Image Generation with Transformers
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows
- Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
- Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation
- Seedance 1.0: Exploring the Boundaries of Video Generation Models
- Quantize-then-Rectify: Efficient VQ-VAE Training
- Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
- Infinite Video Understanding
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- WaiT for the Signal: Simple Frequency-Aware Flow-Matching
- MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
- Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
- A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- MambaVideo for Discrete Video Tokenization with Channel-Split Quantization
- StreamDiT: Real-Time Streaming Text-to-Video Generation
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- Progressive Checkerboards for Autoregressive Multiscale Image Generation
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
- Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
- Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
- RoboScape: Physics-informed Embodied World Model
- Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Highly Compressed Tokenizer Can Generate Without Training
- PhysiX: A Foundation Model for Physics Simulations
- SIDE: Semantic ID Embedding for effective learning from sequences
- Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
- Audio-Sync Video Generation with Multi-Stream Temporal Control
- Align Your Flow: Scaling Continuous-Time Flow Map Distillation
- Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
- VINCIE: Unlocking In-context Image Editing from Video
- DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- ContentV: Efficient Training of Video Generation Models with Limited Compute
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
- Phenotype-Guided Generative Model for High-Fidelity Cardiac MRI Synthesis: Advancing Pretraining and Clinical Applications
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
- D-AR: Diffusion via Autoregressive Models
- PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models
- EF-VI: Enhancing End-Frame Injection for Video Inbetweening
- PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
- multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
- Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
- ReDDiT: Rehashing Noise for Discrete Visual Generation
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- Plug-and-Play Context Feature Reuse for Efficient Masked Generation
- Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation
- Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
- From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- MMaDA: Multimodal Large Diffusion Language Models
- Interspatial Attention for Efficient 4D Human Video Generation
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- Universal Semantic Disentangled Privacy-preserving Speech Representation Learning
- GANCompress: GAN-Enhanced Neural Image Compression with Binary Spherical Quantization
- UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes
- Video-GPT via Next Clip Diffusion
- Generative Pre-trained Autoregressive Diffusion Transformer
- Continuous Visual Autoregressive Generation via Score Maximization
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- Geometric Autoencoder for Diffusion Models
- SOM-VQ: Topology-Aware Tokenization for Interactive Generative Models
- End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
- Ligand-Conditioned Discrete Diffusion for Protein Sequence-Structure Co-Design
- Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation
- PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
- Latent-Compressed Variational Autoencoder for Video Diffusion Models
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- FashionM3: Multimodal, Multitask, and Multiround Fashion Assistant based on Unified Vision-Language Model
- Fast Autoregressive Models for Continuous Latent Generation
- GVC-RT: Towards Real-Time Generative Video Compression at Ultra-Low Bitrates
- Distilling Specialized Orders for Visual Generation
- Boosting Generative Image Modeling via Joint Image-Feature Synthesis
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- Media Meets Communication in 6G: Fundamentals, Key Technologies, and Applications
- MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
- Elucidating the Design Space of Multimodal Protein Language Models
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
Related