Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
2024/10/09 by Sihyun Yu, Yu, Sihyun, Sangkyung Kwak +11 · 2 voices · 268 citations
Computer Science · Engineering · #Artificial intelligence #Computer science #Electrical engineering #Engineering #Physics #Political science #Politics #Reinforcement Learning in Robotics #Representation (politics) #Training (meteorology) #Transformer
paper · pdf · doi:10.48550/arxiv.2410.06940
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/10/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5×, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval.
Cited by
- SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation
- AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
- PILD: Physics-Informed Learning via Diffusion
- Conditioning Residuals for Diffusion Models via Representation Feedback
- Pixel-Space Diffusion Transformers
- Signed Rectified Flow: Negativity-Controlled Generation
- DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer
- REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
- Semantic Anchoring for Robotic Action Representations
- Beyond Converging Representations A Philosophical Response on the Interpretation Risks of Scientific Foundation Models
- Speedrunning ImageNet Diffusion
- MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
- Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
- Perception Encoder: The best visual embeddings are not at the output of the network
- How far can we go with ImageNet for Text-to-Image generation?
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
- T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
- PixelGen: Improving Pixel Diffusion with Perceptual Supervision
- Generalization of Diffusion Models Arises with a Balanced Representation Space
- SemanticGen: Video Generation in Semantic Space
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
- OMP: One-step Meanflow Policy with Directional Alignment
- MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
- MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
- RecTok: Reconstruction Distillation along Rectified Flow
- SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
- Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
- Bidirectional Normalizing Flow: From Data to Noise and Back
- What matters for Representation Alignment: Global Information or Spatial Structure?
- DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
- Repulsor: Accelerating Generative Modeling with a Contrastive Memory Bank
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- D3-Predictor: Noise-Free Deterministic Diffusion for Dense Prediction
- Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion
- DeRA: Decoupled Representation Alignment for Video Tokenization
- SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
- On the Design of One-step Diffusion via Shortcutting Flow Paths
- Network of Theseus (like the ship)
- UniLight: A Unified Representation for Lighting
- CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- LAP: Fast LAtent Diffusion Planner for Autonomous Driving
- Visual Generation Tuning
- SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Overcoming the Curvature Bottleneck in MeanFlow
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior
- Adversarial Flow Models
- MeanFlow Transformers with Representation Autoencoders
- Unlocking Zero-shot Potential of Semi-dense Image Matching via Gaussian Splatting
- DINO-Tok: Adapting DINO for Visual Tokenizers
- PixelDiT: Pixel Diffusion Transformers for Image Generation
- View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
- Flow Map Distillation Without Data
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
- A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- VeCoR -- Velocity Contrastive Regularization for Flow Matching
- CoD: A Diffusion Foundation Model for Image Compression
- Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment
- Decoupling Complexity from Scale in Latent Diffusion Model
- Semantic Context Matters: Improving Conditioning for Autoregressive Models
- Distribution Matching Distillation Meets Reinforcement Learning
- Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
- Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
- ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
- E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- Emu3.5: Native Multimodal Models are World Learners
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion
- Amortized Moment Matching for Visual Generation
- Generative Modeling via Drifting
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
- FARMER: Flow AutoRegressive Transformer over Pixels
- Simple Denoising Diffusion Language Models
- DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching
- Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
- Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video
- Improved Training Technique for Shortcut Models
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- AlphaFlow: Understanding and Improving MeanFlow Models
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
- Latent Diffusion Model without Variational Autoencoder
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
- LayerSync: Self-aligning Intermediate Layers
- BIGFix: Bidirectional Image Generation with Token Fixing
- Diffusion Transformers with Representation Autoencoders
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Asymmetric Flow Models
- Latent Harmony: Synergistic Unified UHD Image Restoration via Latent Space Regularization and Controllable Refinement
- CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving
- Optimal Stopping in Latent Diffusion Models
- Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
- MATRIX: Mask Track Alignment for Interaction-aware Video Generation
- Heptapod: Language Modeling on Visual Signals
- VUGEN: Visual Understanding priors for GENeration
- LARA-Gen: Enabling Continuous Emotion Control for Music Generation Models via Latent Affective Representation Alignment
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Flow-Matching Based Refiner for Molecular Conformer Generation
- REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
- Glocal Information Bottleneck for Time Series Imputation
- RAP: 3D Rasterization Augmented End-to-End Planning
- What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
- Align Your Query: Representation Alignment for Multimodality Medical Object Detection
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Align Your Tangent: Training Better Consistency Models via Manifold-Aligned Tangents
- Selective Underfitting in Diffusion Models
- MorphGen: Controllable and Morphologically Plausible Generative Cell-Imaging
- LUMA: Low-Dimension Unified Motion Alignment with Dual-Path Anchoring for Text-to-Motion Diffusion Model
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow Map Models
- DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
- Scalable GANs with Transformers
- VoiceBridge: Designing Latent Bridge Models for General Speech Restoration at Scale
- HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
- WoW: Towards a World omniscient World model Through Embodied Interaction
- HiGS: History-Guided Sampling for Plug-and-Play Enhancement of Diffusion Models
- Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
- Generation Properties of Stochastic Interpolation under Finite Training Set
- No Alignment Needed for Generation: Learning Linearly Separable Representations in Diffusion Models
- Embodied Representation Alignment with Mirror Neurons
- Flow Matching in the Low-Noise Regime: Pathologies and a Contrastive Remedy
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
- Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
- AToken: A Unified Tokenizer for Vision
- AnyAccomp: Generalizable Accompaniment Generation via Quantized Melodic Bottleneck
- Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting
- Visual Representation Alignment for Multimodal Large Language Models
- Reconstruction Alignment Improves Unified Multimodal Models
- Home-made Diffusion Model from Scratch to Hatch
- Transition Models: Rethinking the Generative Learning Objective
- DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
- GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation
- Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
- Waver: Wave Your Way to Lifelike Video Generation
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Next Visual Granularity Generation
- ENA: Efficient N-dimensional Attention
- Elucidating the Role of Feature Normalization in IJEPA
- Versatile Transition Generation with Image-to-Video Diffusion
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
- DivControl: Knowledge Diversion for Controllable Image Generation
- PixNerd: Pixel Neural Field Diffusion
- A Practical Investigation of Spatially-Controlled Image Generation with Transformers
- Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
- Latent Denoising Makes Good Tokenizers
- LEAF: Latent Diffusion with Efficient Encoder Distillation for Aligned Features in Medical Image Segmentation
- One-step Latent-free Image Generation with Pixel Mean Flows
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
- Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
- Diffuse and Disperse: Image Generation with Representation Regularization
- Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
- Learning Diffusion Models with Flexible Representation Guidance
- Theory-Informed Improvements to Classifier-Free Guidance for Discrete Diffusion Models
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
- Contrastive Conditional-Unconditional Alignment for Long-tailed Diffusion Model
- WaiT for the Signal: Simple Frequency-Aware Flow-Matching
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
- Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspective
- Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
- FACM: Flow-Anchored Consistency Models
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
- A Gift from the Integration of Discriminative and Diffusion-based Generative Learning: Boundary Refinement Remote Sensing Semantic Segmentation
- Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
- Autoregressive Denoising Score Matching is a Good Video Anomaly Detector
- Guidance in the Frequency Domain Enables High-Fidelity Sampling at Low CFG Scales
- Light of Normals: Unified Feature Representation for Universal Photometric Stereo
- UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
- Generative Modeling of Weights: Generalization or Memorization?
- Reimagining Target-Aware Molecular Generation through Retrieval-Enhanced Aligned Diffusion
- Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value
- Contrastive Flow Matching
- Aligning Latent Spaces with Flow Priors
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- Native-Resolution Image Synthesis
- FourierFlow: Frequency-aware Flow Matching for Generative Turbulence Modeling
- D-AR: Diffusion via Autoregressive Models
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
- ACE-Step: A Step Towards Music Generation Foundation Model
- Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
- Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
- No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves
- ReDDiT: Rehashing Noise for Discrete Visual Generation
- Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation
- Plug-and-Play Context Feature Reuse for Efficient Masked Generation
- Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning
- On Denoising Walking Videos for Gait Recognition
- Diffusion Classifiers Understand Compositionality, but Conditions Apply
- TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
- TokBench: Evaluating Your Visual Tokenizer before Visual Generation
- REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
- FLARE: Robot Learning with Implicit World Modeling
- Scaling Diffusion Transformers Efficiently via μP
- Intentional Gesture: Deliver Your Intentions with Gestures for Speech
- CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
- Mean Flows for One-step Generative Modeling
- Training Latent Diffusion Models with Interacting Particle Algorithms
- Neural Thermodynamics: Entropic Forces in Deep and Universal Representation Learning
- VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
- UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
- Where Does Generative Difficulty Reside? An Empirical Study of Target Representations
- Distilling Drifting Transformers with Representation Autoencoders
- Generative Pre-trained Autoregressive Diffusion Transformer
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- Geometric Action Model for Robot Policy Learning
- Scalable and Interpretable Representation Alignment with Ordinal Similarity
- Taming Outlier Tokens in Diffusion Transformers
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens
- CytoSyn: a Foundation Diffusion Model for Histopathology -- Tech Report
- SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
- V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
- Adaptive Protein Tokenization
- Geometric Autoencoder for Diffusion Models
- A Systematic Evaluation of Co-folding Model Representations for Small-Molecule Learning
- DreamWorld: Unified World Modeling in Video Generation
- Unpaired Image-to-Image Translation via a Self-Supervised Semantic Bridge
- End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Image Generation with a Sphere Encoder
- GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- Atom-level Protein Representation Learning Improves Protein Structure Prediction
- X-Fusion: Introducing New Modality to Frozen Large Language Models
- AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- Continuous Adversarial Flow Models
- R3DPA: Leveraging 3D Representation Alignment and RGB Pretrained Priors for LiDAR Scene Generation
- Enhancing Variational Autoencoders with Smooth Robust Latent Encoding
- From Understanding to Erasing: Towards Complete and Stable Video Object Removal
- Boosting Generative Image Modeling via Joint Image-Feature Synthesis
- Diffusion Generative Recommendation with Continuous Tokens
- KVAE: Family of Tokenizers for Multimodal Generative Models
- Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models
- NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
- Seedream 3.0 Technical Report
- Elucidating the Design Space of Multimodal Protein Language Models
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- Efficient Generative Model Training via Embedded Representation Warmup
- GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
- TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
Discussions
Related