Consistency Models
2023/03/02 by Yang Song, Prafulla Dhariwal, Song, Yang +5 · 265 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Image Processing Techniques #Cell Image Analysis Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML)
paper · pdf · doi:10.48550/arxiv.2303.01469
openalex publication_date 2023/03/02 · openalex created_date 2023/03/05 · openalex updated_date 2026/07/28
Abstract
Diffusion models have significantly advanced the fields of image, audio, and video generation, but they depend on an iterative sampling process that causes slow generation. To overcome this limitation, we propose consistency models, a new family of models that generate high quality samples by directly mapping noise to data. They support fast one-step generation by design, while still allowing multistep sampling to trade compute for sample quality. They also support zero-shot data editing, such as image inpainting, colorization, and super-resolution, without requiring explicit training on these tasks. Consistency models can be trained either by distilling pre-trained diffusion models, or as standalone generative models altogether. Through extensive experiments, we demonstrate that they outperform existing distillation techniques for diffusion models in one- and few-step sampling, achieving the new state-of-the-art FID of 3.55 on CIFAR-10 and 6.20 on ImageNet 64x64 for one-step generation. When trained in isolation, consistency models become a new family of generative models that can outperform existing one-step, non-adversarial generative models on standard benchmarks such as CIFAR-10, ImageNet 64x64 and LSUN 256x256.
Cited by
- Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
- ReDiF: Reinforced Distillation for Few Step Diffusion
- Parallel Diffusion Solver via Residual Dirichlet Policy Optimization
- Energy-Guided Flow Matching Enables Few-Step Conformer Generation and Ground-State Identification
- Normalizing Trajectory Models
- FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
- Autoregressive One-Step Generative Modeling for Dynamical System Forecasting
- Residual Prior Diffusion: A Probabilistic Framework Integrating Coarse Latent Priors with Diffusion Models
- PairFlow: Closed-Form Source-Target Coupling for Few-Step Generation in Discrete Flow Models
- Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
- OMP: One-step Meanflow Policy with Directional Alignment
- WaTeRFlow: Watermark Temporal Robustness via Flow Consistency
- Is Your Conditional Diffusion Model Actually Denoising?
- SFBD-OMNI: Bridge models for lossy measurement restoration with limited clean samples
- FlowDet: Unifying Object Detection and Generative Transport Flows
- SoFlow: Solution Flow Models for One-Step Generative Modeling
- DiffusionBrowser: Interactive Diffusion Previews via Multi-Branch Decoders
- Image Diffusion Preview with Consistency Solver
- BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation
- Few-Step Distillation for Text-to-Image Generation: A Practical Guide
- Inverse problems with diffusion models: MAP estimation via mode-seeking loss
- FALCON: Few-step Accurate Likelihoods for Continuous Flows
- CFO: Learning Continuous-Time PDE Dynamics via Flow-Matched Neural Operators
- One-Step Diffusion Samplers via Self-Distillation and Deterministic Flow
- Worst-case generation via minimax optimization in Wasserstein space
- SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
- Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- ICM-SR: Image-Conditioned Manifold Regularization for Image Super-Resolution
- Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
- ReflexFlow: Rethinking Learning Objective for Exposure Bias Alleviation in Flow Matching
- Data-regularized Reinforcement Learning for Diffusion Models at Scale
- On the Design of One-step Diffusion via Shortcutting Flow Paths
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- CSMapping: Scalable Crowdsourced Semantic Mapping and Topology Inference for Autonomous Driving
- Glance: Accelerating Diffusion Models with 1 Sample
- Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
- Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Much Ado About Noising: Dispelling the Myths of Generative Robotic Control
- Syndrome-Flow Consistency Model Achieves One-step Denoising Error Correction Codes
- Accelerating Inference of Masked Image Generators via Reinforcement Learning
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- InvarDiff: Cross-Scale Invariance Caching for Accelerated Diffusion Models
- Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium
- Test-time scaling of diffusions with flow maps
- Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration
- Adversarial Flow Models
- Rethinking Test Time Scaling for Flow-Matching Generative Models
- Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage
- MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
- MeanFlow Transformers with Representation Autoencoders
- FaithFusion: Harmonizing Reconstruction and Generation via Pixel-wise Information Gain
- From Diffusion to One-Step Generation: A Comparative Study of Flow-Based Models with Application to Image Inpainting
- Dynamic Test-Time Compute Scaling in Control Policy: Difficulty-Aware Stochastic Interpolant Policy
- CDLM: Consistency Diffusion Language Models For Faster Sampling
- TReFT: Taming Rectified Flow Models For One-Step Image Translation
- Terminal Velocity Matching
- Flow Map Distillation Without Data
- Understanding, Accelerating, and Improving MeanFlow Training
- FlowSteer: Guiding Few-Step Image Synthesis with Authentic Trajectories
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
- Uni-DAD: Unified Distillation and Adaptation of Diffusion Models for Few-step Few-shot Image Generation
- L1 Sample Flow for Efficient Visuomotor Learning
- One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
- Generative Augmented Reality: Paradigms, Technologies, and Future Applications
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- Complex-Valued 2D Gaussian Representation for Computer-Generated Holography
- Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
- Functional Mean Flow in Hilbert Space
- Learning Straight Flows: Variational Flow Matching for Efficient Generation
- Null-Space Diffusion Distillation Unlocks Speed, Fidelity and Realism in Lensless Imaging
- Free3D: 3D Human Motion Emerges from Single-View 2D Supervision
- Fast Data Attribution for Text-to-Image Models
- One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter for Your Diffusion Models
- SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
- Neodragon: Mobile Video Generation using Diffusion Transformer
- Enhancing Diffusion Model Guidance through Calibration and Regularization
- Multi-agent Coordination via Flow Matching
- Efficient probabilistic surrogate modeling techniques for partially-observed large-scale dynamical systems
- Diffusion Models Bridge Deep Learning and Physics in ENSO Forecasting
- Conditional Diffusion Model-Enabled Scenario-Specific Neural Receivers for Superimposed Pilot Schemes
- Wavelet-Optimized Motion Artifact Correction in 3D MRI Using Pre-trained 2D Score Priors
- Towards One-step Causal Video Generation via Adversarial Self-Distillation
- Improving Robustness to Out-of-Distribution States in Imitation Learning via Deep Koopman-Boosted Diffusion Policy
- Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
- Simplex-to-Euclidean Bijections for Categorical Flow Matching
- Hybrid Consistency Policy: Decoupling Multi-Modal Diversity and Real-Time Efficiency in Robotic Manipulation
- LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
- SplitFlow: Flow Decomposition for Inversion-Free Text-to-Image Editing
- Neural Stochastic Flows: Solver-Free Modelling and Inference for SDE Solutions
- Flow Map Learning via Nongradient Vector Flow
- Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
- DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
- Amortized Moment Matching for Visual Generation
- Multi-scale Autoregressive Models are Laplacian, Discrete, and Latent Diffusion Models in Disguise
- D2 Actor Critic: Diffusion Actor Meets Distributional Critic
- TGCM: Topic-Guided Consistency Modeling for One-Step Disentanglement of Interleaved APT Technique Sequences
- High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation
- Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models
- Balanced conic rectified flow
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
- Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
- Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
- Two-Steps Diffusion Policy for Robotic Manipulation via Genetic Denoising
- Generalised Flow Maps for Few-Step Generative Modelling on Riemannian Manifolds
- Improved Training Technique for Shortcut Models
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
- AutoScape: Geometry-Consistent Long-Horizon Scene Generation
- EchoDistill: Bidirectional Concept Distillation for One-Step Diffusion Personalization
- Poisson Flow Consistency Training
- AlphaFlow: Understanding and Improving MeanFlow Models
- A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
- Diffusion Buffer for Online Generative Speech Enhancement
- Gradient Variance Reveals Failure Modes in Flow-Based Generative Models
- GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
- On Efficiency-Effectiveness Trade-off of Diffusion-based Recommenders
- Adaptive Discretization for Consistency Models
- HumanCM: One Step Human Motion Prediction
- Deep generative priors for 3D brain analysis
- AlignFlow: Improving Flow-based Generative Models with Semi-Discrete Optimal Transport
- Learning an Image Editing Model without Image Editing Pairs
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- Restoring Noisy Demonstration for Imitation Learning With Diffusion Models
- Spatial Computing Communications for Multi-User Virtual Reality in Distributed Mobile Edge Computing Network
- Generative human motion mimicking through feature extraction in denoising diffusion settings
- pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
- DiffLoc: Diffusion Model-Based High-Precision Positioning for 6G Networks
- Shortcutting Pre-trained Flow Matching Diffusion Models is Almost Free Lunch
- FlashWorld: High-quality 3D Scene Generation within Seconds
- Adapting Noise to Data: Generative Flows from 1D Processes
- Mitigating the Noise Shift for Denoising Generative Models via Noise Awareness Guidance
- There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
- Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion Trajectory
- Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
- MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining Dynamics
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- Offline Reinforcement Learning with Generative Trajectory Policies
- Protenix-Mini+: efficient structure prediction model with scalable pairformer
- SoundReactor: Frame-level Online Video-to-Audio Generation
- Fine-Tuning Flow Matching via Maximum Likelihood Estimation of Reconstructions
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- FreqCa: Accelerating Diffusion Models via Frequency-Aware Caching
- Real-Time Motion-Controllable Autoregressive Video Diffusion
- IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation
- FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring
- Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
- SkipSR: Faster Super Resolution with Token Skipping
- Who Said Neural Networks Aren't Linear?
- Efficient High-Resolution Image Editing with Hallucination-Aware Loss and Adaptive Tiling
- SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
- Knowledge Distillation Detection for Open-weights Models
- MACS: Measurement-Aware Consistency Sampling for Inverse Problems
- Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- Distilled Protein Backbone Generation
- LVTINO: LAtent Video consisTency INverse sOlver for High Definition Video Restoration
- Riemannian Consistency Model
- Align Your Tangent: Training Better Consistency Models via Manifold-Aligned Tangents
- Collaborative-Distilled Diffusion Models (CDDM) for Accelerated and Lightweight Trajectory Prediction
- UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
- Diffusion Alignment as Variational Expectation-Maximization
- Continuously Augmented Discrete Diffusion model for Categorical Generative Modeling
- EVODiff: Entropy-aware Variance Optimized Diffusion Inference
- Swift: An Autoregressive Consistency Model for Efficient Weather Forecasting
- Hierarchical Reasoning Models: Perspectives and Misconceptions
- Flow Matching with Semidiscrete Couplings
- Score Distillation of Flow Matching Models
- CMT: Mid-Training for Efficient Learning of Consistency, Mean Flow, and Flow Map Models
- Asymmetric VAE for One-Step Video Super-Resolution Acceleration
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- Training Agents Inside of Scalable World Models
- OAT-FM: Optimal Acceleration Transport for Improved Flow Matching
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis
- Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution
- Foundation Model-Based Adaptive Semantic Image Transmission for Dynamic Wireless Environments
- Stochastic Interpolants via Conditional Dependent Coupling
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Transport Based Mean Flows for Generative Modeling
- Scale-Wise VAR is Secretly Discrete Diffusion
- HiGS: History-Guided Sampling for Plug-and-Play Enhancement of Diffusion Models
- Consistency Models as Plug-and-Play Priors for Inverse Problems
- Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training
- DistillKac: Few-Step Image Generation via Damped Wave Equations
- Score-based Idempotent Distillation of Diffusion Models
- SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
- Learnable Sampler Distillation for Discrete Diffusion Models
- Latent Iterative Refinement Flow: A Geometric-Constrained Approach for Few-Shot Generation
- RFMSR: Residual Flow Matching for Image Super-Resolution
- ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate
- ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
- A Gradient Flow Approach to Solving Inverse Problems with Latent Diffusion Models
- RadarSFD: Single-Frame Diffusion with Pretrained Priors for Radar Point Clouds
- Discrete-Time Diffusion-Like Models for Speech Synthesis
- Virtual Consistency for Audio Editing
- Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation
- Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
- Imagination at Inference: Synthesizing In-Hand Views for Robust Visuomotor Policy Inference
- LowDiff: Efficient Diffusion Sampling with Low-Resolution Condition
- MeanFlowSE: one-step generative speech enhancement via conditional mean flow
- Dual-Actor Fine-Tuning of VLA Models: A Talk-and-Tweak Human-in-the-Loop Approach
- Generative Consistency Models for Estimation of Kinetic Parametric Image Posteriors in Total-Body PET
- PINGS: Physics-Informed Neural Network for Fast Generative Sampling
- Robustifying Diffusion-Denoised Smoothing Against Covariate Shift
- HiCache: Training-free Acceleration of Diffusion Models via Hermite Polynomial-based Feature Caching
- Sig-DEG for Distillation: Making Diffusion Models Faster and Lighter
- CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
- PegasusFlow: Parallel Rolling-Denoising Score Sampling for Robot Diffusion Planner Flow Matching
- ArtifactGen: Benchmarking WGAN-GP vs Diffusion for Label-Aware EEG Artifact Synthesis
- A Generalisable Generative Model for Multi-Detector Calorimeter Simulation
- CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
- Continuous Audio Language Models
- Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
- Pretrained Diffusion Models Are Inherently Skipped-Step Samplers
- Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
- Transition Models: Rethinking the Generative Learning Objective
- Diffusion Generative Models Meet Compressed Sensing, with Applications to Imaging and Finance
- A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions
- OmniCache: A Trajectory-Oriented Global Perspective on Training-Free Cache Reuse for Diffusion Transformer Models
- Sem-RaDiff: Diffusion-Based 3D Radar Semantic Perception in Cluttered Agricultural Environments
- Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling
- Lipschitz-Guided Design of Interpolation Schedules in Generative Models
- A Continuous-Time Consistency Model for 3D Point Cloud Generation
- ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
- VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation
- Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
- Towards 6G Intelligence: The Role of Generative AI in Future Wireless Networks
- Energy-Based Flow Matching for Generating 3D Molecular Structure
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
- Matrix-game 2.0: An open-source real-time and streaming interactive world model
- Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
- Distribution Matching via Generalized Consistency Models
- Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark
- Belief-Conditioned One-Step Diffusion: Real-Time Trajectory Planning with Just-Enough Sensing
- Generic Event Boundary Detection via Denoising Diffusion
- TweezeEdit: Consistent and Efficient Image Editing with Path Regularization
- BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
- PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single Image
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- Prototype-Guided Diffusion: Visual Conditioning without External Memory
- Improving Diversity in Language Models: When Temperature Fails, Change the Loss
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
- DiffPose-Animal: A Language-Conditioned Diffusion Framework for Animal Pose Estimation
- Elucidating Rectified Flow with Deterministic Sampler: Polynomial Discretization Complexity for Multi and One-step Models
- Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
- Towards High-Order Mean Flow Generative Models: Feasibility, Expressivity, and Provably Efficient Criteria
- Elastic Diffusion Transformer
- Fewer Denoising Steps or Cheaper Per-Step Inference: Towards Compute-Optimal Diffusion Model Deployment
- Real-Time Iteration Scheme for Diffusion Policy
- GL-LCM: Global-Local Latent Consistency Models for Fast High-Resolution Bone Suppression in Chest X-Ray Images
- Forecasting When to Forecast: Accelerating Diffusion Models with Confidence-Gated Taylor
- Diffusion models for inverse problems
- On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
- DBLP: Noise Bridge Consistency Distillation For Efficient And Reliable Adversarial Purification
- Jet Image Generation in High Energy Physics Using Diffusion Models
- Video Generators are Robot Policies
- Adaptively Distilled ControlNet: Accelerated Training and Superior Sampling for Medical Image Synthesis
- ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning
- Exploiting Diffusion Prior for Task-driven Image Restoration
- Weighted Conditional Flow Matching
Related