Diffusion Models in Vision: A Survey
2023/03/27 by Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu +1 · 116 citations
Computer Science · Medicine · Physics and Astronomy · #Advanced Neuroimaging Techniques and Applications #Generative Adversarial Networks and Image Synthesis #Model Reduction and Neural Networks
paper · doi:10.1109/tpami.2023.3261988
openalex publication_date 2023/03/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Denoising diffusion models represent a recent emerging topic in computer vision, demonstrating remarkable results in the area of generative modeling. A diffusion model is a deep generative model that is based on two stages, a forward diffusion stage and a reverse diffusion stage. In the forward diffusion stage, the input data is gradually perturbed over several steps by adding Gaussian noise. In the reverse stage, a model is tasked at recovering the original input data by learning to gradually reverse the diffusion process, step by step. Diffusion models are widely appreciated for the quality and diversity of the generated samples, despite their known computational burdens, i.e., low speeds due to the high number of steps involved during sampling. In this survey, we provide a comprehensive review of articles on denoising diffusion models applied in vision, comprising both theoretical and practical contributions in the field. First, we identify and present three generic diffusion modeling frameworks, which are based on denoising diffusion probabilistic models, noise conditioned score networks, and stochastic differential equations. We further discuss the relations between diffusion models and other deep generative models, including variational auto-encoders, generative adversarial networks, energy-based models, autoregressive models and normalizing flows. Then, we introduce a multi-perspective categorization of diffusion models applied in computer vision. Finally, we illustrate the current limitations of diffusion models and envision some interesting directions for future research.
Citations
Cited by
- A Survey of Multimodal Controllable Diffusion Models
- LGCC: Enhancing Flow Matching Based Text-Guided Image Editing with Local Gaussian Coupling and Context Consistency
- Future of AI Models: A Computational perspective on Model collapse
- Through the Lens: Benchmarking Deepfake Detectors Against Moiré-Induced Distortions
- Assessing and Understanding Creativity in Large Language Models
- Scalable Machine Learning Analysis of Parker Solar Probe Solar Wind Data
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers
- Attentive Convolution: Unifying the Expressivity of Self-Attention with Convolutional Efficiency
- Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
- CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation
- Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
- Redundancy as a Structural Information Principle for Learning and Generalization
- Catalyst GFlowNet for electrocatalyst design: A hydrogen evolution reaction case study
- FusionGen: Feature Fusion-Based Few-Shot EEG Data Generation
- Denoised Diffusion for Object-Focused Image Augmentation
- RadioFlow: Efficient Radio Map Construction Framework with Flow Matching
- Permutation-Invariant Spectral Learning via Dyson Diffusion
- Hyperspectral data augmentation with transformer-based diffusion models
- Deep Neural Networks Inspired by Differential Equations
- MONKEY: Masking ON KEY-Value Activation Adapter for Personalization
- EigenScore: OOD Detection using Covariance in Diffusion Models
- A Dynamic Mode Decomposition Approach to Morphological Component Analysis
- Rasterized Steered Mixture of Experts for Efficient 2D Image Regression
- LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- From 2D to 3D, Deep Learning-based Shape Reconstruction in Magnetic Resonance Imaging: A Review
- Translation from Wearable PPG to 12-Lead ECG
- From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis
- Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers
- On the design space between molecular mechanics and machine learning force fields
- Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models
- Comparative Analysis of GAN and Diffusion for MRI-to-CT translation
- Differential-Integral Neural Operator for Long-Term Turbulence Forecasting
- Structure-Attribute Transformations with Markov Chain Boost Graph Domain Adaptation
- Responsible Diffusion: A Comprehensive Survey on Safety, Ethics, and Trust in Diffusion Models
- An Efficient Conditional Score-based Filter for High Dimensional Nonlinear Filtering Problems
- Sobolev acceleration for neural networks
- The future of machine learning for small-molecule drug discovery will be driven by data
- Gener anno : A Genomic Foundation Model for Metagenomic Annotation
- AGSwap: Overcoming Category Boundaries in Object Fusion via Adaptive Group Swapping
- SocialTraj: Two-Stage Socially-Aware Trajectory Prediction for Autonomous Driving via Conditional Diffusion Model
- PRNU-Bench: A Novel Benchmark and Model for PRNU-Based Camera Identification
- A Mutil-conditional Diffusion Transformer for Versatile Seismic Wave Generation
- Discrete Diffusion Models: Novel Analysis and New Sampler Guarantees
- CIDER: A Causal Cure for Brand-Obsessed Text-to-Image Models
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- DiffVL: Diffusion-Based Visual Localization on 2D Maps via BEV-Conditioned GPS Denoising
- AoSRNet: All-in-One Scene Recovery Networks via multi-knowledge integration
- DTGen: Generative Diffusion-Based Few-Shot Data Augmentation for Fine-Grained Dirty Tableware Recognition
- Dual Orthogonal Guidance for Robust Diffusion-based Handwritten Text Generation
- Every Camera Effect, Every Time, All at Once: 4D Gaussian Ray Tracing for Physics-based Camera Effect Data Generation
- A Differentiable Surrogate Model for the Generation of Radio Pulses from In-Ice Neutrino Interactions
- FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution
- Fitting Image Diffusion Models on Video Datasets
- A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions
- A Data-Centric Approach to Pedestrian Attribute Recognition: Synthetic Augmentation via Prompt-driven Diffusion Models
- Efficient Diffusion Models: A Comprehensive Survey From Principles to Practices
- MedGR2: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning
- MicroLad: 2D-to-3D Microstructure Reconstruction and Generation via Latent Diffusion and Score Distillation
- Enhancing Novel View Synthesis from extremely sparse views with SfM-free 3D Gaussian Splatting Framework
- A real-world underwater turbid image enhancement benchmark and beyond
- TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
- DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion Models
- Denoising diffusion models for inverse design of inflatable structures with programmable deformations
- Efficient Image-to-Image Schrödinger Bridge for CT Field of View Extension
- AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- Envisioning Generative Artificial Intelligence in Cartography and Mapmaking
- CD-TVD: Contrastive Diffusion for 3D Super-Resolution with Scarce High-Resolution Time-Varying Data
- LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
- Causal Negative Sampling via Diffusion Model for Out-of-Distribution Recommendation
- Neural Bridge Processes
- Scalable Swin Transformer network for brain tumor segmentation from incomplete MRI modalities
- ContourDiff: Unpaired Medical Image Translation with Structural Consistency
- RAIDX: A Retrieval-Augmented Generation and GRPO Reinforcement Learning Framework for Explainable Deepfake Detection
- VideoGuard: Protecting Video Content from Unauthorized Editing
- READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
- BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models
- MissDDIM: Deterministic and Efficient Conditional Diffusion for Tabular Data Imputation
- Generative AI-Empowered Secure Communications in Space-Air-Ground Integrated Networks: A Survey and Tutorial
- Tackling Ill-posedness of Reversible Image Conversion with Well-posed Invertible Network
- Boosting Generalization Performance in Model-Heterogeneous Federated Learning Using Variational Transposed Convolution
- Diffusion Models for Future Networks and Communications: A Comprehensive Survey
- Semantically-Guided Inference for Conditional Diffusion Models: Enhancing Covariate Consistency in Time Series Forecasting
- Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?
- Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models
- Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
- A Contrastive Diffusion-based Network (CDNet) for Time Series Classification
- CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration
- A Self-training Framework for Semi-supervised Pulmonary Vessel Segmentation and Its Application in COPD
- A Comprehensive Review of Diffusion Models in Smart Agriculture: Progress, Applications, and Challenges
- Generative AI-Driven High-Fidelity Human Motion Simulation
- Leave No One Behind: Fairness-Aware Cross-Domain Recommender Systems for Non-Overlapping Users
- Flow Matching Meets Biology and Life Science: A Survey
- Hyperbolic Deep Learning for Foundation Models: A Survey
- LSDM: LLM-Enhanced Spatio-temporal Diffusion Model for Service-Level Mobile Traffic Prediction
- DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
- RadioDiff-3D: A 3D×3D Radio Map Dataset and Generative Diffusion Based Benchmark for 6G Environment-Aware Communication
- DeepShade: Enable Shade Simulation by Text-conditioned Image Generation
- A Review of Generative AI in Aquaculture: Foundations, Applications, and Future Directions for Smart and Sustainable Farming
- Ocean Diviner: A Diffusion-Augmented Reinforcement Learning Framework for AUV Robust Control in Underwater Tasks
- Try Harder: Hard Sample Generation and Learning for Clothes-Changing Person Re-ID
- Shortening the Trajectories: Identity-Aware Gaussian Approximation for Efficient 3D Molecular Generation
- Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
- AirScape: An Aerial Generative World Model with Motion Controllability
- Diffusion-Guided Knowledge Distillation for Weakly-Supervised Low-Light Semantic Segmentation
- DiffNMR: Diffusion Models for Nuclear Magnetic Resonance Spectra Elucidation
- FedDifRC: Unlocking the Potential of Text-to-Image Diffusion Models in Heterogeneous Federated Learning
- What You Have is What You Track: Adaptive and Robust Multimodal Tracking
- Brownian Bridge Diffusion for Sequential Recommendation
- Cloud Diffusion Part 1: Theory and Motivation
- Machine Learning in Acoustics: A Review and Open-Source Repository
- MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation
Related