2021/08/26 by Max W. Y. Lam, Lam, Max W. Y., Jun Wang +7 · 2 citations
Computer Science · Engineering · Physics and Astronomy · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Music and Audio Processing #Signal Processing (eess.SP) #Sound (cs.SD) #cs.AI #cs.LG #cs.SD #eess.AS #eess.SP #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2108.11514
openalex publication_date 2021/08/26 · arxiv created 2021/09/14 · arxiv updated 2021/09/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Denoising diffusion probabilistic models (DDPMs) have emerged as competitive generative models yet brought challenges to efficient sampling. In this paper, we propose novel bilateral denoising diffusion models (BDDMs), which take significantly fewer steps to generate high-quality samples. From a bilateral modeling objective, BDDMs parameterize the forward and reverse processes with a score network and a scheduling network, respectively. We show that a new lower bound tighter than the standard evidence lower bound can be derived as a surrogate objective for training the two networks. In particular, BDDMs are efficient, simple-to-train, and capable of further improving any pre-trained DDPM by optimizing the inference noise schedules. Our experiments demonstrated that BDDMs can generate high-fidelity samples with as few as 3 sampling steps and produce comparable or even higher quality samples than DDPMs using 1000 steps with only 16 sampling steps (a 62x speedup).