vix.ing · top · new · best · stats

DiffWave: A Versatile Diffusion Model for Audio Synthesis

2020/09/21 by Zhifeng Kong, Wei Ping, Kong, Zhifeng +7 · 235 citations
Computer Science · Engineering · Mathematics · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.CL #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering #stat.ML

paper · pdf · doi:10.48550/arxiv.2009.09761

ICLR 2021 (oral)

openalex publication_date 2020/09/21 · arxiv created 2021/03/30 · arxiv updated 2021/04/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation. The model is non-autoregressive, and converts the white noise signal into structured waveform through a Markov chain with a constant number of steps at synthesis. It is efficiently trained by optimizing a variant of variational bound on the data likelihood. DiffWave produces high-fidelity audios in different waveform generation tasks, including neural vocoding conditioned on mel spectrogram, class-conditional generation, and unconditional generation. We demonstrate that DiffWave matches a strong WaveNet vocoder in terms of speech quality (MOS: 4.44 versus 4.43), while synthesizing orders of magnitude faster. In particular, it significantly outperforms autoregressive and GAN-based waveform models in the challenging unconditional generation task in terms of audio quality and sample diversity from various automatic and human evaluations.

Citations

Cited by

Related