2021/06/05 by Chin-Wei Huang, Huang, Chin-Wei, Jae Hyun Lim +3 · 31 citations
Computer Science · Physics and Astronomy · #Generative Adversarial Networks and Image Synthesis #Machine Learning in Healthcare #Model Reduction and Neural Networks #cs.LG
paper · pdf · doi:10.48550/arxiv.2106.02808
arxiv created 2021/09/30 · arxiv updated 2021/10/01
Discrete-time diffusion-based generative models and score matching methods have shown promising results in modeling high-dimensional image data. Recently, Song et al. (2021) show that diffusion processes that transform data into noise can be reversed via learning the score function, i.e. the gradient of the log-density of the perturbed data. They propose to plug the learned score function into an inverse formula to define a generative diffusion process. Despite the empirical success, a theoretical underpinning of this procedure is still lacking. In this work, we approach the (continuous-time) generative diffusion directly and derive a variational framework for likelihood estimation, which includes continuous-time normalizing flows as a special case, and can be seen as an infinitely deep variational autoencoder. Under this framework, we show that minimizing the score-matching loss is equivalent to maximizing a lower bound of the likelihood of the plug-in reverse SDE proposed by Song et al. (2021), bridging the theoretical gap.