2020/01/30 by Emiel Hoogeboom, Taco S. Cohen, Taco Cohen +4 · 16 citations
Computer Science · Mathematics · #Artificial intelligence #Autoregressive model #Combinatorics #Computer science #Dimension (graph theory) #Distribution (mathematics) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical analysis #Mathematics #Music and Audio Processing #Noise (video) #Statistics #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2001.11235
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/01/30 · openalex publication_date 2020/01/30 · arxiv updated 2020/01/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Media is generally stored digitally and is therefore discrete. Many successful deep distribution models in deep learning learn a density, i.e., the distribution of a continuous random variable. Naïve optimization on discrete data leads to arbitrarily high likelihoods, and instead, it has become standard practice to add noise to datapoints. In this paper, we present a general framework for dequantization that captures existing methods as a special case. We derive two new dequantization objectives: importance-weighted (iw) dequantization and Rényi dequantization. In addition, we introduce autoregressive dequantization (ARD) for more flexible dequantization distributions. Empirically we find that iw and Rényi dequantization considerably improve performance for uniform dequantization distributions. ARD achieves a negative log-likelihood of 3.06 bits per dimension on CIFAR10, which to the best of our knowledge is state-of-the-art among distribution models that do not require autoregressive inverses for sampling.