vix.ing · top · new · best · stats

Information Theoretic Lower Bounds on Negative Log Likelihood

2019/04/12 by Luis Lastras, Luis A. Lastras, Lastras, Luis A. · 2 citations
Computer Science · Mathematics · #Advanced Image Processing Techniques #Algorithm #Applied mathematics #Bayesian probability #Computer science #Data compression #Distortion (music) #Distortion function #Empirical likelihood #Estimation theory #FOS: Computer and information sciences #Function (biology) #Generative Adversarial Networks and Image Synthesis #Information Theory (cs.IT) #Information theory #Latent variable #Latent variable model #Likelihood function #Lossy compression #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Marginal likelihood #Mathematics #Maximum likelihood #Prior probability #Rate–distortion theory #Statistical Methods and Inference #Statistics #Upper and lower bounds #Variable (mathematics) #cs.IT #cs.LG #math.IT #stat.ML

paper · pdf · doi:10.48550/arxiv.1904.06395

published in arXiv (Cornell University) (Cornell University)

arxiv created 2019/04/12 · openalex publication_date 2019/04/12 · arxiv updated 2019/04/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06

Abstract

In this article we use rate-distortion theory, a branch of information theory devoted to the problem of lossy compression, to shed light on an important problem in latent variable modeling of data: is there room to improve the model? One way to address this question is to find an upper bound on the probability (equivalently a lower bound on the negative log likelihood) that the model can assign to some data as one varies the prior and/or the likelihood function in a latent variable model. The core of our contribution is to formally show that the problem of optimizing priors in latent variable models is exactly an instance of the variational optimization problem that information theorists solve when computing rate-distortion functions, and then to use this to derive a lower bound on negative log likelihood. Moreover, we will show that if changing the prior can improve the log likelihood, then there is a way to change the likelihood function instead and attain the same log likelihood, and thus rate-distortion theory is of relevance to both optimizing priors as well as optimizing likelihood functions. We will experimentally argue for the usefulness of quantities derived from rate-distortion theory in latent variable modeling by applying them to a problem in image modeling.

Citations

Cited by

Related