vix.ing · top · new · best · stats · spec

The Information Bottleneck Problem and Its Applications in Machine\n Learning

2020/04/30 by Ziv Goldfeld, Yury Polyanskiy, Goldfeld, Ziv +1 · 15 citations
Computer Science · #Age of Information Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Stochastic Gradient Optimization Techniques

paper · pdf · doi:10.48550/arxiv.2004.14941

openalex publication_date 2020/04/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Inference capabilities of machine learning (ML) systems skyrocketed in recent\nyears, now playing a pivotal role in various aspect of society. The goal in\nstatistical learning is to use data to obtain simple algorithms for predicting\na random variable Y from a correlated observation X. Since the dimension of\nX is typically huge, computationally feasible solutions should summarize it\ninto a lower-dimensional feature vector T, from which Y is predicted. The\nalgorithm will successfully make the prediction if T is a good proxy of Y,\ndespite the said dimensionality-reduction. A myriad of ML algorithms (mostly\nemploying deep learning (DL)) for finding such representations T based on\nreal-world data are now available. While these methods are often effective in\npractice, their success is hindered by the lack of a comprehensive theory to\nexplain it. The information bottleneck (IB) theory recently emerged as a bold\ninformation-theoretic paradigm for analyzing DL systems. Adopting mutual\ninformation as the figure of merit, it suggests that the best representation\nT should be maximally informative about Y while minimizing the mutual\ninformation with X. In this tutorial we survey the information-theoretic\norigins of this abstract principle, and its recent impact on DL. For the\nlatter, we cover implications of the IB problem on DL theory, as well as\npractical algorithms inspired by it. Our goal is to provide a unified and\ncohesive description. A clear view of current knowledge is particularly\nimportant for further leveraging IB and other information-theoretic ideas to\nstudy DL models.\n

Cited by

Related