vix.ing · top · new · best · stats

The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization

2021/06/14 by Daniel LeJeune, Hamid Javadi, LeJeune, Daniel +3
Computer Science · Engineering · Mathematics · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.2106.07769

19 pages, 2 figures. Appeared in NeurIPS 2021. Small typographical correction

openalex publication_date 2021/06/14 · arxiv created 2022/01/03 · arxiv updated 2022/01/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Among the most successful methods for sparsifying deep (neural) networks are those that adaptively mask the network weights throughout training. By examining this masking, or dropout, in the linear case, we uncover a duality between such adaptive methods and regularization through the so-called "η-trick" that casts both as iteratively reweighted optimizations. We show that any dropout strategy that adapts to the weights in a monotonic way corresponds to an effective subquadratic regularization penalty, and therefore leads to sparse solutions. We obtain the effective penalties for several popular sparsification strategies, which are remarkably similar to classical penalties commonly used in sparse optimization. Considering variational dropout as a case study, we demonstrate similar empirical behavior between the adaptive dropout method and classical methods on the task of deep network sparsification, validating our theory.

Citations

Related