vix.ing · top · new · best · stats · spec

Implicit Regularization of Discrete Gradient Dynamics in Linear Neural\n Networks

2019/04/30 by Gauthier Gidel, Francis Bach, Gidel, Gauthier +3 · 9 citations
Computer Science · Engineering · #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Optimization and Control (math.OC) #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques

paper · pdf · doi:10.48550/arxiv.1904.13262

openalex publication_date 2019/04/30 · openalex created_date 2022/07/29 · openalex updated_date 2026/07/28

Abstract

When optimizing over-parameterized models, such as deep neural networks, a\nlarge set of parameters can achieve zero training error. In such cases, the\nchoice of the optimization algorithm and its respective hyper-parameters\nintroduces biases that will lead to convergence to specific minimizers of the\nobjective. Consequently, this choice can be considered as an implicit\nregularization for the training of over-parametrized models. In this work, we\npush this idea further by studying the discrete gradient dynamics of the\ntraining of a two-layer linear network with the least-squares loss. Using a\ntime rescaling, we show that, with a vanishing initialization and a small\nenough step size, this dynamics sequentially learns the solutions of a\nreduced-rank regression with a gradually increasing rank.\n

Citations

Cited by

Related