vix.ing · top · new · best · stats

On the Distributional Properties of Adaptive Gradients

2021/05/15 by Zhang Zhiyi, Liu Ziyin, Zhiyi, Zhang +1 · 4 citations
Computer Science · Mathematics · #Applied mathematics #Artificial intelligence #Artificial neural network #Biology #Bounded function #Computer science #Current (fluid) #Distribution (mathematics) #Divergence (linguistics) #Econometrics #Economics #FOS: Computer and information sciences #Function (biology) #Gaussian Processes and Bayesian Inference #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical analysis #Mathematical optimization #Mathematics #Neural Networks and Applications #Physics #Series (stratigraphy) #Statistical physics #Variance (accounting) #Work (physics) #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.2105.07222

published in arXiv (Cornell University) (Cornell University)

arxiv created 2021/05/15 · openalex publication_date 2021/05/15 · arxiv updated 2021/05/18 · openalex created_date 2021/05/24 · openalex updated_date 2026/07/28

Abstract

Adaptive gradient methods have achieved remarkable success in training deep neural networks on a wide variety of tasks. However, not much is known about the mathematical and statistical properties of this family of methods. This work aims at providing a series of theoretical analyses of its statistical properties justified by experiments. In particular, we show that when the underlying gradient obeys a normal distribution, the variance of the magnitude of the update is an increasing and bounded function of time and does not diverge. This work suggests that the divergence of variance is not the cause of the need for warm up of the Adam optimizer, contrary to what is believed in the current literature.

Related