2019/07/07 by Xiang Cheng, Dong Yin, Cheng, Xiang +5 · 1 citation
Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Markov Chains and Monte Carlo Methods #Neural Networks and Applications #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.1907.03215
openalex publication_date 2019/07/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We prove quantitative convergence rates at which discrete Langevin-like processes converge to the invariant distribution of a related stochastic differential equation. We study the setup where the additive noise can be non-Gaussian and state-dependent and the potential function can be non-convex. We show that the key properties of these processes depend on the potential function and the second moment of the additive noise. We apply our theoretical findings to studying the convergence of Stochastic Gradient Descent (SGD) for non-convex problems and corroborate them with experiments using SGD to train deep neural networks on the CIFAR-10 dataset.