vix.ing · top · new · best · stats

On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective

2021/12/02 by Xiaowu Dai, Dai, Xiaowu, Yuhua Zhu +1 · 4 citations
Computer Science · Mathematics · Physics and Astronomy · #Advanced Thermodynamics and Statistical Mechanics #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Markov Chains and Monte Carlo Methods #Statistics Theory (math.ST) #Stochastic Gradient Optimization Techniques #cs.LG #math.ST #stat.TH

paper · pdf · doi:10.48550/arxiv.2112.00987

arxiv created 2021/12/02 · openalex publication_date 2021/12/02 · arxiv updated 2021/12/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We study the statistical properties of the dynamic trajectory of stochastic gradient descent (SGD). We approximate the mini-batch SGD and the momentum SGD as stochastic differential equations (SDEs). We exploit the continuous formulation of SDE and the theory of Fokker-Planck equations to develop new results on the escaping phenomenon and the relationship with large batch and sharp minima. In particular, we find that the stochastic process solution tends to converge to flatter minima regardless of the batch size in the asymptotic regime. However, the convergence rate is rigorously proven to depend on the batch size. These results are validated empirically with various datasets and models.

Citations

Cited by

Related