2018/01/07 by Chiyuan Zhang, Qianli Liao, Zhang, Chiyuan +9 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Stochastic Gradient Optimization Techniques #cs.LG
paper · pdf · doi:10.48550/arxiv.1801.02254
arxiv created 2018/01/07 · openalex publication_date 2018/01/07 · arxiv updated 2018/01/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In Theory IIb we characterize with a mix of theory and experiments the optimization of deep convolutional networks by Stochastic Gradient Descent. The main new result in this paper is theoretical and experimental evidence for the following conjecture about SGD: SGD concentrates in probability -- like the classical Langevin equation -- on large volume, "flat" minima, selecting flat minimizers which are with very high probability also global minimizers