vix.ing · top · new · best · stats

On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime

2020/05/27 by Stephan Wojtowytsch, Wojtowytsch, Stephan · 6 citations
Computer Science · Mathematics · #35F20 #35Q68 #49Q22 #68T07 #Analysis of PDEs (math.AP) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Neural Networks and Applications #Stochastic Gradient Optimization Techniques #cs.LG #math.AP #msc:35F20 #msc:35Q68 #msc:49Q22 #msc:68T07 #stat.ML

paper · pdf · doi:10.48550/arxiv.2005.13530

arxiv created 2020/05/27 · openalex publication_date 2020/05/27 · arxiv updated 2020/05/28 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

We describe a necessary and sufficient condition for the convergence to minimum Bayes risk when training two-layer ReLU-networks by gradient descent in the mean field regime with omni-directional initial parameter distribution. This article extends recent results of Chizat and Bach to ReLU-activated networks and to the situation in which there are no parameters which exactly achieve MBR. The condition does not depend on the initalization of parameters and concerns only the weak convergence of the realization of the neural network, not its parameter distribution.

Citations

Cited by

Related