vix.ing · top · new · best · stats · spec

Global Convergence of Three-layer Neural Networks in the Mean Field\n Regime

2021/05/11 by Huy Tuan Pham, Pham, Huy Tuan, Phan-Minh Nguyen +1 · 2 citations
Computer Science · Mathematics · Physics and Astronomy · #FOS: Computer and information sciences #FOS: Mathematics #FOS: Physical sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Markov Chains and Monte Carlo Methods #Model Reduction and Neural Networks #Statistical Mechanics (cond-mat.stat-mech) #Statistics Theory (math.ST) #Stochastic Gradient Optimization Techniques

paper · pdf · doi:10.48550/arxiv.2105.05228

openalex publication_date 2021/05/11 · openalex created_date 2021/05/24 · openalex updated_date 2026/07/28

Abstract

In the mean field regime, neural networks are appropriately scaled so that as\nthe width tends to infinity, the learning dynamics tends to a nonlinear and\nnontrivial dynamical limit, known as the mean field limit. This lends a way to\nstudy large-width neural networks via analyzing the mean field limit. Recent\nworks have successfully applied such analysis to two-layer networks and\nprovided global convergence guarantees. The extension to multilayer ones\nhowever has been a highly challenging puzzle, and little is known about the\noptimization efficiency in the mean field regime when there are more than two\nlayers.\n In this work, we prove a global convergence result for unregularized\nfeedforward three-layer networks in the mean field regime. We first develop a\nrigorous framework to establish the mean field limit of three-layer networks\nunder stochastic gradient descent training. To that end, we propose the idea of\na \neuronal embedding, which comprises of a fixed probability space\nthat encapsulates neural networks of arbitrary sizes. The identified mean field\nlimit is then used to prove a global convergence guarantee under suitable\nregularity and convergence mode assumptions, which -- unlike previous works on\ntwo-layer networks -- does not rely critically on convexity. Underlying the\nresult is a universal approximation property, natural of neural networks, which\nimportantly is shown to hold at \any finite training time (not\nnecessarily at convergence) via an algebraic topology argument.\n

Citations

Cited by

Related