2021/06/30 by Arthur Paul Jacot, Arthur Jacot, François Ged +9 · 10 citations
Computer Science · Mathematics · Physics and Astronomy · #Artificial intelligence #Computer science #Dynamics (music) #FOS: Computer and information sciences #Geometry #Initialization #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical optimization #Mathematics #Meteorology #Physics #Quantum chaos and dynamical systems #Saddle #Saddle point #Statistical Mechanics and Entropy #Statistical physics #Symmetry (geometry) #Training (meteorology) #cs.LG #stat.ML #stochastic dynamics and bifurcation
paper · pdf · doi:10.48550/arxiv.2106.15933
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2021/06/30 · arxiv created 2022/01/31 · arxiv updated 2022/02/01 · openalex created_date 2022/07/24 · openalex updated_date 2026/08/06
The dynamics of Deep Linear Networks (DLNs) is dramatically affected by the variance σ2 of the parameters at initialization θ0. For DLNs of width w, we show a phase transition w.r.t. the scaling γ of the variance σ2=w-γ as w→∞: for large variance (γ<1), θ0 is very close to a global minimum but far from any saddle point, and for small variance (γ>1), θ0 is close to a saddle point and far from any global minimum. While the first case corresponds to the well-studied NTK regime, the second case is less understood. This motivates the study of the case γ→ +∞, where we conjecture a Saddle-to-Saddle dynamics: throughout training, gradient descent visits the neighborhoods of a sequence of saddles, each corresponding to linear maps of increasing rank, until reaching a sparse global minimum. We support this conjecture with a theorem for the dynamics between the first two saddles, as well as some numerical experiments.