2022/04/24 by Chao Ma, Ma, Chao, Daniel Kunin +5 · 5 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · Physics and Astronomy · #Advanced Electron Microscopy Techniques and Applications #Advanced Mathematical Modeling in Engineering #Algorithm #Applied mathematics #Artificial intelligence #Artificial neural network #Computer science #Convexity #Function (biology) #Geometry #Gradient descent #Machine learning #Mathematical analysis #Mathematical optimization #Mathematics #Maxima and minima #Model Reduction and Neural Networks #Physics #Quadratic equation #Quadratic function #Quadratic growth #Stability (learning theory) #Statistical physics #cs.LG
paper · pdf · doi:10.48550/arxiv.2204.11326
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2022/04/24 · arxiv created 2022/06/22 · arxiv updated 2022/06/23 · openalex created_date 2022/12/14 · openalex updated_date 2026/08/06
A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks. Though, it usually holds in a very small neighborhood of the minimum, it cannot explain many phenomena observed during the optimization process. In this work, we study the structure of neural network loss functions and its implication on optimization in a region beyond the reach of a good quadratic approximation. Numerically, we observe that neural network loss functions possesses a multiscale structure, manifested in two ways: (1) in a neighborhood of minima, the loss mixes a continuum of scales and grows subquadratically, and (2) in a larger region, the loss shows several separate scales clearly. Using the subquadratic growth, we are able to explain the Edge of Stability phenomenon [5] observed for the gradient descent (GD) method. Using the separate scales, we explain the working mechanism of learning rate decay by simple examples. Finally, we study the origin of the multiscale structure and propose that the non-convexity of the models and the non-uniformity of training data is one of the causes. By constructing a two-layer neural network problem we show that training data with different magnitudes give rise to different scales of the loss function, producing subquadratic growth and multiple separate scales.