2019/08/30 by Benedict Leimkuhler, Charles E. Matthews, Leimkuhler, Benedict +3
Computer Science · Materials Science · Physics and Astronomy · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Materials Science #Model Reduction and Neural Networks #Neural Networks and Applications
paper · pdf · doi:10.48550/arxiv.1908.11843
openalex publication_date 2019/08/30 · openalex created_date 2022/09/28 · openalex updated_date 2026/07/28
Traditionally, neural networks are parameterized using optimization\nprocedures such as stochastic gradient descent, RMSProp and ADAM. These\nprocedures tend to drive the parameters of the network toward a local minimum.\nIn this article, we employ alternative "sampling" algorithms (referred to here\nas "thermodynamic parameterization methods") which rely on discretized\nstochastic differential equations for a defined target distribution on\nparameter space. We show that the thermodynamic perspective already improves\nneural network training. Moreover, by partitioning the parameters based on\nnatural layer structure we obtain schemes with very rapid convergence for data\nsets with complicated loss landscapes.\n We describe easy-to-implement hybrid partitioned numerical algorithms, based\non discretized stochastic differential equations, which are adapted to\nfeed-forward neural networks, including a multi-layer Langevin algorithm,\nAdLaLa (combining the adaptive Langevin and Langevin algorithms) and LOL\n(combining Langevin and Overdamped Langevin); we examine the convergence of\nthese methods using numerical studies and compare their performance among\nthemselves and in relation to standard alternatives such as stochastic gradient\ndescent and ADAM. We present evidence that thermodynamic parameterization\nmethods can be (i) faster, (ii) more accurate, and (iii) more robust than\nstandard algorithms used within machine learning frameworks.\n