vix.ing · top · new · best · stats · spec

Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis

2021/06/04 by Stephan Wojtowytsch, Wojtowytsch, Stephan · 2 citations
Computer Science · Mathematics · Physics and Astronomy · #35K65 #60H30 #Analysis of PDEs (math.AP) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Markov Chains and Monte Carlo Methods #Model Reduction and Neural Networks #Neural Networks and Applications #Primary: 90C26 #Secondary: 68T07

paper · pdf · doi:10.48550/arxiv.2106.02588

openalex publication_date 2021/06/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The representation of functions by artificial neural networks depends on a large number of parameters in a non-linear fashion. Suitable parameters of these are found by minimizing a 'loss functional', typically by stochastic gradient descent (SGD) or an advanced SGD-based algorithm. In a continuous time model for SGD with noise that follows the 'machine learning scaling', we show that in a certain noise regime, the optimization algorithm prefers 'flat' minima of the objective function in a sense which is different from the flat minimum selection of continuous time SGD with homogeneous noise.

Cited by

Related