vix.ing · top · new · best · stats · spec

Approximation Results for Gradient Descent trained Neural Networks

2023/09/09 by Gerrit Welper, Welper, G.
Computer Science · Engineering · Medicine · #41A46 #65K10 #68T07 #Advanced Neural Network Applications #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Medical Imaging and Analysis #Numerical Analysis (math.NA) #Radiomics and Machine Learning in Medical Imaging

paper · pdf · doi:10.48550/arxiv.2309.04860

openalex publication_date 2023/09/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The paper contains approximation guarantees for neural networks that are trained with gradient flow, with error measured in the continuous L2(\mathbbSd-1)-norm on the d-dimensional unit sphere and targets that are Sobolev smooth. The networks are fully connected of constant depth and increasing width. Although all layers are trained, the gradient flow convergence is based on a neural tangent kernel (NTK) argument for the non-convex second but last layer. Unlike standard NTK analysis, the continuous error norm implies an under-parametrized regime, possible by the natural smoothness assumption required for approximation. The typical over-parametrization re-enters the results in form of a loss in approximation rate relative to established approximation methods for Sobolev smooth functions.

Related