vix.ing · top · new · best · stats

Approximation Results for Gradient Descent trained Neural Networks

2023/09/09 by Gerrit Welper, G. Welper, Welper, G. · 1 voice · 1 citation
Computer Science · Engineering · Mathematics · Medicine · #41A46 #65K10 #68T07 #Advanced Neural Network Applications #Applied mathematics #Approximation error #Artificial intelligence #Artificial neural network #Balanced flow #Computer science #FOS: Computer and information sciences #FOS: Mathematics #Geometry #Gradient descent #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical analysis #Mathematics #Medical Imaging and Analysis #Norm (philosophy) #Numerical Analysis (math.NA) #Radiomics and Machine Learning in Medical Imaging #Rate of convergence #Smoothness #Sobolev space #Tangent #Uniform norm #Unit sphere #cs.LG #math.NA #stat.ML

paper · pdf · doi:10.48550/arxiv.2309.04860

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2023/09/09 · arxiv published 2023/09/09 · arxiv updated 2023/09/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The paper contains approximation guarantees for neural networks that are trained with gradient flow, with error measured in the continuous L2(\mathbbSd-1)-norm on the d-dimensional unit sphere and targets that are Sobolev smooth. The networks are fully connected of constant depth and increasing width. Although all layers are trained, the gradient flow convergence is based on a neural tangent kernel (NTK) argument for the non-convex second but last layer. Unlike standard NTK analysis, the continuous error norm implies an under-parametrized regime, possible by the natural smoothness assumption required for approximation. The typical over-parametrization re-enters the results in form of a loss in approximation rate relative to established approximation methods for Sobolev smooth functions.

Cited by

Discussions

Related