vix.ing · top · new · best · stats · spec

SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures

2025/05/14 by Julian Kranz, Davide Gallon, Kranz, Julian +5 · 4 citations
Computer Science · Physics and Astronomy · #03C64 #03C98 #26B40 #FOS: Computer and information sciences #FOS: Mathematics #Logic (math.LO) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Neural Networks and Applications #Optimization and Control (math.OC) #Primary 68T05 #Secondary 68T07

paper · pdf · doi:10.48550/arxiv.2505.09572

openalex publication_date 2025/05/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We study gradient flows for loss landscapes of fully connected feedforward neural networks with commonly used continuously differentiable activation functions such as the logistic, hyperbolic tangent, softplus or GELU function. We prove that the gradient flow either converges to a critical point or diverges to infinity while the loss converges to an asymptotic critical value. Moreover, we prove the existence of a threshold ε>0 such that the loss value of any gradient flow initialized at most ε above the optimal level converges to it. For polynomial target functions and sufficiently big architecture and data set, we prove that the optimal loss value is zero and can only be realized asymptotically. From this setting, we deduce our main result that any gradient flow with sufficiently good initialization diverges to infinity. Our proof heavily relies on the geometry of o-minimal structures. We confirm these theoretical findings with numerical experiments and extend our investigation to more realistic scenarios, where we observe an analogous behavior.

Citations

Cited by

Related