2022/09/02 by Stephan Wojtowytsch, Wojtowytsch, Stephan · 2 citations
Computer Science · #41A30 #65D40 #68T07 #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Neural Networks and Applications #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2209.01173
openalex publication_date 2022/09/02 · openalex created_date 2022/09/06 · openalex updated_date 2026/07/28
In this note, we study how neural networks with a single hidden layer and ReLU activation interpolate data drawn from a radially symmetric distribution with target labels 1 at the origin and 0 outside the unit ball, if no labels are known inside the unit ball. With weight decay regularization and in the infinite neuron, infinite data limit, we prove that a unique radially symmetric minimizer exists, whose weight decay regularizer and Lipschitz constant grow as d and √(d) respectively. We furthermore show that the weight decay regularizer grows exponentially in d if the label 1 is imposed on a ball of radius ε rather than just at the origin. By comparison, a neural networks with two hidden layers can approximate the target function without encountering the curse of dimensionality.