vix.ing · top · new · best · stats · spec

Best k-layer neural network approximations

2019/07/02 by Lek‐Heng Lim, Lim, Lek-Heng, Mateusz Michałek +3 · 2 citations
Mathematics · Physics and Astronomy · #41A30 #41A50 #92B20 #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical Approximation and Integration #Model Reduction and Neural Networks #Tensor decomposition and applications

paper · pdf · doi:10.48550/arxiv.1907.01507

openalex publication_date 2019/07/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We show that the empirical risk minimization (ERM) problem for neural networks has no solution in general. Given a training set s1, …, sn ∈ ℝp with corresponding responses t1,…,tn ∈ ℝq, fitting a k-layer neural network νθ: ℝp → ℝq involves estimation of the weights θ∈ ℝm via an ERM: infθ∈ ℝmi=1n ‖ ti - νθ(si) ‖22. We show that even for k = 2, this infimum is not attainable in general for common activations like ReLU, hyperbolic tangent, and sigmoid functions. A high-level explanation is like that for the nonexistence of best rank-r approximations of higher-order tensors --- the set of parameters is not a closed set --- but the geometry involved for best k-layer neural networks approximations is more subtle. In addition, we show that for smooth activations σ(x)= 1/(1 + exp(-x)) and σ(x)=\tanh(x), such failure to attain an infimum can happen on a positive-measured subset of responses. For the ReLU activation σ(x)=max(0,x), we completely classifying cases where the ERM for a best two-layer neural network approximation attains its infimum. As an aside, we obtain a precise description of the geometry of the space of two-layer neural networks with d neurons in the hidden layer: it is the join locus of a line and the d-secant locus of a cone.

Citations

Cited by

Related