2020/11/06 by Constantin Christof, Christof, Constantin
Engineering · #Advanced Numerical Analysis Techniques #Tribology and Lubrication Engineering #Advanced machining processes and optimization
paper · pdf · doi:10.48550/arxiv.2011.03293
We study the optimization landscape and the stability properties of training\nproblems with squared loss for neural networks and general nonlinear conic\napproximation schemes. It is demonstrated that, if a nonlinear conic\napproximation scheme is considered that is (in an appropriately defined sense)\nmore expressive than a classical linear approximation approach and if there\nexist unrealizable label vectors, then a training problem with squared loss is\nnecessarily unstable in the sense that its solution set depends discontinuously\non the label vector in the training data. We further prove that the same\neffects that are responsible for these instability properties are also the\nreason for the emergence of saddle points and spurious local minima, which may\nbe arbitrarily far away from global solutions, and that neither the instability\nof the training problem nor the existence of spurious local minima can, in\ngeneral, be overcome by adding a regularization term to the objective function\nthat penalizes the size of the parameters in the approximation scheme. The\nlatter results are shown to be true regardless of whether the assumption of\nrealizability is satisfied or not. We demonstrate that our analysis in\nparticular applies to training problems for free-knot interpolation schemes and\ndeep and shallow neural networks with variable widths that involve an arbitrary\nmixture of various activation functions (e.g., binary, sigmoid, tanh, arctan,\nsoft-sign, ISRU, soft-clip, SQNL, ReLU, leaky ReLU, soft-plus, bent identity,\nSILU, ISRLU, and ELU). In summary, the findings of this paper illustrate that\nthe improved approximation properties of neural networks and general nonlinear\nconic approximation instruments are linked in a direct and quantifiable way to\nundesirable properties of the optimization problems that have to be solved in\norder to train them.\n