2018/06/21 by Philipp Petersen, Petersen, Philipp, Mones Raslan +3 · 5 citations
Computer Science · #52A30 #54H99 #68T05 #FOS: Mathematics #Functional Analysis (math.FA) #General Topology (math.GN) #Machine Learning and ELM #Neural Networks and Applications #Topological and Geometric Data Analysis
paper · pdf · doi:10.48550/arxiv.1806.08459
openalex publication_date 2018/06/21 · openalex created_date 2022/09/27 · openalex updated_date 2026/07/28
We analyze the topological properties of the set of functions that can be\nimplemented by neural networks of a fixed size. Surprisingly, this set has many\nundesirable properties. It is highly non-convex, except possibly for a few\nexotic activation functions. Moreover, the set is not closed with respect to\nLp-norms, 0 < p < \∞, for all practically-used activation functions,\nand also not closed with respect to the L^\∞-norm for all\npractically-used activation functions except for the ReLU and the parametric\nReLU. Finally, the function that maps a family of weights to the function\ncomputed by the associated network is not inverse stable for every practically\nused activation function. In other words, if f1, f2 are two functions\nrealized by neural networks and if f1, f2 are close in the sense that\n\‖f1 - f2\‖L^\∞ \≤ \ε for \ε > 0, it is,\nregardless of the size of \ε, usually not possible to find weights\nw1, w2 close together such that each fi is realized by a neural network\nwith weights wi. Overall, our findings identify potential causes for issues\nin the training procedure of deep learning such as no guaranteed convergence,\nexplosion of parameters, and slow convergence.\n