2018/02/10 by Dmitry Yarotsky, Yarotsky, Dmitry · 28 citations
Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #Machine Learning and Algorithms #Model Reduction and Neural Networks #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE) #cs.NE
paper · pdf · doi:10.48550/arxiv.1802.03620
21 pages. In v2: more discussion of phase diagram, described approximation for $p\in(\frac{1}ν,\frac{2}ν)$
openalex publication_date 2018/02/10 · arxiv created 2018/06/06 · arxiv updated 2018/06/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider approximations of general continuous functions on finite-dimensional cubes by general deep ReLU neural networks and study the approximation rates with respect to the modulus of continuity of the function and the total number of weights W in the network. We establish the complete phase diagram of feasible approximation rates and show that it includes two distinct phases. One phase corresponds to slower approximations that can be achieved with constant-depth networks and continuous weight assignments. The other phase provides faster approximations at the cost of depths necessarily growing as a power law L∼ Wα, 0<α≤ 1, and with necessarily discontinuous weight assignments. In particular, we prove that constant-width fully-connected networks of depth L∼ W provide the fastest possible approximation rate ‖f-\widetilde f‖_∞ = O(ωf(O(W-2/ν))) that cannot be achieved with less deep networks.