2019/11/02 by Lei Wu, Wu, Lei, Qingcan Wang +3 · 1 citation
Computer Science · Engineering · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.1911.00645
openalex publication_date 2019/11/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We analyze the global convergence of gradient descent for deep linear residual networks by proposing a new initialization: zero-asymmetric (ZAS) initialization. It is motivated by avoiding stable manifolds of saddle points. We prove that under the ZAS initialization, for an arbitrary target matrix, gradient descent converges to an ε-optimal point in O(L3 log(1/ε)) iterations, which scales polynomially with the network depth L. Our result and the exp(Ω(L)) convergence time for the standard initialization (Xavier or near-identity) [Shamir, 2018] together demonstrate the importance of the residual structure and the initialization in the optimization for deep linear neural networks, especially when L is large.