2023/01/30 by Ayush Bharadwaj, Bharadwaj, Ayush, Serkan Hoşten +1 · 2 citations
Computer Science · Mathematics · #Algebraic Geometry (math.AG) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Stochastic Gradient Optimization Techniques #Tensor decomposition and applications #Topological and Geometric Data Analysis
paper · pdf · doi:10.48550/arxiv.2301.12651
openalex publication_date 2023/01/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We extend the work of Mehta, Chen, Tang, and Hauenstein on computing the complex critical points of the loss function of deep linear neutral networks when the activation function is the identity function. For networks with a single hidden layer trained on a single data point we give an improved bound on the number of complex critical points of the loss function. We show that for any number of hidden layers complex critical points with zero coordinates arise in certain patterns which we completely classify for networks with one hidden layer. We report our results of computational experiments with varying network architectures defining small deep linear networks using HomotopyContinuation.jl.