2022/03/02 by Zhou, Jinxin, Tianyu Ding, Li, Xiao +9 · 9 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #FOS: Mathematics #Information Theory (cs.IT) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Optimization and Control (math.OC) #Sparse and Compressive Sensing Techniques #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2203.01238
openalex publication_date 2022/03/02 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28
When training deep neural networks for classification tasks, an intriguing\nempirical phenomenon has been widely observed in the last-layer classifiers and\nfeatures, where (i) the class means and the last-layer classifiers all collapse\nto the vertices of a Simplex Equiangular Tight Frame (ETF) up to scaling, and\n(ii) cross-example within-class variability of last-layer activations collapses\nto zero. This phenomenon is called Neural Collapse (NC), which seems to take\nplace regardless of the choice of loss functions. In this work, we justify NC\nunder the mean squared error (MSE) loss, where recent empirical evidence shows\nthat it performs comparably or even better than the de-facto cross-entropy\nloss. Under a simplified unconstrained feature model, we provide the first\nglobal landscape analysis for vanilla nonconvex MSE loss and show that the\n(only!) global minimizers are neural collapse solutions, while all other\ncritical points are strict saddles whose Hessian exhibit negative curvature\ndirections. Furthermore, we justify the usage of rescaled MSE loss by probing\nthe optimization landscape around the NC solutions, showing that the landscape\ncan be improved by tuning the rescaling hyperparameters. Finally, our\ntheoretical findings are experimentally verified on practical network\narchitectures.\n