2020/04/22 by Utkarsh Sharma, Jared Kaplan, Sharma, Utkarsh +1 · 1 voice · 4 citations
Computer Science · Mathematics · #FOS: Computer and information sciences #Face and Expression Recognition #Image Processing and 3D Reconstruction #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2004.10802
16+12 pages, 11+11 figures
arxiv created 2020/04/22 · openalex publication_date 2020/04/22 · arxiv published 2020/04/22 · arxiv updated 2020/04/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
When data is plentiful, the loss achieved by well-trained neural networks scales as a power-law L ∝ N-α in the number of network parameters N. This empirical scaling law holds for a wide variety of data modalities, and may persist over many orders of magnitude. The scaling law can be explained if neural models are effectively just performing regression on a data manifold of intrinsic dimension d. This simple theory predicts that the scaling exponents α≈ 4/d for cross-entropy and mean-squared error losses. We confirm the theory by independently measuring the intrinsic dimension and the scaling exponents in a teacher/student framework, where we can study a variety of d and α by dialing the properties of random teacher networks. We also test the theory with CNN image classifiers on several datasets and with GPT-type language models.