Deep Learning Scaling is Predictable, Empirically
2017/12/01 by Joel Hestness, Hestness, Joel, Sharan Narang +15 · 3 voices · 95 citations
Computer Science · #Topic Modeling #Machine Learning and Algorithms #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.1712.00409
Abstract
Deep learning (DL) creates impactful advances following a virtuous recipe: model architecture search, creating large training data sets, and scaling computation. It is widely believed that growing training sets and models should improve accuracy and result in better products. As DL application domains grow, we would like a deeper understanding of the relationships between training set size, computational scale, and model accuracy improvements to advance the state-of-the-art. This paper presents a large scale empirical characterization of generalization error and model size growth as training sets grow. We introduce a methodology for this measurement and test four machine learning domains: machine translation, language modeling, image processing, and speech recognition. Our empirical results show power-law generalization error scaling across a breadth of factors, resulting in power-law exponents---the "steepness" of the learning curve---yet to be explained by theoretical work. Further, model improvements only shift the error but do not appear to affect the power-law exponent. We also show that model size scales sublinearly with data size. These scaling relationships have significant implications on deep learning research, practice, and systems. They can assist model debugging, setting accuracy targets, and decisions about data set growth. They can also guide computing system design and underscore the importance of continued computational scaling.
Cited by
- Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study
- SLT: Robust Quantum Neural Networks for Noisy-Label Medical Image Classification via Supermartingale-based Label Transition
- Content Creation with Spillovers: An Incentive Design Approach
- Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling
- When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
- LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
- Pre-training under infinite compute
- Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
- Fast and Simplex: 2-Simplicial Attention in Triton
- Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
- From Efficiency Gains to Rebound Effects: The Problem of Jevons' Paradox in AI's Polarized Environmental Debate
- The World Is Bigger! A Computationally-Embedded Perspective on the Big World Hypothesis
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- Greedy dynamical meta-learning
- The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation
- Anti-Backdoor Coreset Selection via Cumulative Entropy
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- A Multi-fidelity Double-Delta Wing Dataset and Empirical Scaling Laws for GNN-based Aerodynamic Field Surrogate
- Neural Scaling Laws for Learning-based Identification of Nonlinear Systems
- When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
- On Improving Deep Active Learning with Formal Verification
- Scaling Laws for Code: Every Programming Language Matters
- From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
- ALIGN-FL: Architecture-independent Learning through Invariant Generative component sharing in Federated Learning
- Renormalizable Spectral-Shell Dynamics as the Origin of Neural Scaling Laws
- Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles
- Deep Unsupervised Anomaly Detection in Brain Imaging: Large-Scale Benchmarking and Bias Analysis
- Towards aligned body representations in vision models
- On the Origin of Algorithmic Progress in AI
- MOCLIP: A Foundation Model for Large-Scale Nanophotonic Inverse Design
- Fast Escape, Slow Convergence: Learning Dynamics of Phase Retrieval under Power-Law Data
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- Dynamic Reflections: Probing Video Representations with Text Alignment
- Quantifying the radiative response to surface temperature variability: A critical comparison of current methods
- Neither Consent nor Property: A Policy Lab for Data Law
- Minimizers of the Empirical Risk and Risk Monotonicity
- Efficient Client Contribution Evaluation for Horizontal Federated Learning
- Challenging Images For Minds and Machines
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- Slice Tuner: A Selective Data Acquisition Framework for Accurate and Fair Machine Learning Models
- Mixture-of-Depths Attention
- E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials
- Hi-BEHRT: Hierarchical Transformer-based model for accurate prediction of clinical events using multimodal longitudinal electronic health records
- An Update on a Progressively Expanded Database for Automated Lung Sound Analysis
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- Rapid Knee MRI Acquisition and Analysis Techniques for Imaging Osteoarthritis
- Data Leverage: A Framework for Empowering the Public in its Relationship\n with Technology Companies
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Relative-Based Scaling Law for Neural Language Models
- Parameter Estimation in River Transport Models With Immobile Phase Exchange Using Dimensional Analysis and Reduced-Order Models
- Position: Many generalization measures for deep learning are fragile
- Zero-Shot Performance Prediction for Probabilistic Scaling Laws
- Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
- The De-democratization of AI: Deep Learning and the Compute Divide in Artificial Intelligence Research
- How many samples to label for an application given a foundation model? Chest X-ray classification study
- Scaling Laws and Symmetry, Evidence from Neural Force Fields
- Reusing Overtrained Language Models Saturates Scaling
- Efficient Prediction of Pass@k Scaling in Large Language Models
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
- Beyond Human-Level Accuracy: Computational Challenges in Deep Learning
- Equivariance by Local Canonicalization: A Matter of Representation
- Are neural scaling laws leading quantum chemistry astray?
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- Model Merging Scaling Laws in Large Language Models
- Automated Detection of Aortic Stenosis Using Machine Learning
- Scaling Laws for Neural Material Models
- Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
- How fine can fine-tuning be? Learning efficient language models
- Turing-Universal Learners with Optimal Scaling Laws
- Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
- Physics-informed machine learning: case studies for weather and climate modelling
- Machine Learning For Elliptic PDEs: Fast Rate Generalization Bound, Neural Scaling Law and Minimax Optimality
- Investigating Test-Time Scaling with Reranking for Machine Translation
- Evolution of Concepts in Language Model Pre-Training
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- Why and How Governments Should Monitor AI Development
- FedERL: Federated Efficient and Robust Learning for Common Corruptions
- Revisiting ResNets: Improved Training and Scaling Strategies
- A Discrepancy-Based Perspective on Dataset Condensation
- Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives
- The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
- Language Models are Few-Shot Learners
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Learning with springs and sticks
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- Learning curves for Gaussian process regression with power-law priors and targets
- Engineering Reliable Deep Learning Systems
- Efficient Scaling for LLM-based ASR
- Neural Scaling Laws Surpass Chemical Accuracy for the Many-Electron Schrödinger Equation
- Forecasting West Nile virus with deep graph encoders
- Phi-Ground Tech Report: Advancing Perception in GUI Grounding
Discussions
Related