Understanding deep learning (still) requires rethinking generalization
2021/02/22 by Chiyuan Zhang, Samy Bengio, Moritz Hardt +2 · 137 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Gaussian Processes and Bayesian Inference #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.1145/3446776
openalex publication_date 2021/02/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
Despite their massive size, successful deep artificial neural networks can exhibit a remarkably small gap between training and test performance. Conventional wisdom attributes small generalization error either to properties of the model family or to the regularization techniques used during training. Through extensive systematic experiments, we show how these traditional approaches fail to explain why large neural networks generalize well in practice. Specifically, our experiments establish that state-of-the-art convolutional networks for image classification trained with stochastic gradient methods easily fit a random labeling of the training data. This phenomenon is qualitatively unaffected by explicit regularization and occurs even if we replace the true images by completely unstructured random noise. We corroborate these experimental findings with a theoretical construction showing that simple depth two neural networks already have perfect finite sample expressivity as soon as the number of parameters exceeds the number of data points as it usually does in practice. We interpret our experimental findings by comparison with traditional models. We supplement this republication with a new section at the end summarizing recent progresses in the field since the original version of this paper.
Cited by
- Balancing explainability and privacy in AI systems: A strategic imperative for managers
- A framework based on symbolic regression coupled with eXtended Physics-Informed Neural Networks for gray-box learning of equations of motion from data
- Spectral-transport stability and benign overfitting for minimum norm interpolation
- Not-So-Strange Love: Language Models and Generative Linguistic Theories are More Compatible than They Appear
- Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated Probabilities
- Linguists should learn to love speech-based deep learning models
- You can’t fight in here! This is BBS!
- Deep Learning is Not So Mysterious or Different
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- How linguistics learned to stop worrying and love the language models
- On the Optimal Memorization Power of ReLU Neural Networks
- Procrustean Training for Imbalanced Deep Learning
- Artificial intelligence for modelling infectious disease epidemics
- A PAC-Bayesian approach to generalization for quantum models
- XMIX: Combating Extremely Noisy Labels via Local Smoothness in Self-Supervised Feature Space
- Unsupervised learning enabled label-free single-pixel imaging for resilient information transmission through unknown dynamic scattering media
- On the effect of noise on fitting linear regression models
- Self-Supervised Learning from Noisy and Incomplete Data
- Semantic Space Search Trajectory Networks
- Inferring Communities of Interest in Collaborative Learning-based Recommender Systems
- Stabilizing Multimodal Autoencoders: A Theoretical and Empirical Analysis of Fusion Strategies
- Continuized Nesterov Acceleration for Non-Convex Optimization
- Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference
- From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
- Large language models are not about natural language
- Evaluating Singular Value Thresholds for DNN Weight Matrices based on Random Matrix Theory
- Robustness of Probabilistic Models to Low-Quality Data: A Multi-Perspective Analysis
- Optimized Machine Learning Methods for Studying the Thermodynamic Behavior of Complex Spin Systems
- Winning the Lottery by Preserving Network Training Dynamics with Concrete Ticket Search
- On the Effect of Regularization on Nonparametric Mean-Variance Regression
- Left shifting analysis of Human-Autonomous Team interactions to analyse risks of autonomy in high-stakes AI systems
- Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
- When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
- Lost in Time? A Meta-Learning Framework for Time-Shift-Tolerant Physiological Signal Transformation
- Geometry of Decision Making in Language Models
- Polynomially Over-Parameterized Convolutional Neural Networks Contain Structured Strong Winning Lottery Tickets
- Learning to Clean: Reinforcement Learning for Noisy Label Correction
- Subtract the Corruption: Training-Data-Free Corrective Machine Unlearning using Task Arithmetic
- Morality in AI. A plea to embed morality in LLM architectures and frameworks
- Membership Inference Attacks Beyond Overfitting
- Classification of Hope in Textual Data using Transformer-Based Models
- The first kind of predictability problem of El Niño predictions in a multivariate coupled data‐driven model
- Out-of-Context Misinformation Detection via Variational Domain-Invariant Learning with Test-Time Training
- DenoGrad: A Gradient-Based Framework for Data Refinement in Tabular and Time-Series Learning
- 2.5D Transformer: An Efficient 3D Seismic Interpolation Method without Full 3D Training
- Improving Quantum Neural Networks exploration by Noise-Induced Equalization
- Consistency Change Detection Framework for Unsupervised Remote Sensing Change Detection
- LandSegmenter: Towards a Flexible Foundation Model for Land Use and Land Cover Mapping
- Schedulers for Schedule-free: Theoretically inspired hyperparameters
- PlantTraitNet: An Uncertainty-Aware Multimodal Framework for Global-Scale Plant Trait Inference from Citizen Science Data
- DeepBooTS: Dual-Stream Residual Boosting for Drift-Resilient Time-Series Forecasting
- Non-Asymptotic Optimization and Generalization Bounds for Stochastic Gauss-Newton in Overparameterized Models
- Analyzing the Power of Chain of Thought through Memorization Capabilities
- PDE-SHARP: PDE Solver Hybrids through Analysis and Refinement Passes
- Random neural networks in the infinite width limit as Gaussian processes
- A Generalization Gap Estimation for Overparameterized Models via the Langevin Functional Variance
- From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI
- Multi-label Iterated Learning for Image Classification with Label Ambiguity
- A Practitioner's Guide to Kolmogorov-Arnold Networks
- Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime
- A Derandomization Framework for Structure Discovery: Applications in Neural Networks and Beyond
- Misinformation Detection using Large Language Models with Explainability
- Near-optimal Prediction Error Estimation for Quantum Machine Learning Models
- On the Neural Feature Ansatz for Deep Neural Networks
- Benchmarking noisy label detection methods
- Revisiting Meta-Learning with Noisy Labels: Reweighting Dynamics and Theoretical Guarantees
- ImpMIA: Leveraging Implicit Bias for Membership Inference Attack under Realistic Scenarios
- Automated Evolutionary Optimization for Resource-Efficient Neural Network Training
- Adaptive Gradient Calibration for Single-Positive Multi-Label Learning in Remote Sensing Image Scene Classification
- Contrastive Representations for Label Noise Require Fine-Tuning
- The Effect of Label Noise on the Information Content of Neural Representations
- Generalization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation Regime
- Hybrid Sequential Quantum Computing
- Computing frustration and near-monotonicity in deep neural networks
- Non-Vacuous Generalization Bounds: Can Rescaling Invariances Help?
- Anchored Supervised Fine-Tuning
- Interpretable deep learning: interpretation, interpretability, trustworthiness, and beyond
- SubZeroCore: A Submodular Approach with Zero Training for Coreset Selection
- HyperCore: Coreset Selection under Noise via Hypersphere Models
- Generative Quantile Regression with Variability Penalty
- First-Extinction Law for Resampling Processes
- Are Language Models Models?
- SASD: Self-Attention for Small Datasets—A case study in smart villages
- Structural network measures reveal the emergence of heavy-tailed degree distributions in lottery ticket multilayer perceptrons
- Adversarial machine learning :
- UniSino: Physics-Driven Foundational Model for Universal CT Sinogram Standardization
- Theoretical Foundations of Representation Learning using Unlabeled Data: Statistics and Optimization
- Fundamental Limits of Membership Inference Attacks on Machine Learning Models
- Similarity-based Outlier Detection for Noisy Object Re-Identification Using Beta Mixtures
- Data-driven discovery of dynamical models in biology
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- VASSO: Variance Suppression for Sharpness-Aware Minimization
- Foolish Crowds Support Benign Overfitting
- SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
- What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
- Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
- Guiding Noisy Label Conditional Diffusion Models with Score-based Discriminator Correction
- Is data-efficient learning feasible with quantum models?
- Learning with Noisy Labels Revisited: A Study Using Real-World Human Annotations
- The Computational Complexity of Satisfiability in State Space Models
- Wormhole Dynamics in Deep Neural Networks
- Self-supervised denoising for massive noisy images
- Review of deep learning: concepts, CNN architectures, challenges, applications, future directions
- Probabilistic Forecasting Method for Offshore Wind Farm Cluster under Typhoon Conditions: a Score-Based Conditional Diffusion Model
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
- Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
- Improving Generalization Bounds for VC Classes Using the Hypergeometric\n Tail Inversion
- 4D-PreNet: A Unified Preprocessing Framework for 4D-STEM Data Analysis
- Superior resilience to poisoning and amenability to unlearning in quantum machine learning
- Robust Classification under Noisy Labels: A Geometry-Aware Reliability Framework for Foundation Models
- Teaching the Teacher: Improving Neural Network Distillability for Symbolic Regression via Jacobian Regularization
- Reducing Data Requirements for Sequence-Property Prediction in Copolymer Compatibilizers via Deep Neural Network Tuning
- Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. [europepmc]
- Winsorization for Robust Bayesian Neural Networks. [europepmc]
- Deep Learning-Assisted Burn Wound Diagnosis: Diagnostic Model Development Study. [europepmc]
- Deep Learning for Detecting and Locating Myocardial Infarction by Electrocardiogram: A Literature Review. [europepmc]
- One-shot generalization in humans revealed through a drawing task. [europepmc]
- Trusting our machines: validating machine learning models for single-molecule transport experiments. [europepmc]
- A New Approach for Detecting Fundus Lesions Using Image Processing and Deep Neural Network Architecture Based on YOLO Model. [europepmc]
- Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learning. [europepmc]
- A Parallel Feature Fusion Network Combining GRU and CNN for Motor Imagery EEG Decoding. [europepmc]
- Segmentation by test-time optimization for CBCT-based adaptive radiation therapy. [europepmc]
- Efficient neural codes naturally emerge through gradient descent learning. [europepmc]
- Separation of scales and a thermodynamic description of feature learning in some CNNs. [europepmc]
- Evaluation of Risk of Bias in Neuroimaging-Based Artificial Intelligence Models for Psychiatric Diagnosis: A Systematic Review. [europepmc]
- Assessing the performance of QSP models: biology as the driver for validation. [europepmc]
- Machine learning models for diagnosis and prognosis of Parkinson's disease using brain imaging: general overview, main challenges, and future directions. [europepmc]
- Brain Tumor Detection Based on Deep Learning Approaches and Magnetic Resonance Imaging. [europepmc]
- Risk of data leakage in estimating the diagnostic performance of a deep-learning-based computer-aided system for psychiatric disorders. [europepmc]
- GeneSegNet: a deep learning framework for cell segmentation by integrating gene expression and imaging. [europepmc]
- Enhancements in Radiological Detection of Metastatic Lymph Nodes Utilizing AI-Assisted Ultrasound Imaging Data and the Lymph Node Reporting and Data System Scale. [europepmc]
- Novel applications of Convolutional Neural Networks in the age of Transformers. [europepmc]
- Enhancing Precision in Cardiac Segmentation for Magnetic Resonance-Guided Radiation Therapy Through Deep Learning. [europepmc]
- scTab: Scaling cross-tissue single-cell annotation models. [europepmc]
- Examining the challenges of blood pressure estimation via photoplethysmogram. [europepmc]
- Chisco: An EEG-based BCI dataset for decoding of imagined speech. [europepmc]
- Interpreting single-cell and spatial omics data using deep neural network training dynamics. [europepmc]
Related