Scaling Laws for Autoregressive Generative Modeling
2020/10/28 by Tom Henighan, Jared Kaplan, Henighan, Tom +36 · 126 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #cs.CL #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2010.14701
20+17 pages, 33 figures; added appendix with additional language results
arxiv created 2020/11/06 · arxiv updated 2020/11/09
Abstract
We identify empirical scaling laws for the cross-entropy loss in four domains: generative image modeling, video modeling, multimodal image↔text models, and mathematical problem solving. In all cases autoregressive Transformers smoothly improve in performance as model size and compute budgets increase, following a power-law plus constant scaling law. The optimal model size also depends on the compute budget through a power-law, with exponents that are nearly universal across all data domains. The cross-entropy loss has an information theoretic interpretation as S(True) + DKL(True||Model), and the empirical scaling laws suggest a prediction for both the true data distribution's entropy and the KL divergence between the true and model distributions. With this interpretation, billion-parameter Transformers are nearly perfect models of the YFCC100M image distribution downsampled to an 8× 8 resolution, and we can forecast the model size needed to achieve any given reducible loss (ie DKL) in nats/image for other resolutions. We find a number of additional scaling laws in specific domains: (a) we identify a scaling relation for the mutual information between captions and images in multimodal models, and show how to answer the question "Is a picture worth a thousand words?"; (b) in the case of mathematical problem solving, we identify scaling laws for model performance when extrapolating beyond the training distribution; (c) we finetune generative image models for ImageNet classification and find smooth scaling of the classification loss and error rate, even as the generative loss levels off. Taken together, these results strengthen the case that scaling laws have important implications for neural network performance, including on downstream tasks.
Citations
Cited by
- The Law of Multi-Model Collaboration: Scaling Limits of Model Ensembling for Large Language Models
- When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
- Repurposing 2D Diffusion Models for 3D Shape Completion
- From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
- Renormalizable Spectral-Shell Dynamics as the Origin of Neural Scaling Laws
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
- Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles
- Masked Diffusion for Generative Recommendation
- ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
- Spanning Tree Autoregressive Visual Generation
- Quantifying vacuum-like jets in heavy-ion collisions: a Machine Learning study
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- Overcoming distortion in multidimensional predictive representation
- 8-bit Optimizers via Block-wise Quantization
- The missing data for intelligent scientific instruments
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- Relative-Based Scaling Law for Neural Language Models
- Parameter Estimation in River Transport Models With Immobile Phase Exchange Using Dimensional Analysis and Reduced-Order Models
- High-Resolution Complex Scene Synthesis with Transformers
- Zero-Shot Performance Prediction for Probabilistic Scaling Laws
- Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
- Predicting Task Performance with Context-aware Scaling Laws
- Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- Scaling Laws and Symmetry, Evidence from Neural Force Fields
- AutoPR: Let's Automate Your Academic Promotion!
- Reusing Overtrained Language Models Saturates Scaling
- Membership Inference Attacks on Tokenizers of Large Language Models
- Efficient Prediction of Pass@k Scaling in Large Language Models
- Transformers Discover Molecular Structure Without Graph Priors
- What Scales in Cross-Entropy Scaling Law?
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- A universal compression theory for lottery ticket hypothesis and neural scaling laws
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
- Convergence and Divergence of Language Models under Different Random Seeds
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
- SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement
- Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
- Pretraining Scaling Laws for Generative Evaluations of Language Models
- Semantic Agreement Enables Efficient Open-Ended LLM Cascades
- Scaling Laws are Redundancy Laws
- Scaling Laws for Transfer
- Turing-Universal Learners with Optimal Scaling Laws
- Proof Artifact Co-training for Theorem Proving with Language Models
- Can neural networks do arithmetic? A survey on the elementary numerical skills of state-of-the-art deep learning models
- Training Compute-Optimal Protein Language Models
- Lessons from complex systems science for AI governance
- Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
- Why and How Governments Should Monitor AI Development
- SpecDiff: Accelerating Diffusion Model Inference with Self-Speculation
- Revisiting ResNets: Improved Training and Scaling Strategies
- Embodied Intelligence via Learning and Evolution
- Neural Scaling Laws for Deep Regression
- Scaling Up Vision-Language Pre-training for Image Captioning
- Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph Embedding
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- KG-o1: Enhancing Multi-hop Question Answering in Large Language Models via Knowledge Graph Integration
- Scaling Learned Image Compression Models up to 1 Billion
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Neural Scaling Laws Surpass Chemical Accuracy for the Many-Electron Schrödinger Equation
- Empowering Tabular Data Preparation with Language Models: Why and How?
- Can Language Models Discover Scaling Laws?
- Teaching Autoregressive Language Models Complex Tasks By Demonstration
- Which transformer architecture fits my data? A vocabulary bottleneck in self-attention
- Scaling Scaling Laws with Board Games
- Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
- Search Spaces for Neural Model Training
- Designing lensless imaging systems to maximize information capture
- Improved Scaling Laws in Linear Regression via Data Reuse
- Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
- Beyond Scaling Curves: Internal Dynamics of Neural Networks Through the NTK Lens
- ICAS: Detecting Training Data from Autoregressive Image Generative Models
- Pre-Trained Policy Discriminators are General Reward Models
- Hita: Holistic Tokenizer for Autoregressive Image Generation
- Energy-Based Transformers are Scalable Learners and Thinkers
- In situ fine-tuning of in silico trained Optical Neural Networks
- Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
- Maximal Update Parametrization and Zero-Shot Hyperparameter Transfer for Fourier Neural Operators
- A Scaling Law for Synthetic-to-Real Transfer: How Much Is Your Pre-training Effective?
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- Deep generative models as the probability transformation functions
- Watermarking Autoregressive Image Generation
- Scaling Laws of Motion Forecasting and Planning -- Technical Report
- Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value
- OneRec Technical Report
- Complexity Scaling Laws for Neural Models using Combinatorial Optimization
- Scaling Laws for Uncertainty in Deep Learning
- Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
- Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
- Hopscotch: Discovering and Skipping Redundancies in Language Models
- MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale
- Quiet Feature Learning in Algorithmic Tasks
- Exploring Scaling Laws for EHR Foundation Models
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
- Scaling Laws for Neural Machine Translation
- Is the Number of Trainable Parameters All That Actually Matters?
- Inference Compute-Optimal Video Vision Language Models
- Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
- Superposition Yields Robust Neural Scaling
- Guiding Data Collection via Factored Scaling Curves
- Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- A Reinforcement Learning Environment for Mathematical Reasoning via Program Synthesis
- Internal Data Repetition Destroys Language Models
- Jailbreaking Embodied LLMs via Action-level Manipulation
- LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability
- Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
- Polynomial Speedup in Diffusion Models with the Multilevel Euler-Maruyama Method
- Deriving Neural Scaling Laws from the statistics of natural language
- MAYA: Addressing Inconsistencies in Generative Password Guessing through a Unified Benchmark
- Data Scaling Laws for End-to-End Autonomous Driving
- Compute-Optimal LLMs Provably Generalize Better With Scale
- Learning to Reason under Off-Policy Guidance
- Scalable multilayer diffractive neural network with all-optical nonlinear activation
- Scaling Laws for Data-Efficient Visual Transfer Learning
- Can Pre-training Indicators Reliably Predict Fine-tuning Outcomes of LLMs?
- Better artificial intelligence does not mean better models of biology
Related