Scaling Laws for Autoregressive Generative Modeling
2020/10/28 by Tom Henighan, Jared Kaplan, Henighan, Tom +35 · 73 citations
Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Domain Adaptation and Few-Shot Learning
paper · pdf · doi:10.48550/arxiv.2010.14701
Abstract
We identify empirical scaling laws for the cross-entropy loss in four domains: generative image modeling, video modeling, multimodal image↔text models, and mathematical problem solving. In all cases autoregressive Transformers smoothly improve in performance as model size and compute budgets increase, following a power-law plus constant scaling law. The optimal model size also depends on the compute budget through a power-law, with exponents that are nearly universal across all data domains. The cross-entropy loss has an information theoretic interpretation as S(True) + DKL(True||Model), and the empirical scaling laws suggest a prediction for both the true data distribution's entropy and the KL divergence between the true and model distributions. With this interpretation, billion-parameter Transformers are nearly perfect models of the YFCC100M image distribution downsampled to an 8× 8 resolution, and we can forecast the model size needed to achieve any given reducible loss (ie DKL) in nats/image for other resolutions. We find a number of additional scaling laws in specific domains: (a) we identify a scaling relation for the mutual information between captions and images in multimodal models, and show how to answer the question "Is a picture worth a thousand words?"; (b) in the case of mathematical problem solving, we identify scaling laws for model performance when extrapolating beyond the training distribution; (c) we finetune generative image models for ImageNet classification and find smooth scaling of the classification loss and error rate, even as the generative loss levels off. Taken together, these results strengthen the case that scaling laws have important implications for neural network performance, including on downstream tasks.
Citations
Cited by
- The Law of Multi-Model Collaboration: Scaling Limits of Model Ensembling for Large Language Models
- When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
- Repurposing 2D Diffusion Models for 3D Shape Completion
- From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis
- Renormalizable Spectral-Shell Dynamics as the Origin of Neural Scaling Laws
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
- Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles
- Masked Diffusion for Generative Recommendation
- ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
- Spanning Tree Autoregressive Visual Generation
- Quantifying vacuum-like jets in heavy-ion collisions: a Machine Learning study
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- Overcoming distortion in multidimensional predictive representation
- 8-bit Optimizers via Block-wise Quantization
- The missing data for intelligent scientific instruments
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- Relative-Based Scaling Law for Neural Language Models
- Parameter Estimation in River Transport Models With Immobile Phase Exchange Using Dimensional Analysis and Reduced-Order Models
- High-Resolution Complex Scene Synthesis with Transformers
- Zero-Shot Performance Prediction for Probabilistic Scaling Laws
- Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
- Predicting Task Performance with Context-aware Scaling Laws
- Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- Scaling Laws and Symmetry, Evidence from Neural Force Fields
- AutoPR: Let's Automate Your Academic Promotion!
- Reusing Overtrained Language Models Saturates Scaling
- Membership Inference Attacks on Tokenizers of Large Language Models
- Efficient Prediction of Pass@k Scaling in Large Language Models
- Transformers Discover Molecular Structure Without Graph Priors
- What Scales in Cross-Entropy Scaling Law?
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- A universal compression theory for lottery ticket hypothesis and neural scaling laws
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
- Convergence and Divergence of Language Models under Different Random Seeds
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
- SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement
- Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
- Pretraining Scaling Laws for Generative Evaluations of Language Models
- Semantic Agreement Enables Efficient Open-Ended LLM Cascades
- Scaling Laws are Redundancy Laws
- Scaling Laws for Transfer
- Turing-Universal Learners with Optimal Scaling Laws
- Proof Artifact Co-training for Theorem Proving with Language Models
- Can neural networks do arithmetic? A survey on the elementary numerical skills of state-of-the-art deep learning models
- Training Compute-Optimal Protein Language Models
- Lessons from complex systems science for AI governance
- Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
- Why and How Governments Should Monitor AI Development
- SpecDiff: Accelerating Diffusion Model Inference with Self-Speculation
- Revisiting ResNets: Improved Training and Scaling Strategies
- Embodied intelligence via learning and evolution
- Neural Scaling Laws for Deep Regression
- Scaling Up Vision-Language Pre-training for Image Captioning
- Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph Embedding
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- KG-o1: Enhancing Multi-hop Question Answering in Large Language Models via Knowledge Graph Integration
- Scaling Learned Image Compression Models up to 1 Billion
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Neural Scaling Laws Surpass Chemical Accuracy for the Many-Electron Schrödinger Equation
- Empowering Tabular Data Preparation with Language Models: Why and How?
Related