Training Compute-Optimal Large Language Models
2022/03/29 by Jordan Hoffmann, Sebastian Borgeaud, Hoffmann, Jordan +41 · 13 voices · 429 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Machine Learning and Algorithms
paper · pdf · doi:10.48550/arxiv.2203.15556
Abstract
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4× more more data. Chinchilla uniformly and significantly outperforms Gopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, greater than a 7% improvement over Gopher.
Cited by
- Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study
- Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
- Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data
- Position: Quantum Program Generation Must Prioritize Validity Over Probabilistic Scaling
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
- Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
- Pixel-Space Diffusion Transformers
- MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization
- Post-Training in Time Series Foundation Models: A Unifying Framework
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- Circuit Claims Depend on What Is Extracted and How It Is Compared
- Mobius Learning: Cyclic Depth Folding in Transformers
- Enhancing Small Language Models Reasoning through Knowledge Graph Grounding
- Capability from Access Structure, Not Scale: Lower Bounds and Pre-Registered Tests for Hybrid Sequence Models
- Loop the Loopies!
- Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
- An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation
- Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models
- HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency
- Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors
- Scaling Point-in-Time Language Models
- An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
- Understanding Reasoning from Pretraining to Post-Training
- xHC: Expanded Hyper-Connections
- Do Transformers Need Three Projections? Systematic Study of QKV Variants
- Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
- Response drift across frontier large language models
- Domyn-Small: A European 10B Reasoning Language Model
- Information-Theoretic Limits of Reliability and Scaling in Language Models
- There Will Be a Scientific Theory of Deep Learning
- Sparser, Faster, Lighter Transformer Language Models
- Spelling Bee Embeddings for Language Modeling
- In-Context Probing for Membership Inference in Fine-Tuned Language Models
- Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
- Epistemological Fault Lines Between Human and Artificial Intelligence
- TRINITY: An Evolved LLM Coordinator
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- Kimi Linear: An Expressive, Efficient Attention Architecture
- LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
- Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- On the Theoretical Limitations of Embedding-Based Retrieval
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- Whither symbols in the era of advanced neural networks?
- The wall confronting large language models
- Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
- Fast and Simplex: 2-Simplicial Attention in Triton
- Small Language Models are the Future of Agentic AI
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- Reinforcement Pre-Training
- Post-Post-API Age: Studying Digital Platforms in Scant Data Access Times
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Protein Structure Tokenization: Benchmarking and New Recipe
- Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
- Deep Learning is Not So Mysterious or Different
- Large Language Diffusion Models
- Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
- Value-Based Deep RL Scales Predictably
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
- The World Is Bigger! A Computationally-Embedded Perspective on the Big World Hypothesis
- Theoretical Foundations of Scaling Law in Familial Models
- The Law of Multi-Model Collaboration: Scaling Limits of Model Ensembling for Large Language Models
- Understanding the Mechanisms of Fast Hyperparameter Transfer
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models
- Attention Residuals
- Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
- Scale Weight Decay and Train Better
- Generative Artificial Intelligence for Software Engineering -- A Research Agenda
- The Entropic Bound for Transformers: Why Static Rank Fails and Attention-Native Rank Recovers
- Bridging Compute- and Data-Optimal Pretraining
- When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale, Architecture, and Output Parsing Robustness
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning
- Hierarchical Grading in Large Language Models
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
- Market Design for AI: Beyond the Copyright Binary
- How Transformers Learn to Plan via Multi-Token Prediction
- SDUM: A Scalable Deep Unrolled Model for Universal MRI Reconstruction
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- Beyond Context: Large Language Models Failure to Grasp Users Intent
- A Multi-fidelity Double-Delta Wing Dataset and Empirical Scaling Laws for GNN-based Aerodynamic Field Surrogate
- Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
- The AI Scaling Wall of Diminishing Returns: Of LLMs, Electric Dogs, and General Relativity
- Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
- DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations
- AraMix: Recycling, Refiltering, and Deduplicating to Deliver the Largest Arabic Pretraining Corpus
- HARBOR: Holistic Adaptive Risk assessment model for BehaviORal healthcare
- When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
- CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
- Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
- TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
- LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
- Bolmo: Byteifying the Next Generation of Language Models
- Dual-objective Language Models: Training Efficiency Without Overfitting
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- Scaling Laws for Code: Every Programming Language Matters
- MiniLingua: A Small Open-Source LLM for European Languages
- PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
- BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
- The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
- Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- On the Dynamics of Multi-Agent LLM Communities Driven by Value Diversity
- Renormalizable Spectral-Shell Dynamics as the Origin of Neural Scaling Laws
- On Learning-Curve Monotonicity for Maximum Likelihood Estimators
- XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs
- DINOv2: Learning Robust Visual Features without Supervision
- Scaling Behavior of Discrete Diffusion Language Models
- Neurosymbolic Information Extraction from Transactional Documents
- Graph Deep Learning for Intracranial Aneurysm Blood Flow Simulation and Risk Assessment
- Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
- Is GPT-OSS All You Need? Benchmarking Large Language Models for Financial Intelligence and the Surprising Efficiency Paradox
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
- FOAM: Blocked State Folding for Memory-Efficient LLM Training
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- A Latent Variable Framework for Scaling Laws in Large Language Models
- LLM Harms: A Taxonomy and Discussion
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
- A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- From FLOPs to Footprints: The Resource Cost of Artificial Intelligence
- Data Curation Through the Lens of Spectral Dynamics: Static Limits, Dynamic Acceleration, and Practical Oracles
- Humanity in the Age of AI: Reassessing 2025's Existential-Risk Narratives
- Parameter Reduction Improves Vision Transformers: A Comparative Study of Sharing and Width Reduction
- Start Making Sense(s): A Developmental Probe of Attention Specialization Using Lexical Ambiguity
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- LUMOS: Large User MOdels for User Behavior Prediction
- Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium
- Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code
- Complex QA and language models hybrid architectures, Survey
- Fluid Intelligence: A Forward Look on AI Foundation Models in Computational Fluid Dynamics
- Robot-Powered Data Flywheels: Deploying Robots in the Wild for Continual Data Collection and Foundation Model Adaptation
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Fast Escape, Slow Convergence: Learning Dynamics of Phase Retrieval under Power-Law Data
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- Selective Rotary Position Embedding
- Energy Scaling Laws for Diffusion Models: Quantifying Compute and Carbon Emissions in Image Generation
- Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
- From generative AI to the brain: five takeaways
- GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI
- Developing a Grounded View of AI
- Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
- DAP: A Discrete-token Autoregressive Planner for Autonomous Driving
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- Can large language models be a cardinality estimator? An empirical study
- Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
- Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis
- FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
- Virtual Width Networks
- Quantifying vacuum-like jets in heavy-ion collisions: a Machine Learning study
- Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
- Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
- A Generalized Spectral Framework to Expain Neural Scaling and Compression Dynamics
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
- Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
- Model Agreement via Anchoring
- LLM Driven Processes to Foster Explainable AI
- The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
- Scaling Laws and In-Context Learning: A Unified Theoretical Framework
- BiPETE: A Bi-Positional Embedding Transformer Encoder for Risk Assessment of Alcohol and Substance Use Disorder with Electronic Health Records
- Deep Progressive Training: scaling up depth capacity of zero/one-layer models
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- Diffusion Language Models are Super Data Learners
- Towards Multi-Fidelity Scaling Laws of Neural Surrogates in CFD
- AI Progress Should Be Measured by Capability-Per-Resource, Not Scale Alone: A Framework for Gradient-Guided Resource Allocation in LLMs
- ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Neither Consent nor Property: A Policy Lab for Data Law
- An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- Beyond Benchmarks: The Economics of AI Inference
- Towards Scaling Laws for Symbolic Regression
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
- A social path to human-like artificial intelligence
- From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs
- Mixture-of-Depths Attention
- How far can you go? Extrapolating values of catalytic activity from known protein landscapes in natural and directed evolution
- ProGen2: Exploring the boundaries of protein language models
- Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
- Will Scaling Improve Social Simulation with LLMs?
- Can Large Language Models Transform Computational Social Science?
- HRM-Text: Efficient Pretraining Beyond Scaling
- The Rise of AI in Weather and Climate Information and its Impact on Global Inequality
- Large language models encode clinical knowledge
- MIN-Merging: Merge the Important Neurons for Model Merging
- Spectral imaginings and sympoietic creativity: AI hallucinations and the ethics of posthuman creativity
- What Really Matters in Matrix-Whitening Optimizers?
- The Economics of AI Training Data: A Research Agenda
- The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- Larger and more instructable language models become less reliable
- Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks
- Frustratingly Easy Task-aware Pruning for Large Language Models
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
- Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
- Boosting Accuracy and Efficiency of Budget Forcing in LLMs via Reinforcement Learning for Mathematical Reasoning
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Fluidity Index: Next-Generation Super-intelligence Benchmarks
- Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- Relative-Based Scaling Law for Neural Language Models
- Video Consistency Distance: Enhancing Temporal Consistency for Image-to-Video Generation via Reward-Based Fine-Tuning
- Parameter Estimation in River Transport Models With Immobile Phase Exchange Using Dimensional Analysis and Reduced-Order Models
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- Position: Many generalization measures for deep learning are fragile
- Unbiased Gradient Low-Rank Projection
- Elastic ViTs from Pretrained Models without Retraining
- Protein Language Models: Is Scaling Necessary?
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- Zero-Shot Performance Prediction for Probabilistic Scaling Laws
- Computational Budget Should Be Considered in Data Selection
- Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
- Midtraining Bridges Pretraining and Posttraining Distributions
- Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
- VaultGemma: A Differentially Private Gemma Model
- Position: Require Frontier AI Labs To Release Small "Analog" Models
- Scaling Vision Transformers for Functional MRI with Flat Maps
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
- The Art of Scaling Reinforcement Learning Compute for LLMs
- SAGE: Streaming Agreement-Driven Gradient Sketches for Representative Subset Selection
- Structured Sparsity and Weight-adaptive Pruning for Memory and Compute efficient Whisper models
- BIGFix: Bidirectional Image Generation with Token Fixing
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- Indoor Localization using Compact, Telemetry-Agnostic, Transfer-Learning Enabled Decoder-Only Transformer
- Demystifying Numerosity in Diffusion Models -- Limitations and Remedies
- Automating Structural Engineering Workflows with Large Language Model Agents
- Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- Variational Open-Domain Question Answering
- DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Forget Attention: Importance-Aware Attention Is All You Need
- On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
- Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
- Token Is All You Price
- Scaling Laws and Symmetry, Evidence from Neural Force Fields
- KORMo: Korean Open Reasoning Model for Everyone
- MeSH: Memory-as-State-Highways for Recursive Transformers
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- Mid-Training of Large Language Models: A Survey
- Reusing Overtrained Language Models Saturates Scaling
- Optimal Stopping vs Best-of-N for Inference Time Optimization
- Test-Time Efficient Pretrained Model Portfolios for Time Series Forecasting
- Membership Inference Attacks on Tokenizers of Large Language Models
- Latent Speech-Text Transformer
- D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Increasing LLM response trustworthiness using voting ensembles
- What Makes Diffusion Language Models Super Data Learners?
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- Optimal Scaling Needs Optimal Norm
- TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Self-Speculative Masked Diffusions
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- Leave No TRACE: Black-box Detection of Copyrighted Dataset Usage in Large Language Models via Watermarking
- Accuracy Law for the Future of Deep Time Series Forecasting
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
- The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
- Composer: A Search Framework for Hybrid Neural Architecture Design
- Generalized Parallel Scaling with Interdependent Generations
- BroRL: Scaling Reinforcement Learning via Broadened Exploration
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
- Improving Metacognition and Uncertainty Communication in Language Models
- Are neural scaling laws leading quantum chemistry astray?
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- LLM-Powered Code Analysis and Optimization for Gaussian Splatting Kernels
- Per-example gradients: a new frontier for understanding and improving optimizers
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling View
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
- Efficient Hyperparameter Tuning via Trajectory Invariance Principle
- Vision Function Layer in Multimodal LLMs
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
- Model Merging Scaling Laws in Large Language Models
- Learning to Ponder: Adaptive Reasoning in Latent Space
- Muon: Training and Trade-offs with Latent Attention and MoE
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Reinforcement Mid-Training
- DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
- Towards a Comprehensive Scaling Law of Mixture-of-Experts
- Training Optimal Large Diffusion Language Models
- Evaluating large language models on business process modeling: framework, benchmark, and self-improvement analysis
- Beyond Outliers: A Study of Optimizers Under Quantization
- Mapping Overlaps in Benchmarks through Perplexity in the Wild
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Effective Quantization of Muon Optimizer States
- Tracing the Representation Geometry of Language Models from Pretraining to Post-training
- Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
- The (In)Effectiveness of Psychological Targeting: A Meta‐Analytic Review
- Scale-Wise VAR is Secretly Discrete Diffusion
- Compute-Optimal Quantization-Aware Training
- Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
- Predicting LLM Reasoning Performance with Small Proxy Model
- A short survey on almost orthogonal vectors in a few specific large dimensions
- Scaling Laws are Redundancy Laws
- Look Before you Leap: Estimating LLM Benchmark Scores from Descriptions
- Artificial Intelligence’s new clothes? A system technology perspective
- SPARQ: An Optimization Framework for the Distribution of AI-Intensive Applications under Non-Linear Delay Constraints
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
- SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions
- Towards joint scaling laws with optimal batch size schedules
- A foundation model of numerical intelligence with cross-disciplinary generalization
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- Explaining Data Mixing Scaling Laws
- Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
- Generative AI for Economic Research: Use Cases and Implications for Economists
- Training Compute-Optimal Protein Language Models
- Linguistic properties and model scale in brain encoding: from small to compressed language models
- Auditing large language models: a three-layered approach
- Bridging the data gap between children and large language models
- Reinforcement Learning on Pre-Training Data
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
- Investigating Test-Time Scaling with Reranking for Machine Translation
- nDNA -- the Semantic Helix of Artificial Cognition
- Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
- MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
- On the Convergence of Muon and Beyond
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
- FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
- Learning the natural history of human disease with generative transformers
- VerilogMonkey: Exploring Parallel Scaling for Automated Verilog Code Generation with LLMs
- Large Language Models and the Reverse Turing Test
- Provable Generalization in Overparameterized Neural Nets
- Large Language Model Scaling Laws for Neural Quantum States in Quantum Chemistry
- Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- Instant prediction of relaxation in moiré superlattices using neural networks
- Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
- Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
- Neural Scaling Laws for Deep Regression
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- LLM Architecture, Scaling Laws, and Economics: A Quick Summary
- CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
- Generative Data Refinement: Just Ask for Better Data
- RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector
- Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
- Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
- Effectively obtaining acoustic, visual and textual data from videos
- Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives
- Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
- Can VLMs Recall Factual Associations From Visual References?
- Understanding sparse autoencoder scaling in the presence of feature manifolds
- Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
- OneRec-V2 Technical Report
- Training Transformers for Mesh-Based Simulations
- Compute-Optimal Scaling for Value-Based Deep RL
- In2x at WMT25 Translation Task
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
- Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
- ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
- Note on Selection Bias in Observational Estimates of Algorithmic Progress
- DeepFleet: Multi-Agent Foundation Models for Mobile Robots
- Scaling Learned Image Compression Models up to 1 Billion
- Towards Scalable Training for Handwritten Mathematical Expression Recognition
- Pushing the Envelope of LLM Inference on AI-PC
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- Position: Ideas Should be the Center of Machine Learning Research
- Multimodal learning with next-token prediction for large multimodal models
- Leveraging LLMs for Smart Cities Qualitative Data Analysis
- Intuition emerges in Maximum Caliber models at criticality
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- AGI for the Earth, the path, possibilities and how to evaluate intelligence of models that work with Earth Observation Data?
- MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
- Small transformer architectures for task switching
- Chain of Questions: Guiding Multimodal Curiosity in Language Models
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Survey of Large Language Models in Extended Reality: Technical Paradigms and Application Frontiers
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- Lessons from complex systems science for AI governance
- Co-Producing AI: Toward an Augmented, Participatory Lifecycle
Discussions
- arxiv.org/abs/2203.15556 [bsky, 17 points, 1 comments]
- For reference: [bsky, 13 points, 1 comments]
- Training Compute-Optimal Large Language Models [hn, 8 points, 0 comments]
- 70B language model outperforms Gopher and GPT-3 [hn, 5 points, 0 comments]
- "We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant." arxiv.org [bsky, 5 points, 0 comments]
- Training Compute-Optimal Large Language Models [hn, 3 points, 0 comments]
- Two to start. Hoffmann et al. 2022 (Chinchilla) worked out the compute/data tradeoff: arxiv.org/abs/2203.15556. Villalobos et al. 2022 project public human text data running out somewhere 2026-2032: a [bsky, 3 points, 1 comments]
- Did that paper’s results get revised by the Chinchilla paper below, which suggested existing models were undertrained for their size? arxiv.org/abs/2203.15556 [bsky, 2 points, 0 comments]
- The Chinchilla paper (Training Compute-Optimal Large Language Models, Hoffmann et al., 2022) is widely regarded as a gold standard for empirically characterizing and optimizing scaling laws for large [bsky, 1 points, 1 comments]
- Training Compute-Optimal Large Language Models [hn, 1 points, 0 comments]
- ... Chinchilla-Papers (arxiv.org/abs/2203.15556): Minimum-Envelope-Methode (untere Hüllkurve), Iso-Flop-Analyse (am zuverlässigsten) und Kurvenanpassung an ein parametrisches Modell. 3/3 youtu.be/6Q-E [bsky, 0 points, 0 comments]
- Discover the latest advancements in AI as researchers explore innovative applications and ethical considerations surrounding artificial intelligence. This article delves into how AI is reshaping vario [bsky, 0 points, 0 comments]
- Chinchilla: Smaller models trained on more data* are what you need.
*10x more compute should be spent on 3.2x larger model and 3.2x more tokens
https://arxiv.org/abs/2203.15556 [bsky, 0 points, 1 comments]
Related