Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
2023/04/03 by Stella Biderman, Biderman, Stella, Hailey Schoelkopf +23 · 3 voices · 216 citations
#cs.CL
paper · pdf · doi:10.48550/arxiv.2304.01373
Abstract
How do large language models (LLMs) develop and evolve over the course of training? How do these patterns change as models scale? To answer these questions, we introduce Pythia, a suite of 16 LLMs all trained on public data seen in the exact same order and ranging in size from 70M to 12B parameters. We provide public access to 154 checkpoints for each one of the 16 models, alongside tools to download and reconstruct their exact training dataloaders for further study. We intend Pythia to facilitate research in many areas, and we present several case studies including novel results in memorization, term frequency effects on few-shot performance, and reducing gender bias. We demonstrate that this highly controlled setup can be used to yield novel insights toward LLMs and their training dynamics. Trained models, analysis code, training code, and training data can be found at \urlhttps://github.com/EleutherAI/pythia.
Cited by
- Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable
- Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
- Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
- surprisal is Not a Theory
- Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions
- Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
- Abstraction Induces the Brain Alignment of Language and Speech Models
- No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training
- Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition
- Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
- First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
- On the Interpretability of Whisper Encodings Using Sparse Autoencoders
- Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance
- More Than Memory: Task-Conditioned Signed FFN Writes in Long-Context Retrieval
- Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
- Tokenisation via Convex Relaxations
- Attention to Mamba: A Recipe for Cross-Architecture Distillation
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs
- Lost in Backpropagation: The LM Head is a Gradient Bottleneck
- A Human-Centric Framework for Data Attribution in Large Language Models
- Extracting books from production language models
- In-Context Probing for Membership Inference in Fine-Tuned Language Models
- To model human linguistic prediction, make LLMs less superhuman
- How Do LLMs Use Their Depth?
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Scaling language model size yields diminishing returns for single-message political persuasion
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
- What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
- Enough Coin Flips Can Make LLMs Act Bayesian
- Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
- Language Models Grow Less Humanlike beyond Phase Transition
- TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
- Enforcing Orderedness to Improve Feature Consistency
- Rethinking Intrinsic Dimension Estimation in Neural Representations
- Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
- A Brain-like Synergistic Core in LLMs Drives Behaviour and Learning
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- Representational Curvature Modulates Behavioral Uncertainty in Large Language Models
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
- Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures
- Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
- Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models
- Semantic Deception: When Reasoning Models Can't Compute an Addition
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
- The Deleuzian Representation Hypothesis
- From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
- Dedicated Searches for Millicharged Particles at Intensity-Frontier Facilities: SpinQuest and SHiP
- Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient Shorthand
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- Towards better dense rewards in Reinforcement Learning Applications
- Nexus: Higher-Order Attention Mechanisms in Transformers
- What Is Preference Optimization Doing, How and Why?
- UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
- Start Making Sense(s): A Developmental Probe of Attention Specialization Using Lexical Ambiguity
- Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
- A Unified Understanding of Offline Data Selection and Online Self-refining Generation for Post-training LLMs
- TrackList: Tracing Back Query Linguistic Diversity for Head and Tail Knowledge in Open Large Language Models
- Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
- Geometry of Decision Making in Language Models
- Large Language Models for Sentiment Analysis to Detect Social Challenges: A Use Case with South African Languages
- Developmental Atlas of Attention Head Specialization: Spacing, Stranding, and the Capacity Tax of BPE Tokenization
- UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition
- Steganographic Backdoor Attacks in NLP: Ultra-Low Poisoning and Defense Evasion
- GRPO Privacy Is at Risk: A Membership Inference Attack Against Reinforcement Learning With Verifiable Rewards
- Which Sparse Autoencoder Features Are Real? Model-X Knockoffs for False Discovery Rate Control
- What can LLMs tell us about the mechanisms behind polarity illusions in humans? Experiments across model scales and training steps
- Graph Representation-based Model Poisoning on the Heterogeneous Internet of Agents
- Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- In-Context Learning Without Copying
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- SCALE: Upscaled Continual Learning of Large Language Models
- FlashEVA: Accelerating LLM inference via Efficient Attention
- EL-MIA: Quantifying Membership Inference Risks of Sensitive Entities in LLMs
- Detecting Data Contamination in LLMs via In-Context Learning
- Causal Masking on Spatial Data: An Information-Theoretic Case for Learning Spatial Datasets with Unimodal Language Models
- Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
- MossNet: Mixture of State-Space Experts is a Multi-Head Attention
- E-Scores for (In)Correctness Assessment of Generative Model Outputs
- Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
- CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG
- Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers
- Will Scaling Improve Social Simulation with LLMs?
- Why Open Source? A Game-Theoretic Analysis of the AI Race
- France or Spain or Germany or France: A Neural Account of Non-Redundant Redundant Disjunctions
- A Survey on Unlearning in Large Language Models
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Relative Scaling Laws for LLMs
- PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
- Transformer Based Linear Attention with Optimized GPU Kernel Implementation
- Explaining and Mitigating Crosslingual Tokenizer Inequities
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- Correlation Dimension of Auto-Regressive Large Language Models
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- Context-level Language Modeling by Learning Predictive Context Embeddings
- An Empirical Study of Sample Selection Strategies for Large Language Model Repair
- Relative-Based Scaling Law for Neural Language Models
- CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage
- Machine Text Detectors are Membership Inference Attacks
- AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
- Blackbox Model Provenance via Palimpsestic Membership Inference
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Watermark Robustness and Radioactivity May Be at Odds in Federated Learning
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- Finding Manifolds With Bilinear Autoencoders
- Constantly Improving Image Models Need Constantly Improving Benchmarks
- Midtraining Bridges Pretraining and Posttraining Distributions
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
- End-to-End Multi-Modal Diffusion Mamba
- Towards Understanding Valuable Preference Data for Large Language Model Alignment
- The Mechanistic Emergence of Symbol Grounding in Language Models
- Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
- Influence Dynamics and Stagewise Data Attribution
- CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- Early Detection and Reduction of Memorisation for Domain Adaptation and Instruction Tuning
- A-IPO: Adaptive Intent-driven Preference Optimization
- Representational Alignment Across Model Layers and Brain Regions with Hierarchical Optimal Transport
- On the Representations of Entities in Auto-regressive Large Language Models
- Provable Training Data Identification for Large Language Models
- Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Vocabulary embeddings organize linguistic structure early in language model training
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
- Reusing Overtrained Language Models Saturates Scaling
- Training Dynamics Impact Post-Training Quantization Robustness
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
- Membership Inference Attacks on Tokenizers of Large Language Models
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
- Instability in Downstream Task Performance During LLM Pretraining
- What Scales in Cross-Entropy Scaling Law?
- Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Characterizing Model Behavior Under Synthetic Data Training: An Empirical Study Across Scales and Mixing Ratios
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
- Decomposing Attention To Find Context-Sensitive Neurons
- Convergence and Divergence of Language Models under Different Random Seeds
- Bayesian Influence Functions for Hessian-Free Data Attribution
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Mitigating Biases in Language Models via Bias Unlearning
- Negative Pre-activations Differentiate Syntax
- Measuring Sparse Autoencoder Feature Sensitivity
- Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
- Pretraining Scaling Laws for Generative Evaluations of Language Models
- Mapping Overlaps in Benchmarks through Perplexity in the Wild
- Train Once, Answer All: Many Pretraining Experiments for the Cost of One
- LLM Interpretability with Identifiable Temporal-Instantaneous Representation
- Adaptive Token-Weighted Differential Privacy for LLMs: Not All Tokens Require Equal Protection
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Tracing the Representation Geometry of Language Models from Pretraining to Post-training
- Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
- Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
- Best-of-∞ -- Asymptotic Performance of Test-Time Compute
- Towards Atoms of Large Language Models
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- Failure Modes of Maximum Entropy RLHF
- Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
- The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
- Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
- Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
- Similarity Field Theory: A Mathematical Framework for Intelligence
- Evolution of Concepts in Language Model Pre-Training
- The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- Toward Efficient Influence Function: Dropout as a Compression Tool
- Modeling Transformers as complex networks to analyze learning dynamics
- Transcoder-based Circuit Analysis for Interpretable Single-Cell Foundation Models
- Large Language Model probabilities cannot distinguish between possible and impossible language
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- Scrub It Out! Erasing Sensitive Memorization in Code Language Models via Machine Unlearning
- Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings
- Fluid Language Model Benchmarking
- KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Neural Breadcrumbs: Membership Inference Attacks on LLMs Through Hidden State and Attention Pattern Analysis
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
- Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
- Distribution-Aware Feature Selection for SAEs
- RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching
- Even Heads Fix Odd Errors: Mechanistic Discovery and Surgical Repair in Transformer Attention
- Predicting the Order of Upcoming Tokens Improves Language Modeling
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- EMSEdit: Efficient Multi-Step Meta-Learning-based Model Editing
- DTPA: Dynamic Token-level Prefix Augmentation for Controllable Text Generation
- Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
- Learning Dynamics of Meta-Learning in Small Model Pretraining
- A Survey on Data Security in Large Language Models
- How Does Controllability Emerge In Language Models During Pretraining?
- Win-k: Improved Membership Inference Attacks on Small Language Models
- LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration
- Meaning-infused grammar: Gradient Acceptability Shapes the Geometric Representations of Constructions in LLMs
- How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?
Discussions
Related