Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
2026/06/03 by Stella Biderman, Mohammad Aflah Khan, Niloofar Mireshghallah +3 · 1 voice
#cs.AI #cs.CL
paper · pdf
Abstract
What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data, objectives, architectures, and optimization dynamics. Yet much of AI research treats models as fixed artifacts, analyzing behaviors after training rather than asking why they emerge. This position paper argues that a science of AI must move beyond post-hoc fixes and study the training dynamics that produce model behavior. Such a science should support progressively stronger forms of understanding: predicting outcomes from early training signals, intervening when trajectories go wrong, and ultimately designing training procedures that more reliably produce desired properties. Scaling laws have made prediction routine for loss; the challenge is extending this success to capabilities, biases, robustness, and safety-relevant behaviors. We articulate requirements for such theories grounded in the history and philosophy of science, examine progress in mechanistic interpretability, fairness, memorization, and simplicity bias, and identify concrete open problems.
Citations
- Towards end-to-end automation of AI research
- How I Met Your Bias: Investigating Bias Amplification in Diffusion Models
- The Dead Salmons of AI Interpretability
- Start Making Sense(s): A Developmental Probe of Attention Specialization Using Lexical Ambiguity
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Explaining and Mitigating Crosslingual Tokenizer Inequities
- Hubble: a Model Suite to Advance the Study of LLM Memorization
- Blackbox Model Provenance via Palimpsestic Membership Inference
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
- Better Training Data Attribution via Better Inverse Hessian-Vector Products
- Can Interpretation Predict Behavior on Unseen Data?
- Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
- Hidden Breakthroughs in Language Model Training
- Distributional Training Data Attribution: What do Influence Functions Sample?
- Fairness Dynamics During Training
- Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural Networks
- Characterizing Pattern Matching and Its Limits on Compositional Task Structures
- A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
- MAGIC: Near-Optimal Data Attribution for Deep Learning
- Bigram Subnetworks: Mapping to Next Tokens in Transformer Language Models
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- More of the Same: Persistent Representational Harms Under Increased Representation
- Which Attention Heads Matter for In-Context Learning?
- Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders
- Sparse Autoencoders Trained on the Same Data Learn Different Features
- Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
- Why do language models perform worse for morphologically complex languages?
- Scaling Laws for Predicting Downstream Performance in LLMs
- Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
dattri: A Library for Efficient Data Attribution- The Llama 3 Herd of Models
- Demystifying Verbatim Memorization in Large Language Models
- Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
- Does Refusal Training in LLMs Generalize to the Past Tense?
- LLM Circuit Analyses Are Consistent Across Training and Scale
- A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
- Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon
- Would Deep Generative Models Amplify Bias in Future Models?
- Do Membership Inference Attacks Work on Large Language Models?
- Neural Networks Learn Statistics of Increasing Complexity
- Probing Critical Learning Dynamics of PLMs for Hate Speech Detection
- Large Language Models Relearn Removed Concepts
- LLM360: Towards Fully Transparent Open-Source LLMs
- The Transient Nature of Emergent In-Context Learning in Transformers
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Low-Resource Languages Jailbreak GPT-4
- Physics of Language Models: Part 3.2, Knowledge Manipulation
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
- FLIRT: Feedback Loop In-context Red Teaming
- The Bias Amplification Paradox in Text-to-Image Generation
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Jailbroken: How Does LLM Safety Training Fail?
- Self-Consuming Generative Models Go MAD
- On the special role of class-selective neurons in early training
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
- GPT-4 Technical Report
- A Sign That Spells: DALL-E 2, Invisual Images and The Racial Politics of Feature Space
- Linear Connectivity Reveals Generalization Strategies
- Training Compute-Optimal Large Language Models
- Training language models to follow instructions with human feedback
- An Empirical Study on Explanations in Out-of-Domain Settings
- Datamodels: Predicting Predictions from Training Data
- A Systematic Study of Bias Amplification
- The MultiBERTs: BERT Reproductions for Robustness Analysis
- An Interpretability Illusion for BERT
- On the Dangers of Stochastic Parrots
- Language (Technology) is Power: A Critical Survey of "Bias" in NLP
- Selectivity considered harmful: evaluating the causal impact of class selectivity in DNNs
- Scaling Laws for Neural Language Models
- A Constructive Prediction of the Generalization Error Across Scales
- Explanations can be manipulated and geometry is to blame
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- Sanity Checks for Saliency Maps
- Are All Languages Equally Hard to Language-Model?
- Deep Learning Scaling is Predictable, Empirically
- Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word\n Embeddings
- Deep Residual Learning for Image Recognition
- Object Detectors Emerge in Deep Scene CNNs
- ImageNet classification with deep convolutional neural networks
- Random Scaling of Emergent Capabilities
Discussions
Related