Are Emergent Abilities of Large Language Models a Mirage?
2023/04/28 by Rylan Schaeffer, Brando Miranda, Schaeffer, Rylan +3 · 20 voices · 151 citations
Computer Science · Engineering · Mathematics · Psychology · #Artificial intelligence #Cognitive psychology #Computer science #Engineering #Epistemology #Explainable Artificial Intelligence (XAI) #Mathematics #Metric (unit) #Natural Language Processing Techniques #Property (philosophy) #Psychology #Scale (ratio) #Scaling #Simple (philosophy) #Task (project management) #Test (biology) #Topic Modeling #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2304.15004
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recent work claims that large language models display emergent abilities, abilities not present in smaller-scale models that are present in larger-scale models. What makes emergent abilities intriguing is two-fold: their sharpness, transitioning seemingly instantaneously from not present to present, and their unpredictability, appearing at seemingly unforeseeable model scales. Here, we present an alternative explanation for emergent abilities: that for a particular task and model family, when analyzing fixed model outputs, emergent abilities appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale. Specifically, nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes in model performance. We present our alternative explanation in a simple mathematical model, then test it in three complementary ways: we (1) make, test and confirm three predictions on the effect of metric choice using the InstructGPT/GPT-3 family on tasks with claimed emergent abilities; (2) make, test and confirm two predictions about metric choices in a meta-analysis of emergent abilities on BIG-Bench; and (3) show to choose metrics to produce never-before-seen seemingly emergent abilities in multiple vision tasks across diverse deep networks. Via all three analyses, we provide evidence that alleged emergent abilities evaporate with different metrics or with better statistics, and may not be a fundamental property of scaling AI models.
Cited by
- Context Is King: How In-Context Specification Shapes the Geometry of Concepts
- Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
- The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory
- Parameter-Efficient Continual Fine-Tuning: A Survey
- The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
- Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning
- Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
- RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
- Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
- Lost in Context: Addressing Context Anxiety in Large Language Models
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- The Illusion of Insight in Reasoning Models
- Is In-Context Learning Learning?
- Large Language Models and Emergence: A Complex Systems Perspective
- Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
- Greedy dynamical meta-learning
- Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
- Hierarchical Grading in Large Language Models
- Market Design for AI: Beyond the Copyright Binary
- Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
- Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
- Teaching and Critiquing Conceptualization and Operationalization in NLP
- Scaling Laws for Code: Every Programming Language Matters
- Curriculum Guided Massive Multi Agent System Solving For Robust Long Horizon Tasks
- Single-Agent Scaling Fails Multi-Agent Intelligence: Towards Foundation Models with Native Multi-Agent Intelligence
- Sequential Enumeration in Large Language Models
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
- Beyond Mimicry: Preference Coherence in LLMs
- Evidence of Phase Transitions in Small Transformer-Based Language Models
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
- Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
- Importance-Aware Data Selection for Efficient LLM Instruction Tuning
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
- CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
- Will Scaling Improve Social Simulation with LLMs?
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
- Relative Scaling Laws for LLMs
- The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- Relative-Based Scaling Law for Neural Language Models
- Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
- Evaluating LLM Reasoning Beyond Correctness and CoT
- UniCode: Augmenting Evaluation for Code Reasoning
- Position: Require Frontier AI Labs To Release Small "Analog" Models
- The Mechanistic Emergence of Symbol Grounding in Language Models
- KORMo: Korean Open Reasoning Model for Everyone
- Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
- Inductive Bias and Spectral Properties of Single-Head Attention in High Dimensions
- Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
- Pretraining Scaling Laws for Generative Evaluations of Language Models
- Review of Hallucination Understanding in Large Language and Vision Models
- Predicting LLM Reasoning Performance with Small Proxy Model
- The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
- ALIMA – Ein RAG-basiertes System zur LLM-gestützten Sacherschließung: Prototypentwicklung und erste Erfahrungen aus der Praxis
- On the Edge of Memorization in Diffusion Models
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
- From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
- Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
- Artificial or Human Intelligence?
- Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
- MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
- The Ramon Llull's Thinking Machine for Automated Ideation
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Equinox: Holistic Fair Scheduling in Serving Large Language Models
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- A Survey on Agentic Service Ecosystems: Measurement, Analysis, and Optimization
- Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
- How Does Controllability Emerge In Language Models During Pretraining?
- Metric assessment protocol in the context of answer fluctuation on MCQ tasks
- Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
- What Does it Mean for a Neural Network to Learn a "World Model"?
- Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
- CASCADE: LLM-Powered JavaScript Deobfuscator at Google
- Thinking Isn't an Illusion: Overcoming the Limitations of Reasoning Models via Tool Augmentations
- A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
- Olica: Efficient Structured Pruning of Large Language Models without Retraining
- AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs
- Asymptotic theory of in-context learning by linear attention
- Meta-Learning Transformers to Improve In-Context Generalization
- A validity-guided workflow for robust large language model research in psychology
- Energy-Based Transformers are Scalable Learners and Thinkers
- The Thin Line Between Comprehension and Persuasion in LLMs
- Decomposing Prediction Mechanisms for In-Context Recall
- Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
- Not All Explanations for Deep Learning Phenomena Are Equally Valuable
- Positioning AI Tools to Support Online Harm Reduction Practice: Applications and Design Directions
- Emergence of Text Readability in Vision Language Models
- Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
- Bayesian Social Deduction with Graph-Informed Language Models
- From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology
- Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
- Complexity Scaling Laws for Neural Models using Combinatorial Optimization
- Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
- Can Theoretical Physics Research Benefit from Language Agents?
- Distillation Robustifies Unlearning
- Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning
- The Guanyin Protocol: A Framework for Immediately Establishing an Understanding of Both Causality and Compassion in LLM Systems Using Semantic Anchoring
- Fodor and Pylyshyn's Legacy: Still No Human-like Systematic Compositionality in Neural Networks
- Quiet Feature Learning in Algorithmic Tasks
- The Steganographic Potentials of Language Models
- Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors
- Curse of High Dimensionality Issue in Transformer for Long-context Modeling
- Learning Extrapolative Sequence Transformations from Markov Chains
- Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
- Towards Semantic Integration of Opinions: Unified Opinion Concepts Ontology and Extraction Task
- Small Models, Smarter Learning: The Power of Joint Task Training
- The emergence of sparse attention: impact of data distribution and benefits of repetition
- Pixels and Predictions: Potential of GPT-4V in Meteorological Imagery Analysis and Forecast Communication
- Small-to-Large Generalization: Data Influences Models Consistently Across Scale
- The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models
- Scaling Laws for State Dynamics in Large Language Models
- Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
- No Consciousness? No Meaning (and no AGI!)
- Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
- CurveShift: Is Agent Progress Scalar? Separating Level from Shape
- Towards Contamination Resistant Benchmarks
- A Hamiltonian-Inspired Local-Operator Ansatz for Slimming Large Language Models
- The production of meaning in the processing of natural language
- The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail
- Leaderboard Incentives: Model Rankings under Strategic Post-Training
- The Brain Abstracted
- Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
- Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining
- On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
- Reasoning Capabilities and Invariability of Large Language Models
- LENSLLM: Unveiling Fine-Tuning Dynamics for LLM Selection
- Arithmetic Pedagogy for Language Models
- Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
- Brevity Constraints Reverse Performance Hierarchies in Language Models
- The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next
- Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance
- Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock
- Nested Learning: The Illusion of Deep Learning Architectures
- Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation
- A Model Zoo on Phase Transitions in Neural Networks
- When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
- Protoreasoning in Tiny Transformers
- PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
- Audit Cards: Contextualizing AI Evaluations
- The Ignition Index: Measuring Global Workspace Dynamics in Language Models
- Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
- Large Language Models Could Be Rote Learners
- Large language model [wikipedia]
Discussions
- Are emergent abilities of large language models a mirage? [hn, 154 points, 130 comments]
- Diese Studie zweifelt die "emergenten Fähigkeiten" von LLMs, also plötzliche, sprunghafte Leistungssteigerungen, sogar an und hält sie für eine reine Methoden-Täuschung. arxiv.org/abs/2304.150... [bsky, 20 points, 1 comments]
- When the title of the paper is a question, you already know the answer. :). A best paper at NeurIPS, providing useful and insightful analysis: arxiv.org/abs/2304.15004 [bsky, 14 points, 0 comments]
- Great read: part of the research on LLM remains poor, no doubt because of the incentives and the private nature of some of the research [bsky, 6 points, 0 comments]
- arxiv.org/abs/2304.15004 ? [bsky, 4 points, 2 comments]
- okay this paper rules https://arxiv.org/abs/2304.15004 [bsky, 3 points, 3 comments]
- Jeg troede Emergent Abilities var blevet debunked af paperet "Are Emergent Abilities of Large Language Models a Mirage?" (arxiv.org/abs/2304.15004). Folk referer dog stadigvæk til Emergent Abilities [bsky, 3 points, 1 comments]
- "We present an alternative explanation that emergent abilities appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale." arxiv.org/abs/2304 [bsky, 2 points, 0 comments]
- https://bsky.app/profile/nsaphra.bsky.social/post/3li2xws4qjc2l [bsky, 2 points, 1 comments]
- Yes, extrapolation goes beyond D, and i suppose that LLMs can't go beyond D, because how should they? Training data is fix, latent space is fix, for a model there is no beyond D. See also "Are Emergen [bsky, 2 points, 1 comments]
- arxiv.org/abs/2304.15004 [bsky, 2 points, 0 comments]
- Except that scaling has hit a plateau. Also: [bsky, 1 points, 2 comments]
- Are Emergent Abilities of Large Language Models a Mirage? (2023) [hn, 1 points, 0 comments]
- "[W]e provide evidence that alleged emergent abilities evaporate with different metrics or with better statistics, and may not be a fundamental property of scaling AI models." #AI #LargeLanguageModel [bsky, 1 points, 0 comments]
- You're assigning a lot of characteristics to AI that simply aren't there at the moment, nor is there any guarantee they will be. Emergent abilities may not even exist with LLMs. arxiv.org/abs/2304.1 [bsky, 0 points, 1 comments]
- Are Emergent Abilities of Large Language Models a Mirage? Presents an alternative explanation for emergent abilities: one can choose a metric which leads to the inference of an emergent ability or a [bsky, 0 points, 0 comments]
- A recent paper on emergence questions if LLMs really have emergent properties after all: https:// arxiv.org/abs/2304.15004 # LLM # NLP # NLProc # metrics # arxiv # arxiv_2304_15004 [mastodon, 0 points, 0 comments]
- Research paper casts doubt on Emergent Abilities of Large Language Models. Using 3 different analyses ... "we find strong supporting evidence that emergent abilities may not be a fundamental property [bsky, 0 points, 0 comments]
- Sudden emergence of new capabilities (e.g. arithmetic) in LLMs might just be a measurement artifact, find Schaeffer et al. https://arxiv.org/abs/2304.15004 [bsky, 0 points, 0 comments]
- Donc, non, les très grands modèles ne sont pas plus créatifs. D'ailleurs, des travaux établissent qu'avec des métriques et une analyse sérieuses, les "capacités émergentes" que l'on prête aux très gra [bsky, 0 points, 1 comments]
Related