Are Emergent Abilities of Large Language Models a Mirage?
2023/04/28 by Rylan Schaeffer, Brando Miranda, Schaeffer, Rylan +3 · 20 voices · 76 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2304.15004
openalex publication_date 2023/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recent work claims that large language models display emergent abilities, abilities not present in smaller-scale models that are present in larger-scale models. What makes emergent abilities intriguing is two-fold: their sharpness, transitioning seemingly instantaneously from not present to present, and their unpredictability, appearing at seemingly unforeseeable model scales. Here, we present an alternative explanation for emergent abilities: that for a particular task and model family, when analyzing fixed model outputs, emergent abilities appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale. Specifically, nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes in model performance. We present our alternative explanation in a simple mathematical model, then test it in three complementary ways: we (1) make, test and confirm three predictions on the effect of metric choice using the InstructGPT/GPT-3 family on tasks with claimed emergent abilities; (2) make, test and confirm two predictions about metric choices in a meta-analysis of emergent abilities on BIG-Bench; and (3) show to choose metrics to produce never-before-seen seemingly emergent abilities in multiple vision tasks across diverse deep networks. Via all three analyses, we provide evidence that alleged emergent abilities evaporate with different metrics or with better statistics, and may not be a fundamental property of scaling AI models.
Cited by
- Context Is King: How In-Context Specification Shapes the Geometry of Concepts
- Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
- The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory
- Parameter-Efficient Continual Fine-Tuning: A Survey
- The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
- Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning
- Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
- RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
- Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
- Lost in Context: Addressing Context Anxiety in Large Language Models
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- The Illusion of Insight in Reasoning Models
- Is In-Context Learning Learning?
- Large Language Models and Emergence: A Complex Systems Perspective
- Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
- Greedy dynamical meta-learning
- Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
- Hierarchical Grading in Large Language Models
- Market Design for AI: Beyond the Copyright Binary
- Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
- Rethinking Knowledge Distillation in Collaborative Machine Learning: Memory, Knowledge, and Their Interactions
- Teaching and Critiquing Conceptualization and Operationalization in NLP
- Scaling Laws for Code: Every Programming Language Matters
- Curriculum Guided Massive Multi Agent System Solving For Robust Long Horizon Tasks
- Single-Agent Scaling Fails Multi-Agent Intelligence: Towards Foundation Models with Native Multi-Agent Intelligence
- Sequential Enumeration in Large Language Models
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
- Beyond Mimicry: Preference Coherence in LLMs
- Evidence of Phase Transitions in Small Transformer-Based Language Models
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
- Optimal Attention Temperature Enhances In-Context Learning under Distribution Shift
- Importance-Aware Data Selection for Efficient LLM Instruction Tuning
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
- CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
- Will Scaling Improve Social Simulation with LLMs?
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
- Relative Scaling Laws for LLMs
- The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
- Capability Ceilings in Autoregressive Language Models: Empirical Evidence from Knowledge-Intensive Tasks
- Relative-Based Scaling Law for Neural Language Models
- Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
- Evaluating LLM Reasoning Beyond Correctness and CoT
- UniCode: A Framework for Generating High Quality Competitive Coding Problems
- Position: Require Frontier AI Labs To Release Small "Analog" Models
- The Mechanistic Emergence of Symbol Grounding in Language Models
- KORMo: Korean Open Reasoning Model for Everyone
- Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
- Inductive Bias and Spectral Properties of Single-Head Attention in High Dimensions
- Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
- Pretraining Scaling Laws for Generative Evaluations of Language Models
- Review of Hallucination Understanding in Large Language and Vision Models
- Predicting LLM Reasoning Performance with Small Proxy Model
- The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models
- ALIMA – Ein RAG-basiertes System zur LLM-gestützten Sacherschließung: Prototypentwicklung und erste Erfahrungen aus der Praxis
- On the Edge of Memorization in Diffusion Models
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
- From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
- Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
- Artificial or Human Intelligence?
- Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
- MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
- The Ramon Llull's Thinking Machine for Automated Ideation
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Equinox: Holistic Fair Scheduling in Serving Large Language Models
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- A Survey on Agentic Service Ecosystems: Measurement, Analysis, and Optimization
- Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
- How Does Controllability Emerge In Language Models During Pretraining?
- Large language model [wikipedia]
Discussions
- Are emergent abilities of large language models a mirage? [hn, 154 points, 130 comments]
- Diese Studie zweifelt die "emergenten Fähigkeiten" von LLMs, also plötzliche, sprunghafte Leistungssteigerungen, sogar an und hält sie für eine reine Methoden-Täuschung. arxiv.org/abs/2304.150... [bsky, 20 points, 1 comments]
- When the title of the paper is a question, you already know the answer. :). A best paper at NeurIPS, providing useful and insightful analysis: arxiv.org/abs/2304.15004 [bsky, 14 points, 0 comments]
- Great read: part of the research on LLM remains poor, no doubt because of the incentives and the private nature of some of the research [bsky, 6 points, 0 comments]
- arxiv.org/abs/2304.15004 ? [bsky, 4 points, 2 comments]
- okay this paper rules https://arxiv.org/abs/2304.15004 [bsky, 3 points, 3 comments]
- Jeg troede Emergent Abilities var blevet debunked af paperet "Are Emergent Abilities of Large Language Models a Mirage?" (arxiv.org/abs/2304.15004). Folk referer dog stadigvæk til Emergent Abilities [bsky, 3 points, 1 comments]
- "We present an alternative explanation that emergent abilities appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale." arxiv.org/abs/2304 [bsky, 2 points, 0 comments]
- https://bsky.app/profile/nsaphra.bsky.social/post/3li2xws4qjc2l [bsky, 2 points, 1 comments]
- Yes, extrapolation goes beyond D, and i suppose that LLMs can't go beyond D, because how should they? Training data is fix, latent space is fix, for a model there is no beyond D. See also "Are Emergen [bsky, 2 points, 1 comments]
- arxiv.org/abs/2304.15004 [bsky, 2 points, 0 comments]
- Except that scaling has hit a plateau. Also: [bsky, 1 points, 2 comments]
- Are Emergent Abilities of Large Language Models a Mirage? (2023) [hn, 1 points, 0 comments]
- "[W]e provide evidence that alleged emergent abilities evaporate with different metrics or with better statistics, and may not be a fundamental property of scaling AI models." #AI #LargeLanguageModel [bsky, 1 points, 0 comments]
- You're assigning a lot of characteristics to AI that simply aren't there at the moment, nor is there any guarantee they will be. Emergent abilities may not even exist with LLMs. arxiv.org/abs/2304.1 [bsky, 0 points, 1 comments]
- Are Emergent Abilities of Large Language Models a Mirage? Presents an alternative explanation for emergent abilities: one can choose a metric which leads to the inference of an emergent ability or a [bsky, 0 points, 0 comments]
- A recent paper on emergence questions if LLMs really have emergent properties after all: https:// arxiv.org/abs/2304.15004 # LLM # NLP # NLProc # metrics # arxiv # arxiv_2304_15004 [mastodon, 0 points, 0 comments]
- Research paper casts doubt on Emergent Abilities of Large Language Models. Using 3 different analyses ... "we find strong supporting evidence that emergent abilities may not be a fundamental property [bsky, 0 points, 0 comments]
- Sudden emergence of new capabilities (e.g. arithmetic) in LLMs might just be a measurement artifact, find Schaeffer et al. https://arxiv.org/abs/2304.15004 [bsky, 0 points, 0 comments]
- Donc, non, les très grands modèles ne sont pas plus créatifs. D'ailleurs, des travaux établissent qu'avec des métriques et une analyse sérieuses, les "capacités émergentes" que l'on prête aux très gra [bsky, 0 points, 1 comments]
Related