AI models collapse when trained on recursively generated data
2024/07/24 by Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao +3 · 6 voices · 157 citations
Computer Science · #Topic Modeling #Generative Adversarial Networks and Image Synthesis #Domain Adaptation and Few-Shot Learning
paper · pdf · doi:10.1038/s41586-024-07566-y
Abstract
) demonstrated high performance across a variety of language tasks. ChatGPT introduced such language models to the public. It is now clear that generative artificial intelligence (AI) such as large language models (LLMs) is here to stay and will substantially change the ecosystem of online text and images. Here we consider what may happen to GPT-n once LLMs contribute much of the text found online. We find that indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear. We refer to this effect as 'model collapse' and show that it can occur in LLMs as well as in variational autoencoders (VAEs) and Gaussian mixture models (GMMs). We build theoretical intuition behind the phenomenon and portray its ubiquity among all learned generative models. We demonstrate that it must be taken seriously if we are to sustain the benefits of training from large-scale data scraped from the web. Indeed, the value of data collected about genuine human interactions with systems will be increasingly valuable in the presence of LLM-generated content in data crawled from the Internet.
Cited by
- AI as Governance
- Can AI weather models predict out-of-distribution gray swan tropical cyclones?
- Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- AI Contagion in Social Networks
- The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis
- Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture
- Data and trained models for "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication"
- Optimal Self-Distillation for Rectified Flow via Linear Probing
- A framework for single and multi-agent human-AI curiosity ecosystems
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- False authorship: an explorative case study around an AI-generated article published under my name
- Misinformation as strategy: Epistemic consequences and the undermining of shared truth
- AI-Mediated Communication Can Steer Collective Opinion
- Machine understanding
- Not-So-Strange Love: Language Models and Generative Linguistic Theories are More Compatible than They Appear
- Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing
- LLM hallucinations in the wild: Large-scale evidence from non-existent citations
- Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated Probabilities
- Mathematical methods and human thought in the age of AI
- The billion-dollar case for sustaining palaeontology’s digital databases
- Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network
- Linguists should learn to love speech-based deep learning models
- You can’t fight in here! This is BBS!
- Iterative Compositional Data Generation for Robot Control
- LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
- Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
- Pre-training under infinite compute
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- The wall confronting large language models
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial Mislabeling
- Wikipedia Contributions in the Wake of ChatGPT
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
- How linguistics learned to stop worrying and love the language models
- AI is trash
- "An Endless Stream of AI Slop": How Developers Discuss the Burden of AI-Assisted Software Development
- From revolution to evolution: What generative AI really means for language learning
- From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research
- Understanding how users may work around algorithmic bias
- State-dependent error correlations shape voting thresholds in committees of AI agents
- The One-Word Census: Answer-Choice Conformity Across 44 Language Models
- Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI
- Bridging the Gap on AI-Assisted Scientific Software Development Through Transparency and Traceability
- GoldenFuzz: Generative Golden Reference Hardware Fuzzing
- The Silent Scholar Problem: A Probabilistic Framework for Breaking Epistemic Asymmetry in LLM Agents
- Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
- Adaptive Probability Flow Residual Minimization for High-Dimensional Fokker-Planck Equations
- Sharing Knowledge without Sharing Data: Stitches can improve ensembles of disjointly trained models
- Epistemic diversity across language models mitigates knowledge collapse
- Large language models are not about natural language
- ALIGN-FL: Architecture-independent Learning through Invariant Generative component sharing in Federated Learning
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Entropy Collapse: A Universal Failure Mode of Intelligent Systems
- On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
- CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation
- Chasing Shadows: Pitfalls in LLM Security Research
- ValuePilot: A Two-Phase Framework for Value-Driven Decision-Making
- Detecting Perspective Shifts in Multi-agent Systems
- Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
- Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
- Artificial Intelligence / Human Intelligence: Who Controls Whom?
- The Evolutionary Ecology of Software: Constraints, Innovation, and the AI Disruption
- Preventing Model Collapse via Contraction-Conditioned Neural Filters
- Guiding Generative Models for Protein Design: Prompting, Steering and Aligning
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- Large Language Models as Search Engines: Societal Challenges
- Generative AI in Sociological Research: State of the Discipline
- Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
- Rethinking Data Value: Asymmetric Data Shapley for Structure-Aware Valuation in Data Markets and Machine Learning Pipelines
- Science Behind a Paywall: Restricted Access Limits the Promise of Artificial Intelligence
- Detecting Generated Images by Fitting Natural Image Distributions
- Optimizing Diversity and Quality through Base-Aligned Model Collaboration
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- Why Less is More (Sometimes): A Theory of Data Curation
- Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
- A Criminology of Machines
- The Efficiency Costs of Information Assurance in AI-Enabled Labor Markets: Evidence from LinkedIn's Policy Changes
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
- F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill
- Governance of Generative AI
- GAIDeT (Generative AI Delegation Taxonomy): A taxonomy for humans to delegate tasks to generative artificial intelligence in scientific research and publishing
- Engaging the unengaged: Differential effects of AI-driven climate communication across audiences
- Large language models and the problem of rhetorical debt
- The tragedy of the cognitive commons: collective intelligence beyond AI-induced knowledge collapse
- A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
- JASPAR 2026: expansion of transcription factor binding profiles and integration of deep learning models
- Survey of Cultural Awareness in Language Models: Text and Beyond
- Artificial intelligence, calculative reason, and technical domination: lessons from Husserl, Heidegger, and Marcuse
- The rise of the research automaton: science as process or product in the era of generative AI?
- The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
- Retrieval Collapses When AI Pollutes the Web
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- The promise and limitations of using GenAI to reduce climate scepticism
- Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation
- When Your AI Agent Succumbs to Peer-Pressure: Studying Opinion-Change Dynamics of LLMs
- Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
- Diffusion Models as Dataset Distillation Priors
- Fine-tuning Flow Matching Generative Models with Intermediate Feedback
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Making Power Explicable in AI: Analyzing, Understanding, and Redirecting Power to Operationalize Ethics in AI Technical Practice
- The social consequences of AI delegation
- KORMo: Korean Open Reasoning Model for Everyone
- Beyond Real Data: Synthetic Data through the Lens of Regularization
- High-dimensional Analysis of Synthetic Data Selection
- Evaluating generative AI’s potential to dispel misinformation about wind farms
- Artificial Intelligence in Detecting Statistical Errors: Implications for Authors, Reviewers, and Editors
- When Do Credal Sets Stabilize? Fixed-Point Theorems for Credal Set Updates
- AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
- On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
- Neon: Negative Extrapolation From Self-Training Improves Image Generation
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- LLM-NAS: LLM-driven Hardware-Aware Neural Architecture Search
- Think Less, Label Better: Multi-Stage Domain-Grounded Synthetic Data Generation for Fine-Tuning Large Language Models in Telecommunications
- PCPO: Proportionate Credit Policy Optimization for Aligning Image Generation Models
- Learning in an Echo Chamber: Online Learning with Replay Adversary
- Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
- Generative AI and the information commons: controversy, copyright, and closure
- Using conversational AI to reduce science skepticism
- Theory Is All You Need: AI, Human Cognition, and Causal Reasoning
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge Regression
- Bridging the Gap Between Scientific Laws Derived by AI Systems and Canonical Knowledge via Abductive Inference with AI-Noether
- First-Extinction Law for Resampling Processes
- Are Language Models Models?
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- Lessons from complex systems science for AI governance
- Generative AI Robs Students of the Joy of Learning
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- Generative artificial intelligence (AI) and its performance at indexing tasks
- Ghost in the cache: How data decay shapes the unseen landscape of AI memory
- Purpose before policy: academic integrity, generative AI, and rhetorical stance
- Hybrid Data can Enhance the Utility of Synthetic Data for Training Anti-Money Laundering Models
- The Narcissus Hypothesis: Descending to the Rung of Illusion
- A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
- The Even Sheen of AI: Kitsch, LLMs, and Homogeneity
- Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
- Human intelligence versus artificial intelligence in classifying economics research articles: exploratory evidence
- Multi-Task Diffusion Approach For Prediction of Glioma Tumor Progression
- ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
- Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
- A biologically inspired separable learning vision model for real-time traffic object perception in Dark
- Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
- Generative KI für TA
- The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
- FuXi-TC: A generative framework integrating deep learning and physics-based models for improved tropical cyclone forecasts
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
- The Ramon Llull's Thinking Machine for Automated Ideation
- A Theory of Information, Variation, and Artificial Intelligence
- The Epistemic Impact of Large Language Models on Policymaking
- DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
- Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
Discussions
- AI models will likely be trained on AI-generated content; yet new research shows this creates a destructive cycle called model collapse where each generation loses ability to capture rare events and h [bsky, 62 points, 6 comments]
- Recent reading #academicSky Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., Gal, Y., 2024. AI models collapse when trained on recursively generated data. Nature 631, 755–759. doi. [bsky, 11 points, 0 comments]
- We have seen technology improve. (Think VHS -> DVD -> Blue-Ray -> 4K.) There are reasons to think generative AI will NOT follow this trend. One of those reasons is called "Model collapse." doi.org/1 [bsky, 5 points, 0 comments]
- I'm sad you or anyone has to put up with this. It does appear to be a battle of LLM-based bots which I'm worried may make the future of informed commentary quite difficult. Which I guess in the end is [bsky, 1 points, 0 comments]
- Shumailov, I., Shumaylov, Z., Zhao, Y. et al. AI models collapse when trained on recursively generated data. Nature 631, 755–759 (2024). doi.org/10.1038/s415... [bsky, 0 points, 0 comments]
- Plus: there is emerging evidence that the quality of LLMs degrades when they are trained on lower quality material, especially when it is created through LLMs. This suggests a persistent need for high [bsky, 0 points, 1 comments]
Related