Extracting Training Data from Diffusion Models
2023/01/30 by Nicholas Carlini, Jamie Hayes, Carlini, Nicholas +15 · 9 voices · 173 citations
Computer Science · #Artificial intelligence #Computer science #Computer vision #Data mining #Data science #Diffusion #Filter (signal processing) #Generative Adversarial Networks and Image Synthesis #Generative grammar #Generative model #Geography #Machine learning #Pipeline (software) #Training (meteorology) #Training set
paper · pdf · doi:10.48550/arxiv.2301.13188
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/01/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Image diffusion models such as DALL-E 2, Imagen, and Stable Diffusion have attracted significant attention due to their ability to generate high-quality synthetic images. In this work, we show that diffusion models memorize individual images from their training data and emit them at generation time. With a generate-and-filter pipeline, we extract over a thousand training examples from state-of-the-art models, ranging from photographs of individual people to trademarked company logos. We also train hundreds of diffusion models in various settings to analyze how different modeling and data decisions affect privacy. Overall, our results show that diffusion models are much less private than prior generative models such as GANs, and that mitigating these vulnerabilities may require new advances in privacy-preserving training.
Cited by
- DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection
- Agentic Evaluation of Copyright Law Compliance
- Membership Inference Attacks for Unseen Classes
- Dominant vs. Dominated: Concept-Level Generative Collapse in Diffusion Models
- Signed Rectified Flow: Negativity-Controlled Generation
- RRAM-DP: Device-Calibrated Differential Privacy for In-Memory Edge Learning
- Points as Tori: Fast Pointwise Signed Distance for Point Clouds
- Semi-Supervised Conditional Diffusion via Label Augmentation
- Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection
- On the Interpolation Effect of Score Smoothing in Diffusion Models
- A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
- How much do language models memorize?
- On the Closed-Form of Flow Matching: Generalization Does Not Arise from Target Stochasticity
- Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
- Memorization in 3D Shape Generation: An Empirical Study
- Taxing Artificial Intelligence
- A Reinforcement Learning Approach to Synthetic Data Generation
- Generalization of Diffusion Models Arises with a Balanced Representation Space
- AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
- Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting
- Quantifying Return on Security Controls in LLM Systems
- LCMem: A Universal Model for Robust Image Memorization Detection
- Beyond Memorization: Selective Learning for Copyright-Safe Diffusion Model Training
- CAPTAIN: Semantic Feature Injection for Memorization Mitigation in Text-to-Image Diffusion Models
- Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
- Membership and Dataset Inference Attacks on Large Audio Generative Models
- PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning
- Prediction with Expert Advice under Local Differential Privacy
- SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities
- Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective
- Privacy Preserving Diffusion Models for Mixed-Type Tabular Data Generation
- Latent Diffusion Inversion Requires Understanding the Latent Space
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models
- The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models
- A PDE Perspective on Generative Diffusion Models
- A Parallel Region-Adaptive Differential Privacy Framework for Image Pixelization
- Dialogues Towards Sociologies of Generative AI
- BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation
- Cyclic Denoising Reveals Ultrastable Memories in Diffusion Models
- RA-Det: Towards Universal Detection of AI-Generated Images via Robustness Asymmetry
- Learning to Attack: Uncovering Privacy Risks in Sequential Data Releases
- On the Anisotropy of Score-Based Generative Models
- Black Box Absorption: LLMs Undermining Innovative Ideas
- Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
- Nonparametric Data Attribution for Diffusion Models
- Local Differential Privacy for Federated Learning with Fixed Memory Usage and Per-Client Privacy
- DiffEM: Learning from Corrupted Data with Diffusion Models via Expectation Maximization
- A Black-Box Debiasing Framework for Conditional Sampling
- Sensitivity, Specificity, and Consistency: A Tripartite Evaluation of Privacy Filters for Synthetic Data Generation
- Diffusion Models and the Manifold Hypothesis: Log-Domain Smoothing is Geometry Adaptive
- Adjusting Initial Noise to Mitigate Memorization in Text-to-Image Diffusion Models
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
- Knowledge Distillation Detection for Open-weights Models
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Selective Underfitting in Diffusion Models
- How Diffusion Models Memorize
- Score-based Membership Inference on Diffusion Models
- A phase transition in diffusion models reveals the hierarchical nature of data
- Copyright Infringement Detection in Text-to-Image Diffusion Models via Differential Privacy
- A Law of Data Reconstruction for Random Features (and Beyond)
- Review of Hallucination Understanding in Large Language and Vision Models
- A Unified Framework for Diffusion Model Unlearning with f-Divergence
- RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
- Responsible Diffusion: A Comprehensive Survey on Safety, Ethics, and Trust in Diffusion Models
- Latent Iterative Refinement Flow: A Geometric-Constrained Approach for Few-Shot Generation
- Adversarial machine learning :
- On the Edge of Memorization in Diffusion Models
- A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis
- A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
- Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
- Enterprise AI Must Enforce Participant-Aware Access Control
- MIA-EPT: Membership Inference Attack via Error Prediction for Tabular Data
- ReTrack: Data Unlearning in Diffusion Models through Redirecting the Denoising Trajectory
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- Fundamental Limits of Membership Inference Attacks on Machine Learning Models
- Robust Concept Erasure in Diffusion Models: A Theoretical Perspective on Security and Robustness
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- Prompt Pirates Need a Map: Stealing Seeds helps Stealing Prompts
- ArtifactGen: Benchmarking WGAN-GP vs Diffusion for Label-Aware EEG Artifact Synthesis
- Towards Post-mortem Data Management Principles for Generative AI
- Imitative Membership Inference Attack
- On the MIA Vulnerability Gap Between Private GANs and Diffusion Models
- On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
- AMCR: A Framework for Assessing and Mitigating Copyright Risks in Generative Models
- Localizing and Mitigating Memorization in Image Autoregressive Models
- The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang
- Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
- Towards 6G Intelligence: The Role of Generative AI in Future Wireless Networks
- Mitigating Data Exfiltration Attacks through Layer-Wise Learning Rate Decay Fine-Tuning
- Demystifying Foreground-Background Memorization in Diffusion Models
- SoK: Data Minimization in Machine Learning
- Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
- Data Exfiltration by Compression Attack: Definition and Evaluation on Medical Image Data
- FedMP: Tackling Medical Feature Heterogeneity in Federated Learning from a Manifold Perspective
- Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- A Survey on Data Security in Large Language Models
- LeakyCLIP: Extracting Training Data from CLIP
- LoReUn: Data Itself Implicitly Provides Cues to Improve Machine Unlearning
- Cascading and Proxy Membership Inference Attacks
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- A New One-Shot Federated Learning Framework for Medical Imaging Classification with Feature-Guided Rectified Flow and Knowledge Distillation
- AEDR: Training-Free AI-Generated Image Attribution via Autoencoder Double-Reconstruction
- FADE: Adversarial Concept Erasure in Flow Models
- 3S-Attack: Spatial, Spectral and Semantic Invisible Backdoor Attack Against DNN Models
- Reconstructing Template-Memorized Images from Natural Prompts
- FlexOlmo: Open Language Models for Flexible Data Use
- Concept-TRAK: Understanding how diffusion models learn concepts through concept-level attribution
- Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge
- Backdoors in Conditional Diffusion: Threats to Responsible Synthetic Data Pipelines
- Differentially Private Auditing Under Strategic Response
- Enhanced Generative Model Evaluation with Clipped Density and Coverage
- Low-Perplexity LLM-Generated Sequences and Where To Find Them
- Securing AI Systems: A Guide to Known Attacks and Impacts
- Inpainting is All You Need: A Diffusion-based Augmentation Method for Semi-supervised Medical Image Segmentation
- Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment
- Exploring Image Generation via Mutually Exclusive Probability Spaces and Local Correlation Hypothesis
- SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt Learning
- Diffusion models under low-noise regime
- Large-Scale Training Data Attribution for Music Generative Models via Unlearning
- When Diffusion Models Memorize: Inductive Biases in Probability Flow of Minimum-Norm Shallow Neural Nets
- Black-Box Privacy Attacks on Shared Representations in Multitask Learning
- Discrete Diffusion in Large Language and Multimodal Models: A Survey
- Restoring Gaussian Blurred Face Images for Deanonymization Attacks
- GaussMarker: Robust Dual-Domain Watermark for Diffusion Models
- Prompt-Guided Latent Diffusion with Predictive Class Conditioning for 3D Prostate MRI Generation
- The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing
- Learning to Weight Parameters for Training Data Attribution
- PixCell: A generative foundation model for digital histopathology images
- Membership Inference Attacks on Sequence Models
- Quantifying Cross-Modality Memorization in Vision-Language Models
- Survey on the Evaluation of Generative Models in Music
- Safer Prompts: Reducing Risks from Memorization in Visual Generative AI
- SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples
- Trade-offs in Data Memorization via Strong Data Processing Inequalities
- Adversarial Attacks in Multimodal Systems: A Practitioner's Survey
- Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis
- Differential privacy for medical deep learning: methods, tradeoffs, and deployment implications
- Interpreting Large Text-to-Image Diffusion Models with Dictionary Learning
- Bayesian Perspective on Memorization and Reconstruction
- FPAN: Mitigating Replication in Diffusion Models through the Fine-Grained Probabilistic Addition of Noise to Token Embeddings
- Kernel-Smoothed Scores for Denoising Diffusion: A Bias-Variance Study
- MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection
- What is Adversarial Training for Diffusion Models?
- Understanding Generalization in Diffusion Distillation via Probability Flow Distance
- CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
- A Survey on Progress in LLM Alignment from the Perspective of Reward Design
- Querying Kernel Methods Suffices for Reconstructing their Training Data
- Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
- Deeper Diffusion Models Amplify Bias
- Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
- One-Step Offline Distillation of Diffusion-based Models via Koopman Modeling
- On Membership Inference Attacks in Knowledge Distillation
- Conditionally Strongly Log-Concave Generative Models
- CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs
- AC-LoRA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs
- Identifying Memorization of Diffusion Models through p-Laplace Analysis
- Demystifying Diffusion Policies: Action Memorization and Simple Lookup Table Alternatives
- The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
- Morphological Addressing of Identity Basins in Text-to-Image Diffusion Models
- We Should Separate Memorization from Copyright
- Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
- Equalized Generative Treatment: Matching f-divergences for Fairness in Generative Models
- Beyond Anonymization: Object Scrubbing for Privacy-Preserving 2D and 3D Vision Tasks
- DualOptim: Enhancing Efficacy and Stability in Machine Unlearning with Dual Optimizers
- Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds
- AI Safety Should Prioritize the Future of Work
- ControlNET: A Firewall for RAG-based LLM System
- Large Language Models Could Be Rote Learners
- GenEAva: Generating Cartoon Avatars with Fine-Grained Facial Expressions from Realistic Diffusion-based Faces
- Code Generation with Small Language Models: A Codeforces-Based Study
Discussions
- Extracting training data from diffusion models [hn, 163 points, 309 comments]
- no. arxiv.org/abs/2301.13188 [bsky, 7 points, 2 comments]
- this research paper from 2023 should be similarly interesting. i would go as far as to say the only thing modern “ai” companies have brought to the table to forward the tech is their disregard for cop [bsky, 6 points, 0 comments]
- He do not address the fact that in the paper they address the copyright point and support it independent of the opinion a posterior participants do not change it. and he ignores the teorical corpus ov [bsky, 5 points, 1 comments]
- Extracting Training Data from Diffusion Models [hn, 3 points, 0 comments]
- 潜在拡散モデルで同じ画像を学習時に100回使うと記憶して酷似画像を生成するという論文。ピクセルベースの拡散モデルでは記憶する割合がさらに増えると書かれています。 arxiv.org/abs/2301.13188 私がSD1.5を使用した経験では、データセットに多数画像のあるマリオやプーチンは、名前を書くだけで何度も生成できました。 記憶して再現できるAIを悪用させないために、再現させない技術と法規 [bsky, 2 points, 0 comments]
- Stable Diffusion Memorizes Training Points [bsky, 0 points, 0 comments]
- Esto es falso, las IAGs no imitan ningún proceso mental. Son algoritmos que funcionan a base de scaping data. arxiv.org/abs/2301.13188 [bsky, 0 points, 1 comments]
- I don't want to go too far into this because I don't have time right now, but there's plenty of instances of modern generative models replicating artists work. But for now, here's a paper on a couple [bsky, 0 points, 1 comments]
Related