vix.ing · top · new · best · stats

Extracting Training Data from Diffusion Models

2023/01/30 by Nicholas Carlini, Jamie Hayes, Carlini, Nicholas +15 · 9 voices · 173 citations
Computer Science · #Artificial intelligence #Computer science #Computer vision #Data mining #Data science #Diffusion #Filter (signal processing) #Generative Adversarial Networks and Image Synthesis #Generative grammar #Generative model #Geography #Machine learning #Pipeline (software) #Training (meteorology) #Training set

paper · pdf · doi:10.48550/arxiv.2301.13188

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2023/01/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Image diffusion models such as DALL-E 2, Imagen, and Stable Diffusion have attracted significant attention due to their ability to generate high-quality synthetic images. In this work, we show that diffusion models memorize individual images from their training data and emit them at generation time. With a generate-and-filter pipeline, we extract over a thousand training examples from state-of-the-art models, ranging from photographs of individual people to trademarked company logos. We also train hundreds of diffusion models in various settings to analyze how different modeling and data decisions affect privacy. Overall, our results show that diffusion models are much less private than prior generative models such as GANs, and that mitigating these vulnerabilities may require new advances in privacy-preserving training.

Cited by

Discussions

Related