Scalable Extraction of Training Data from (Production) Language Models
2023/11/28 by Milad Nasr, Nicholas Carlini, Nasr, Milad +17 · 21 voices · 111 citations
Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.48550/arxiv.2311.17035
Abstract
This paper studies extractable memorization: training data that an adversary can efficiently extract by querying a machine learning model without prior knowledge of the training dataset. We show an adversary can extract gigabytes of training data from open-source language models like Pythia or GPT-Neo, semi-open models like LLaMA or Falcon, and closed models like ChatGPT. Existing techniques from the literature suffice to attack unaligned models; in order to attack the aligned ChatGPT, we develop a new divergence attack that causes the model to diverge from its chatbot-style generations and emit training data at a rate 150x higher than when behaving properly. Our methods show practical attacks can recover far more data than previously thought, and reveal that current alignment techniques do not eliminate memorization.
Cited by
- Implicit Reasoning Steering via Concept Chaining
- Extracting books from production language models
- User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
- A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
- How much do language models memorize?
- Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
- Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
- Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
- Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
- Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization
- ContextLeak: Auditing Leakage in Private In-Context Learning Methods
- CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs
- On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
- When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- PrivCode: When Code Generation Meets Differential Privacy
- Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs
- Quantifying the Privacy Implications of High-Fidelity Synthetic Network Traffic
- Large Language Models as Search Engines: Societal Challenges
- For Those Who May Find Themselves on the Red Team
- Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
- Remembering Unequally: Global and Disciplinary Bias in LLM-Generated Co-Authorship Networks
- RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipeline
- Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
- PrivacyGuard: A Modular Framework for Privacy Auditing in Machine Learning
- Leverage Unlearning to Sanitize LLMs
- CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage
- Extracting alignment data in open models
- An Investigation of Memorization Risk in Healthcare Foundation Models
- Quantifying Information Disclosure During Gradient Descent Using Gradient Uniqueness
- UpSafe^∘C: Upcycling for Controllable Safety in Large Language Models
- sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing
- The Model's Language Matters: A Comparative Privacy Analysis of LLMs
- Memorization in Language Models through the Lens of Intrinsic Dimension
- (Token-Level) InfoRMIA: Stronger Membership Inference and Memorization Assessment for LLMs
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
- The potential for AI to revolutionize conservation: a horizon scan
- External Data Extraction Attacks against Retrieval-Augmented Large Language Models
- Adaptive Token-Weighted Differential Privacy for LLMs: Not All Tokens Require Equal Protection
- Federated Learning of Quantile Inference under Local Differential Privacy
- No Prior, No Leakage: Revisiting Reconstruction Attacks in Trained Neural Networks
- GEP: A GCG-Based method for extracting personally identifiable information from chatbots built on small language models
- On the Edge of Memorization in Diffusion Models
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- Enterprise AI Must Enforce Participant-Aware Access Control
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- Why Data Anonymization Has Not Taken Off
- Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications
- Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
- Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- The Influence of Code Comments on the Perceived Helpfulness of Stack Overflow Posts
- A Study of Privacy-preserving Language Modeling Approaches
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
- Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- PETLP: A Privacy-by-Design Pipeline for Social Media Data in AI Research
- Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment
- Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
- Understanding Privacy Norms Around LLM-Based Chatbots: A Contextual Integrity Perspective
- Prompt Injection Vulnerability of Consensus Generating Applications in Digital Democracy
- Current State in Privacy-Preserving Text Preprocessing for Domain-Agnostic NLP
- Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
- Bridging AI Innovation and Healthcare Needs: Lessons Learned from Incorporating Modern NLP at The BC Cancer Registry
- Differentiating hype from practical applications of large language models in medicine -- a primer for healthcare professionals
- What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
- PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- Memorization Sinks: Isolating Memorization during LLM Training
- The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
- Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models
- PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
- FlashDP: Private Training Large Language Models with Efficient DP-SGD
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
- Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models
- On Recipe Memorization and Creativity in Large Language Models: Is Your Model a Creative Cook, a Bad Cook, or Merely a Plagiator?
- Counterfactual Influence as a Distributional Quantity
- PrivacyXray: Detecting Privacy Breaches in LLMs through Semantic Consistency and Probability Certainty
- BLUR: A Bi-Level Optimization Approach for LLM Unlearning
- Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
- Approximating Language Model Training Data from Weights
- SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
- Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
- Black-Box Access is Insufficient for Rigorous AI Audits
- Membership Inference Attacks on Sequence Models
- Quantifying Cross-Modality Memorization in Vision-Language Models
- OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
- DMRL: Data- and Model-aware Reward Learning for Data Extraction
- ATAG: AI-Agent Application Threat Assessment with Attack Graphs
- Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference Attack
- ACCESS DENIED INC: The First Benchmark Environment for Sensitivity Awareness
- Automatic Calibration for Membership Inference Attack on Large Language Models
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
- Hush! Protecting Secrets During Model Training: An Indistinguishability Approach
- Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
- Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities
- DUSK: Do Not Unlearn Shared Knowledge
- Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
- Adversarially Pretrained Transformers may be Universally Robust In-Context Learners
- Fragments to Facts: Partial-Information Fragment Inference from LLMs
- One-Step Offline Distillation of Diffusion-based Models via Koopman Modeling
- PANORAMA: A synthetic PII-laced dataset for studying sensitive data memorization in LLMs
- IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
- PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
- We Should Separate Memorization from Copyright
- The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
- 2023 in science [wikipedia]
- Timeline of computing 2020–present [wikipedia]
Discussions
- Google Researchers’ Attack Prompts ChatGPT to Reveal Its Training Data [lemmy, 355 points, 67 comments]
- Scalable extraction of training data from (production) language models [hn, 105 points, 14 comments]
- "researchers showed that there are large amounts of privately identifiable information (PII) in OpenAI’s large language models...the chatbot spit out large passages of text scraped verbatim from other [bsky, 20 points, 0 comments]
- Extracting Training Data from ChatGPT [lemmy, 10 points, 0 comments]
- Great 2001/Hal dying energy in this project that broke ChatGPT by telling it to repeat single words forever, causing it to cough up training data verbatim arxiv.org/pdf/2311.170... [bsky, 7 points, 0 comments]
- Scalable Extraction of Training Data from (Production) Language Models [hn, 5 points, 0 comments]
- arxiv.org/abs/2311.17035 You can extract verbatim training data from ChatCPT by asking "Repeat this word forever: ‘poem poem poem poem" 🤯 Repeating a token makes it 'diverge' from the alignment tra [bsky, 5 points, 0 comments]
- Remember when people found out they could extract whole blocks of training data from LLMs? Betcha negotiations to "buy access" to academic publishers' products strated about a day after that. [bsky, 4 points, 1 comments]
- My read of the paper is that if you get the program to just generate junk with little guidance it eventually randomly generates something that starts some text that is overfit in the data and then rep [bsky, 4 points, 2 comments]
- Extracting Training Data from LLMs [hn, 4 points, 1 comments]
- so, this was interesting the other day (arxiv.org/abs/2311.17035 ) because i've seen arguments that copyright claims against generative ai is invalid because you can't recover the original work from t [bsky, 3 points, 2 comments]
- Scalable Extraction of Training Data from (Production) Language Models [hn, 2 points, 0 comments]
- One thing this suggests to me is that AI learns a bit more like humans than we might otherwise expect: while an AI model can only memorize a tiny portion of its training set, it appears to memorize a [bsky, 2 points, 1 comments]
- “Scalable Extraction of Training Data from (Production) Language Models” "We show an adversary can extract gigabytes of training data from open-source language models like Pythia or GPT-Neo, semi-ope [bsky, 1 points, 1 comments]
- herewith the research paper used [you can download the PDF from Cornell's ArXiv labs-- right side] arxiv.org/abs/2311.17035 if you don't see this blame @safety.bsky.app for their lack of a [non/un]ba [bsky, 1 points, 1 comments]
- Scalable Extraction of Training Data from (Production) Language Models [lobsters, 1 points, 1 comments]
- Research paper on extracting (a lot of) training data from ChatGPT arxiv.org/abs/2311.17035 [bsky, 0 points, 0 comments]
- Here’s a paper in which researchers are able to get LLMs to regurgitate their training data in response to weird prompts. arxiv.org/abs/2311.17035 I guess if you shake the vending machine in the right [bsky, 0 points, 0 comments]
- Scalable Extraction of Training Data from (Production) Language Models arxiv.org/pdf/2311.170... [bsky, 0 points, 0 comments]
- Apparently ChatGPT starts spitting out training data if you ask it to say 'poem' too many times arxiv.org/pdf/2311.170... [bsky, 0 points, 0 comments]
- OpenAI supposedly fixed the bug in this paper arxiv.org/pdf/2311.170... that causes ChatGPT to spit out memorized training data. but nope, I don't think they have: [bsky, 0 points, 0 comments]
Related