A Survey on Privacy Risks and Protection in Large Language Models
2025/05/04 by Chen, Kang, Zhou, Xiuze, Lin, Yuanguo +3 · 6 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2505.01976
Abstract
Although Large Language Models (LLMs) have become increasingly integral to diverse applications, their capabilities raise significant privacy concerns. This survey offers a comprehensive overview of privacy risks associated with LLMs and examines current solutions to mitigate these challenges. First, we analyze privacy leakage and attacks in LLMs, focusing on how these models unintentionally expose sensitive information through techniques such as model inversion, training data extraction, and membership inference. We investigate the mechanisms of privacy leakage, including the unauthorized extraction of training data and the potential exploitation of these vulnerabilities by malicious actors. Next, we review existing privacy protection against such risks, such as inference detection, federated learning, backdoor mitigation, and confidential computing, and assess their effectiveness in preventing privacy leakage. Furthermore, we highlight key practical challenges and propose future research directions to develop secure and privacy-preserving LLMs, emphasizing privacy risk assessment, secure knowledge transfer between models, and interdisciplinary frameworks for privacy governance. Ultimately, this survey aims to establish a roadmap for addressing escalating privacy challenges in the LLMs domain.
Citations
- PII-Bench: Evaluating Query-Aware Privacy Protection Systems
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Model Inversion in Split Learning for Personalized LLMs: New Insights from Information Bottleneck Theory
- SeSeMI: Secure Serverless Model Inference on Sensitive Data
- Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
- LLM-PBE: Assessing Data Privacy in Large Language Models
- Enhancing Data Privacy in Large Language Models through Private Association Editing
- GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
- Knowledge Distillation in Federated Learning: a Survey on Long Lasting Challenges and New Solutions
- Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
- A Comparative Analysis of Word-Level Metric Differential Privacy: Benchmarking The Privacy-Utility Trade-off
- Differentially Private Next-Token Prediction of Large Language Models
- BadEdit: Backdooring large language models by model editing
- On Protecting the Data Privacy of Large Language Models (LLMs): A Survey
- LLMGuard: Guarding Against Unsafe LLM Behavior
- Prompt Stealing Attacks Against Large Language Models
- Privacy-Preserving Language Model Inference with Instance Obfuscation
- Do Membership Inference Attacks Work on Large Language Models?
- Security and Privacy Challenges of Large Language Models: A Survey
- Security and Privacy Challenges of Large Language Models: A Survey
- TrustLLM: Trustworthiness in Large Language Models
- SecFormer: Fast and Accurate Privacy-Preserving Inference for Transformer Models via SMPC
- A Comprehensive Survey of Attack Techniques, Implementation, and Mitigation Strategies in Large Language Models
- Privacy Issues in Large Language Models: A Survey
- PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models
- PriPrune: Quantifying and Preserving Privacy in Pruned Federated Learning
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
- Privacy Preserving Large Language Models: ChatGPT Case Study Based Vision and Framework
- Beyond Memorization: Violating Privacy Via Inference with Large Language Models
- PrivacyMind: Large Language Models Can Be Contextual Privacy Protection Learners
- FedBPT: Efficient Federated Black-box Prompt Tuning for Large Language Models
- LinGCN: Structural Linearized Graph Convolutional Network for Homomorphically Encrypted Inference
- Large language models can accurately predict searcher preferences
- A Survey on Model Compression for Large Language Models
- ProPILE: Probing Privacy Leakage in Large Language Models
- Flocks of Stochastic Parrots: Differentially Private Prompt Learning for Large Language Models
- Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility
- ChatGPT Evaluation on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations
- Enhancing Fine-Tuning Based Backdoor Defense with Sharpness-Aware Minimization
- ChatGPT for good? On opportunities and challenges of large language models for education
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- Fine-Tuning Is All You Need to Mitigate Backdoor Attacks
- Holistic risk assessment of inference attacks in machine learning
- Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers
- FedPrompt: Communication-Efficient and Privacy Preserving Prompt Tuning in Federated Learning
- THE-X: Privacy-Preserving Transformer Inference with Homomorphic Encryption
- You Are What You Write: Preserving Privacy in the Era of Large Language Models
- Training language models to follow instructions with human feedback
- Towards a Data Privacy-Predictive Performance Trade-off
- Differentially Private Fine-tuning of Language Models
- Backdoor Attacks on Pre-trained Models by Layerwise Weight Poisoning
- On the Opportunities and Risks of Foundation Models
- Killing One Bird with Two Stones: Model Extraction and Attribute Inference Attacks against BERT-based APIs
- On the (In)Feasibility of Attribute Inference Attacks on Machine Learning Models
- Extracting Training Data from Large Language Models
- Data-Free Model Extraction
- Weight Poisoning Attacks on Pre-trained Models
- Advances and Open Problems in Federated Learning
- The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural Networks
- Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks
- A Survey on Homomorphic Encryption Schemes: Theory and Implementation
- Membership Inference Attacks against Machine Learning Models
- Unique Security and Privacy Threats of Large Language Models: A Comprehensive Survey
- InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
- Integration of Large Language Models and Federated Learning
Cited by
Related