Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
2026/04/08 by Qian Ma, Sarah Rajtmajer · 1 voice
Computer Science · #cs.CR #cs.AI
paper · pdf · doi:10.18653/v1/2026.findings-acl.10
arxiv published 2026/04/08 · arxiv updated 2026/07/13
Abstract
Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy and utility. We propose Realistic and Privacy-Preserving Synthetic Data Generation (RPSG), which uses private seeds and integrates privacy-preserving strategies, including a formal differential privacy (DP) mechanism in the candidate selection, to generate realistic synthetic data. Comprehensive experiments against state-of-the-art private synthetic data generation methods demonstrate that RPSG achieves high fidelity to private data while providing strong privacy protection.
Citations
- Knowledge Distillation Using Frontier Open-source LLMs: Generalizability and the Role of Synthetic Data
- Masked Clinical Modelling: A Framework for Synthetic and Augmented Survival Data Generation
- Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
- MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
- Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data
- LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
- IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization
- On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
- A Synthetic Dataset for Personal Attribute Inference
- PrE-Text: Training Language Models on Private Federated Data in the Age of LLMs
- Privacy Preserving Prompt Engineering: A Survey
- Differentially Private Synthetic Data via Foundation Model APIs 2: Text
- Reducing Privacy Risks in Online Self-Disclosures with Language Models
- DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
- Harnessing large-language models to generate private synthetic text
- Membership Inference Attacks against Language Models via Neighbourhood Comparison
- Fine-Tuning Language Models with Just Forward Passes
- Does Synthetic Data Generation of LLMs Help Clinical Text Mining?
- A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
- Machine Learning for Synthetic Data Generation: A Review
- Synthetic Text Generation with Differential Privacy: A Simple and Practical Recipe
- Differentially Private Language Models for Secure Data Sharing
- More than a Feeling: Accuracy and Application of Sentiment Analysis
- Memorization in NLP Fine-tuning Methods
- Membership Inference Attacks From First Principles
- Differentially Private Fine-tuning of Language Models
- Extracting Training Data from Large Language Models
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language\n Generation, Translation, and Comprehension
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- The Curious Case of Neural Text Degeneration
- Jointly Measuring Diversity and Quality in Text Generation Models
- Scalable Private Learning with PATE
- Texygen: A Benchmarking Platform for Text Generation Models
- Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record
- Membership Inference Attacks against Machine Learning Models
- Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data
- The Algorithmic Foundations of Differential Privacy
Discussions