vix.ing · top · new · best · stats · spec

Differentially Private Low-dimensional Synthetic Data from High-dimensional Datasets

2023/05/26 by Yiyun He, He, Yiyun, Thomas Strohmer +5
Computer Science · #Chaos-based Image/Signal Encryption #Cryptography and Data Security #Cryptography and Security (cs.CR) #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data #Probability (math.PR) #Statistics Theory (math.ST)

paper · pdf · doi:10.48550/arxiv.2305.17148

openalex publication_date 2023/05/26 · openalex created_date 2023/05/31 · openalex updated_date 2026/07/28

Abstract

Differentially private synthetic data provide a powerful mechanism to enable data analysis while protecting sensitive information about individuals. However, when the data lie in a high-dimensional space, the accuracy of the synthetic data suffers from the curse of dimensionality. In this paper, we propose a differentially private algorithm to generate low-dimensional synthetic data efficiently from a high-dimensional dataset with a utility guarantee with respect to the Wasserstein distance. A key step of our algorithm is a private principal component analysis (PCA) procedure with a near-optimal accuracy bound that circumvents the curse of dimensionality. Unlike the standard perturbation analysis, our analysis of private PCA works without assuming the spectral gap for the covariance matrix.

Related