2025/10/02 by Oussama Ounissi, Ounissi, Oussama, Nicklas Jävergård +4
Computer Science · #15-03 #15A18 #47A55 #Computational Physics and Python Applications #FOS: Computer and information sciences #FOS: Mathematics #Generative Adversarial Networks and Image Synthesis #Machine Learning (stat.ML) #Methodology (stat.ME) #Statistics Theory (math.ST) #Time Series Analysis and Forecasting
paper · pdf · doi:10.48550/arxiv.2510.02405
openalex publication_date 2025/10/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Synthetic data generation is increasingly used in applications involving privacy preservation, data sharing, and data scarcity. In many situations, preserving the dependence structure of the original data is of central interest. In this work, we propose a lightweight postprocessing methodology for synthetic tabular data based on the Orthogonal Procrustes problem. Starting from an already generated synthetic dataset, our approach constructs the closest dataset that restores the Pearson correlation structure of the original data. On the theoretical side, we show that preserving Pearson correlation is equivalent to the action of linear orthogonal maps in the centered-data subspace, and then deploy the Orthogonal Procrustes problem. However, in order for this to hold, we first establish a result ensuring that applying the Orthogonal Procrustes step remains in the aforementioned subspace under suitable assumptions. Applications to several datasets and synthetic data generators illustrate the effectiveness of the proposed approach. In particular, the numerical experiments indicate that the correlation structure can be restored while largely preserving the individual feature distributions, the geometry of the data, and the performance of downstream classification tasks.