2026/03/04 by Hakime Öztürk, Tejumade Afonja, Joonas Jälkö +19 · 1 voice
Biochemistry, Genetics and Molecular Biology · Computer Science · Medicine · #Single-cell and spatial transcriptomics #Privacy-Preserving Technologies in Data #Ferroptosis and cancer prognosis
paper · doi:10.64898/2026.03.02.707794
openalex publication_date 2026/03/04 · openalex created_date 2026/03/05 · openalex updated_date 2026/07/22
Abstract Background The synthesis of anonymized data derived from real-world cohorts offers a promising strategy for regulatory-compliant and privacy-preserving biological data sharing, potentially facilitating model development that can improve predictive performance. However, the extent to which generative models can preserve biological signals while remaining resilient to adversarial privacy attacks in high-dimensional omics contexts remains underexplored. To address this gap, the CAMDA 2025 Health Privacy Challenge launched a community-driven effort to systematically benchmark synthetic and privacy-preserving data generation for bulk RNA-seq cohorts. Results Building on this initiative, we systematically benchmarked 11 generative methods across two cancer cohorts (∼1,000 and ∼5,000 patients) over 978 landmark genes. Methods were evaluated across complementary axes of distributional fidelity, downstream utility, biological plausibility and empirical privacy risk, with emphasis on trade-offs between vulnerability to membership inference attacks (MIA) and other evaluation dimensions. Expressive deep generative models achieved strong predictive utility and differential expression recovery, but were often more vulnerable to membership inference risk. Differentially private methods improved resistance to attacks at the cost of reduced utility, while simpler statistical approaches offered competitive utility with moderate privacy risk and fast training. Conclusions Synthetic bulk RNA-seq quality is inherently multi-dimensional and shaped by trade-offs between utility, biological preservation and privacy. Our results indicate that differences in model architecture drive distinct trade-offs across these axes, suggesting that model choice should align with dataset characteristics, intended downstream use and privacy requirements. Privacy risk should also be assessed using multiple complementary attack methods and, where possible, formal differential privacy protection.