2020/11/01 by Tejas Kulkarni, Joonas Jälkö, Kulkarni, Tejas +7 · 5 citations
Computer Science · Mathematics · #Artificial intelligence #Bayesian inference #Bayesian probability #Bayesian statistics #Computer science #Cryptography and Security (cs.CR) #Econometrics #FOS: Computer and information sciences #Generalized linear model #Inference #Linear model #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Mathematics #Privacy-Preserving Technologies in Data #Random Matrices and Applications #Statistical Methods and Bayesian Inference #cs.CR #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2011.00467
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2020/11/01 · arxiv created 2021/05/12 · arxiv updated 2021/05/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Generalized linear models (GLMs) such as logistic regression are among the most widely used arms in data analyst's repertoire and often used on sensitive datasets. A large body of prior works that investigate GLMs under differential privacy (DP) constraints provide only private point estimates of the regression coefficients, and are not able to quantify parameter uncertainty. In this work, with logistic and Poisson regression as running examples, we introduce a generic noise-aware DP Bayesian inference method for a GLM at hand, given a noisy sum of summary statistics. Quantifying uncertainty allows us to determine which of the regression coefficients are statistically significantly different from zero. We provide a previously unknown tight privacy analysis and experimentally demonstrate that the posteriors obtained from our model, while adhering to strong privacy guarantees, are close to the non-private posteriors.