2021/06/28 by Martin Jankowiak, Jankowiak, Martin
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Bayesian Methods and Mixture Models #Computation (stat.CO) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Metabolomics and Mass Spectrometry Studies #Methodology (stat.ME) #Statistical Methods and Bayesian Inference #stat.CO #stat.ME #stat.ML
paper · pdf · doi:10.48550/arxiv.2106.14981
18 pages; this work is superseded by arXiv:2208.01180
openalex publication_date 2021/06/28 · arxiv created 2022/09/12 · arxiv updated 2022/09/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Bayesian variable selection is a powerful tool for data analysis, as it offers a principled method for variable selection that accounts for prior information and uncertainty. However, wider adoption of Bayesian variable selection has been hampered by computational challenges, especially in difficult regimes with a large number of covariates or non-conjugate likelihoods. Generalized linear models for count data, which are prevalent in biology, ecology, economics, and beyond, represent an important special case. Here we introduce an efficient MCMC scheme for variable selection in binomial and negative binomial regression that exploits Tempered Gibbs Sampling (Zanella and Roberts, 2019) and that includes logistic regression as a special case. In experiments we demonstrate the effectiveness of our approach, including on cancer data with seventeen thousand covariates.