2022/12/26 by Travis Canida, Hongjie Ke, Canida, Travis +4
Biochemistry, Genetics and Molecular Biology · #FOS: Computer and information sciences #Gene expression and cancer classification #Genetic Mapping and Diversity in Plants and Animals #Genetic and phenotypic traits in livestock #Methodology (stat.ME)
paper · pdf · doi:10.48550/arxiv.2212.13294
openalex publication_date 2022/12/26 · openalex created_date 2023/01/06 · openalex updated_date 2026/07/28
Variable selection has played a critical role in modern statistical learning and scientific discoveries. Numerous regularization and Bayesian variable selection methods have been developed in the past two decades for variable selection, but most of these methods consider selecting variables for only one response. As more data is being collected nowadays, it is common to analyze multiple related responses from the same study. Existing multivariate variable selection methods select variables for all responses without considering the possible heterogeneity across different responses, i.e. some features may only predict a subset of responses but not the rest. Motivated by the multi-trait fine mapping problem in genetics to identify the causal variants for multiple related traits, we developed a novel multivariate Bayesian variable selection method to select critical predictors from a large number of grouped predictors that target at multiple correlated and possibly heterogeneous responses. Our new method is featured by its selection at multiple levels, its incorporation of prior biological knowledge to guide selection and identification of best subset of responses predictors target at. We showed the advantage of our method via extensive simulations and a real fine mapping example to identify causal variants associated with different subsets of addictive behaviors.