2017/07/07 by Stephen Burgess, Burgess, Stephen, Verena Zuber +7 · 1 citation
Biochemistry, Genetics and Molecular Biology · #FOS: Computer and information sciences #Genetic Associations and Epidemiology #Genetic Mapping and Diversity in Plants and Animals #Genetic and phenotypic traits in livestock #Methodology (stat.ME)
paper · pdf · doi:10.48550/arxiv.1707.02215
openalex publication_date 2017/07/07 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
Mendelian randomization uses genetic variants to make causal inferences about\nthe effect of a risk factor on an outcome. With fine-mapped genetic data, there\nmay be hundreds of genetic variants in a single gene region any of which could\nbe used to assess this causal relationship. However, using too many genetic\nvariants in the analysis can lead to spurious estimates and inflated Type 1\nerror rates. But if only a few genetic variants are used, then the majority of\nthe data is ignored and estimates are highly sensitive to the particular choice\nof variants. We propose an approach based on summarized data only (genetic\nassociation and correlation estimates) that uses principal components analysis\nto form instruments. This approach has desirable theoretical properties: it\ntakes the totality of data into account and does not suffer from numerical\ninstabilities. It also has good properties in simulation studies: it is not\nparticularly sensitive to varying the genetic variants included in the analysis\nor the genetic correlation matrix, and it does not have greatly inflated Type 1\nerror rates. Overall, the method gives estimates that are not so precise as\nthose from variable selection approaches (such as using a conditional analysis\nor pruning approach to select variants), but are more robust to seemingly\narbitrary choices in the variable selection step. Methods are illustrated by an\nexample using genetic associations with testosterone for 320 genetic variants\nto assess the effect of sex hormone-related pathways on coronary artery disease\nrisk, in which variable selection approaches give inconsistent inferences.\n