2009/11/01 by Charles Kooperberg, Michael LeBlanc, James Y. Dai +1 · 26 citations
Biochemistry, Genetics and Molecular Biology · Mathematics · Medicine · #Biology #Computational biology #Computer science #Disease #Endoplasmic Reticulum Stress and Disease #Gene #Genetic Associations and Epidemiology #Genetic association #Genetics #Genome-wide association study #Genotype #Identification (biology) #Machine learning #Medicine #RNA regulation and disease #SNP #Selection (genetic algorithm) #Single-nucleotide polymorphism #q-bio.GN #stat.ME
paper · pdf · doi:10.1214/09-sts287
published in Statistical Science 24(4), 472-488 (Institute of Mathematical Statistics) · Published in at http://dx.doi.org/10.1214/09-STS287 the Statistical Science (http://www.imstat.org/sts/) by the Institute of Mathematical Statistics (http://www.imstat.org)
openalex publication_date 2009/11/01 · arxiv created 2010/10/22 · arxiv updated 2010/10/25 · openalex created_date 2020/11/23 · openalex updated_date 2026/08/06
Genome-wide association studies, in which as many as a million single nucleotide polymorphisms (SNP) are measured on several thousand samples, are quickly becoming a common type of study for identifying genetic factors associated with many phenotypes. There is a strong assumption that interactions between SNPs or genes and interactions between genes and environmental factors substantially contribute to the genetic risk of a disease. Identification of such interactions could potentially lead to increased understanding about disease mechanisms; drug × gene interactions could have profound applications for personalized medicine; strong interaction effects could be beneficial for risk prediction models. In this paper we provide an overview of different approaches to model interactions, emphasizing approaches that make specific use of the structure of genetic data, and those that make specific modeling assumptions that may (or may not) be reasonable to make. We conclude that to identify interactions it is often necessary to do some selection of SNPs, for example, based on prior hypothesis or marginal significance, but that to identify SNPs that are marginally associated with a disease it may also be useful to consider larger numbers of interactions.