2009/11/01 by Nilanjan Chatterjee, Yi‐Hau Chen, Yi-Hau Chen +2 · 24 citations
Biochemistry, Genetics and Molecular Biology · Mathematics · Medicine · #Ambiguity #Bioinformatics and Genomic Networks #Biology #Computational biology #Computer science #Gene #Genetic Associations and Epidemiology #Genetic Mapping and Diversity in Plants and Animals #Genetic association #Genetics #Genotype #Haplotype #Imputation (statistics) #Logistic regression #Mathematics #Medicine #Missing data #Population #Single-nucleotide polymorphism #Statistics #stat.ME
paper · pdf · doi:10.1214/09-sts297
published in Statistical Science 24(4), 489-502 (Institute of Mathematical Statistics) · Published in at http://dx.doi.org/10.1214/09-STS297 the Statistical Science (http://www.imstat.org/sts/) by the Institute of Mathematical Statistics (http://www.imstat.org)
openalex publication_date 2009/11/01 · arxiv created 2010/10/22 · arxiv updated 2010/10/25 · openalex created_date 2020/11/23 · openalex updated_date 2026/08/06
Although prospective logistic regression is the standard method of analysis for case-control data, it has been recently noted that in genetic epidemiologic studies one can use the "retrospective" likelihood to gain major power by incorporating various population genetics model assumptions such as Hardy-Weinberg-Equilibrium (HWE), gene-gene and gene-environment independence. In this article, we review these modern methods and contrast them with the more classical approaches through two types of applications (i) association tests for typed and untyped single nucleotide polymorphisms (SNPs) and (ii) estimation of haplotype effects and haplotype-environment interactions in the presence of haplotype-phase ambiguity. We provide novel insights to existing methods by construction of various score-tests and pseudo-likelihoods. In addition, we describe a novel two-stage method for analysis of untyped SNPs that can use any flexible external algorithm for genotype imputation followed by a powerful association test based on the retrospective likelihood. We illustrate applications of the methods using simulated and real data.