2020/04/13 by Brady T. West, West, Brady T., Roderick J. A. Little +11
Biochemistry, Genetics and Molecular Biology · Social Sciences · #Applications (stat.AP) #Evolution and Genetic Dynamics #FOS: Computer and information sciences #Genetic Associations and Epidemiology #Media Influence and Politics #Methodology (stat.ME)
paper · pdf · doi:10.48550/arxiv.2004.06139
openalex publication_date 2020/04/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Selection bias is a serious potential problem for inference about\nrelationships of scientific interest based on samples without well-defined\nprobability sampling mechanisms. Motivated by the potential for selection bias\nin (a) estimated relationships of polygenic scores (PGSs) with phenotypes in\ngenetic studies of volunteers, and (b) estimated differences in subgroup means\nin surveys of smartphone users, we derive novel measures of selection bias for\nestimates of the coefficients in linear and probit regression models fitted to\nnon-probability samples, when aggregate-level auxiliary data are available for\nthe selected sample and the target population. The measures arise from normal\npattern-mixture models that allow analysts to examine the sensitivity of their\ninferences to assumptions about non-ignorable selection in these samples. We\nexamine the effectiveness of the proposed measures in a simulation study, and\nthen use them to quantify the selection bias in (a) estimated PGS-phenotype\nrelationships in a large study of volunteers recruited via Facebook, and (b)\nestimated subgroup differences in mean past-year employment duration in a\nnon-probability sample of low-educated smartphone users. We evaluate the\nperformance of the measures in these applications using benchmark estimates\nfrom large probability samples.\n