2020/01/13 by Lee S. McDaniel, McDaniel, Lee S., Jonathan S. Schildcrout +5
Decision Sciences · Mathematics · #FOS: Computer and information sciences #Methodology (stat.ME) #Optimal Experimental Design Methods #Statistical Methods and Bayesian Inference #Statistical Methods in Clinical Trials #stat.ME
paper · pdf · doi:10.48550/arxiv.2001.04444
arxiv created 2020/01/13 · openalex publication_date 2020/01/13 · arxiv updated 2020/01/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Biased sampling designs can be highly efficient when studying rare (binary) or low variability (continuous) endpoints. We consider longitudinal data settings in which the probability of being sampled depends on a repeatedly measured response through an outcome-related, auxiliary variable. Such auxiliary variable- or outcome-dependent sampling improves observed response and possibly exposure variability over random sampling, even though the auxiliary variable is not of scientific interest. For analysis, we propose a generalized linear model based approach using a sequence of two offsetted regressions. The first estimates the relationship of the auxiliary variable to response and covariate data using an offsetted logistic regression model. The offset hinges on the (assumed) known ratio of sampling probabilities for different values of the auxiliary variable. Results from the auxiliary model are used to estimate observation-specific probabilities of being sampled conditional on the response and covariates, and these probabilities are then used to account for bias in the second, target population model. We provide asymptotic standard errors accounting for uncertainty in the estimation of the auxiliary model, and perform simulation studies demonstrating substantial bias reduction, correct coverage probability, and improved design efficiency over simple random sampling designs. We illustrate the approaches with two examples.