2010/05/03 by Krista J. Gile, Mark S. Handcock · 1 citation
Health Professions · Mathematics · Medicine · #Computer science #Demography #Econometrics #Estimator #Fraction (chemistry) #HIV, Drug Use, Sexual Risk #HIV/AIDS Research and Interventions #Homelessness and Social Issues #Mathematics #Population #Respondent #Sample (material) #Sample size determination #Sampling (signal processing) #Sampling bias #Statistics #Survey sampling #Telecommunications
paper · doi:10.1111/j.1467-9531.2010.01223.x
crossref issued 2010/05/03 · crossref published 2010/05/03 · crossref published-online 2010/05/03 · openalex publication_date 2010/05/03 · crossref created 2010/05/04 · crossref published-print 2010/08/01 · openalex created_date 2025/10/10 · crossref deposited 2026/05/01 · crossref indexed 2026/07/31 · openalex updated_date 2026/08/01
Respondent-Driven Sampling (RDS) employs a variant of a link-tracing network sampling strategy to collect data from hard-to-reach populations. By tracing the links in the underlying social network, the process exploits the social structure to expand the sample and reduce its dependence on the initial (convenience) sample.The current estimators of population averages make strong assumptions in order to treat the data as a probability sample. We evaluate three critical sensitivities of the estimators: to bias induced by the initial sample, to uncontrollable features of respondent behavior, and to the without-replacement structure of sampling.Our analysis indicates: (1) that the convenience sample of seeds can induce bias, and the number of sample waves typically used in RDS is likely insufficient for the type of nodal mixing required to obtain the reputed asymptotic unbiasedness; (2) that preferential referral behavior by respondents leads to bias; (3) that when a substantial fraction of the target population is sampled the current estimators can have substantial bias.This paper sounds a cautionary note for the users of RDS. While current RDS methodology is powerful and clever, the favorable statistical properties claimed for the current estimates are shown to be heavily dependent on often unrealistic assumptions. We recommend ways to improve the methodology.