2023/12/11 by Lauren Kennedy, Aki Vehtari, Kennedy, Lauren +3
Mathematics · #FOS: Computer and information sciences #Methodology (stat.ME) #Survey Sampling and Estimation Techniques
paper · pdf · doi:10.48550/arxiv.2312.06334
openalex publication_date 2023/12/11 · openalex created_date 2023/12/14 · openalex updated_date 2026/07/28
Generalization to new samples is a fundamental rationale for statistical modeling. For this purpose, model validation is particularly important, but recent work in survey inference has suggested that simple aggregation of individual prediction scores does not give a good measure of the score for population aggregate estimates. In this manuscript we explain why this occurs, propose two scoring metrics designed specifically for this problem, and demonstrate their use in three different ways. We show that these scoring metrics correctly order models when compared to the true score, although they do underestimate the magnitude of the score. We demonstrate with a problem in survey research, where multilevel regression and poststratification (MRP) has been used extensively to adjust convenience and low-response surveys to make population and subpopulation estimates.