vix.ing · top · new · best · stats · spec

Regression with missing Ys: An improved strategy for analyzing multiply imputed data

2007/05/18 by Paul T. von Hippel · 6 citations
Mathematics · #Advanced Statistical Methods and Models #Statistical Methods and Bayesian Inference #Statistical Methods and Inference #stat.ME

paper · pdf · doi:10.1111/j.1467-9531.2007.00180.x

published as Sociological Methodology (2007) volume 37, pp. 83-117

openalex publication_date 2007/05/18 · crossref created 2007/05/18 · crossref issued 2007/08/01 · crossref published 2007/08/01 · crossref published-online 2007/08/01 · crossref published-print 2007/08/01 · arxiv created 2016/05/03 · arxiv updated 2017/03/27 · openalex created_date 2025/10/10 · crossref deposited 2026/05/01 · openalex updated_date 2026/08/04 · crossref indexed 2026/08/04

Abstract

When fitting a generalized linear model -- such as a linear regression, a logistic regression, or a hierarchical linear model -- analysts often wonder how to handle missing values of the dependent variable Y. If missing values have been filled in using multiple imputation, the usual advice is to use the imputed Y values in analysis. We show, however, that using imputed Ys can add needless noise to the estimates. Better estimates can usually be obtained using a modified strategy that we call multiple imputation, then deletion (MID). Under MID, all cases are used for imputation, but following imputation cases with imputed Y values are excluded from the analysis. When there is something wrong with the imputed Y values, MID protects the estimates from the problematic imputations. And when the imputed Y values are acceptable, MID usually offers somewhat more efficient estimates than an ordinary MI strategy.

Citations

Cited by

Related