vix.ing · top · new · best · stats · spec

The fallacy of single imputation for trait databases: Use multiple imputation instead

2025/02/18 by Simon P. Blomberg, Orlin S. Todorov · 2 voices · 3 citations
Biochemistry, Genetics and Molecular Biology · Mathematics · #Genetic and phenotypic traits in livestock #Statistical Methods and Bayesian Inference #Genetic Mapping and Diversity in Plants and Animals

paper · pdf · doi:10.1111/2041-210x.14494

Abstract

Abstract The past few years have seen the publication of many new trait databases. A common problem with large databases is a lack of completeness, or inversely, the high prevalence of missing values. Biologists have developed several methods to impute (fill in) missing values. This allows ordinary statistical procedures to be used in analyses and the use of only complete cases, with a concomitant loss of power and accuracy, can be avoided. Often, biologists use simulation to test new methods by deleting values from a dataset and recording how well the imputed values match the known, removed values. Here we argue that this is a poor measure of the strength of an imputation method. We also describe the importance and logic of the statistical procedure of multiple imputation, which requires that the imputations need not be precise or accurate estimates of the missing data.

Citations

Cited by

Discussions

Related