vix.ing · top · new · best · stats · spec

A Recipe for a Good π. How to Properly Estimate Population Genetics Summary Statistics and Why we Should Systematically Report Them

2026/05/04 by Maxence Brault, Thomas Brazier, Alexander Mackintosh +3 · 2 voices · 2 citations
Biochemistry, Genetics and Molecular Biology · Environmental Science · #Genetic diversity and population structure #Environmental DNA in Biodiversity Studies #Isotope Analysis in Ecology

paper · doi:10.1093/gbe/evag103

Abstract

Many long-standing questions in population genomics can now be addressed through comparative analyses and by leveraging the vast amount of genomic data being generated. In the context of questioning the utility of producing such a large amount of genomic data, whether for ecological or economic reasons, we argue that data publication should be standardized to ensure long-term reusability. Based on a literature review and key examples, we emphasize that despite the growing volume of available data, the lack of methodological documentation and the absence of metadata make most published polymorphism datasets incomparable, preventing the calculation of meaningful statistics and the application of FAIR (Findable, Accessible, Interoperable, Reusable) principles. We stress that the Variant Calling Format (VCF) as it is used and published today is insufficient, as it does not report the number of monomorphic sites, which are required to compute basic statistics such as pairwise nucleotide diversity (π) or Watterson's θ. We further propose guidelines and best practices to provide sufficient information to allow the proper calculation of these statistics while accounting for sources of bias and misestimation frequently observed in the literature. Finally, we underscore the need for the systematic reporting of standardized statistics, coupled with transparent documentation of data processing steps, to ensure the reproducibility and comparability of population genomic research.

Citations

Cited by

Discussions

Related