2026/01/20 by William Hemstrom, Mark R. Christie · 1 voice
Biochemistry, Genetics and Molecular Biology · #Genetic Associations and Epidemiology #Genetic diversity and population structure #Forensic and Genetic Research
paper · doi:10.1093/jhered/esag005
openalex publication_date 2026/01/20 · openalex created_date 2026/01/23 · openalex updated_date 2026/08/01
Estimators for the numbers of polymorphic loci or alleles, such as private allele counts or allelic richness, are not as commonly used or well developed for single nucleotide polymorphism (SNP) data as are genetic estimators which rely on direct measurements of allele frequencies (such as expected heterozygosity, nucleotide diversity, and Tajima's θ). The number of segregating sites (S), the number of nucleotide sites that have more than one allele across the genome (e.g. SNPs), is one such estimator which does not rely on estimates of allele frequencies. S can provide informative estimates of genetic diversity across multiple scales, from genes to chromosomes to entire genomes, and is particularly informative when used in conjunction with allele frequency dependent estimators such as expected heterozygosity. However, segregating site counts are rarely adjusted to correct for unequal sample sizes or differences in missing data among populations or sample groups, and when they are, they typically fail to account for deviations from Hardy-Weinberg Proportions (HWP). Here, we introduce an unbiased estimator for the number of segregating sites expected in a sample group following rarefaction (S') and we use simulated data sets to illustrate that S' allows for accurate comparisons of the number of segregating sites among multiple sample groups with varied sample sizes and deviations from HWP. Lastly, by re-analyzing two existing empirical datasets which calculated S we show that S' produces different and less variable rank-order genetic diversity estimates than either S or Watterson's θ, which may have potential impacts on downstream biological inference.