2002/07/05 by A. Eriksson, Anders Eriksson, B. Haubold +5
Agricultural and Biological Sciences · Biochemistry, Genetics and Molecular Biology · Physics and Astronomy · #Biological Physics (physics.bio-ph) #Chromosomal and Genetic Variations #FOS: Biological sciences #FOS: Physical sciences #Genomics and Phylogenetic Studies #Quantitative Biology (q-bio) #RNA and protein synthesis mechanisms #physics.bio-ph #q-bio
paper · pdf · doi:10.48550/arxiv.physics/0207024
17 pages, 5 figures
arxiv created 2002/07/05 · openalex publication_date 2002/07/05 · arxiv updated 2016/09/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Single nucleotide polymorphisms (SNPs) often appear in clusters along the length of a chromosome. This is due to variation in local coalescent times caused by,for example, selection or recombination. Here we investigate whether recombination alone (within a neutral model) can cause statistically significant SNP clustering. We measure the extent of SNP clustering as the ratio between the variance of SNPs found in bins of length l, and the mean number of SNPs in such bins, σ2l/μl. For a uniform SNP distribution σ2l/μl=1, for clustered SNPs σ2l/μl > 1. Apart from the bin length, three length scales are important when accounting for SNP clustering: The mean distance between neighboring SNPs, Δ, the mean length of chromosome segments with constant time to the most recent common ancestor, \el, and the total length of the chromosome, L. We show that SNP clustering is observed if Δ< \el ≪ L. Moreover, if l≪ \el ≪ L, clustering becomes independent of the rate of recombination. We apply our results to the analysis of SNP data sets from mice, and human chromosomes 6 and X. Of the three data sets investigated, the human X chromosome displays the most significant deviation from neutrality.