2007/09/04 by Elizabeth S. Allman, Cecile Ane, Allman, Elizabeth S. +4 · 1 citation
Biochemistry, Genetics and Molecular Biology · Mathematics · #62P10 #92D15 #Evolution and Genetic Dynamics #FOS: Biological sciences #FOS: Mathematics #Populations and Evolution (q-bio.PE) #Statistics Theory (math.ST) #math.ST #msc:62P10 #msc:92D15 #q-bio.PE #stat.TH
paper · pdf · doi:10.48550/arxiv.0709.0531
35 pages, 3 figures; Minor revisions and reformatting to reflect version to be published
openalex publication_date 2007/09/04 · arxiv created 2008/02/01 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Inference of evolutionary trees and rates from biological sequences is commonly performed using continuous-time Markov models of character change. The Markov process evolves along an unknown tree while observations arise only from the tips of the tree. Rate heterogeneity is present in most real data sets and is accounted for by the use of flexible mixture models where each site is allowed its own rate. Very little has been rigorously established concerning the identifiability of the models currently in common use in data analysis, although non-identifiability was proven for a semi-parametric model and an incorrect proof of identifiability was published for a general parametric model (GTR+Gamma+I). Here we prove that one of the most widely used models (GTR+Gamma) is identifiable for generic parameters, and for all parameter choices in the case of 4-state (DNA) models. This is the first proof of identifiability of a phylogenetic model with a continuous distribution of rates.