2025/07/30 by Caroline Puente-Lelièvre, Ashar J. Malik, Jordan Douglas · 1 voice · 17 citations
Biochemistry, Genetics and Molecular Biology · #Biology #Computational biology #Computer science #Data science #Evolutionary biology #Gene #Genetics #Genomics and Phylogenetic Studies #Machine Learning in Bioinformatics #Phylogenetic tree #Phylogenetics #Protein Structure and Dynamics #Sequence (biology)
paper · doi:10.1093/gbe/evaf139
published in Genome Biology and Evolution 17(8) (Oxford University Press)
openalex publication_date 2025/07/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Protein structural phylogenetics is an interdisciplinary branch of molecular evolution that (i) uses 3D structural data to trace evolutionary histories, and (ii) uses these evolutionary relationships to explore the diversity of protein structures and their ancestral functions. The appeal in extracting phylogenetic information from protein structure lies in the greater conservation of protein structure compared with sequence, reflecting its resilience to mutation over long evolutionary timescales. Leveraging this information is particularly useful for examining relationships within the "twilight zone"-a region of low protein sequence similarity where it becomes challenging to resolve noise from signal. Historically, the field has been constrained by the limited availability of high-resolution structural data. However, recent breakthroughs in artificial intelligence have made high-quality protein structural data widely accessible. Although the methods for constructing phylogenetic trees from protein structures have progressed significantly from distance-based approaches used since the 1970s, this area of research still lags behind the advanced probabilistic models employed in sequence-based phylogenetics; particularly Bayesian and maximum likelihood approaches. This article reviews the current state of protein structural phylogenetics, outlines methods for extracting evolutionary insights from structural data, and highlights key applications and future directions. Due to the surge of newly available structural information, it is anticipated that sequence and structural data will become routinely integrated in phylogenetic analysis; poising us to venture further into the twilight zone and form cross-disciplinary and translational collaborations.