vix.ing · top · new · best · stats · spec

Advancing plant phylogenetics and evolution through genomic data bursts and methodological innovations

2025/07/01 by Da-Yu Wu, Kangshan Mao · 1 voice
Agricultural and Biological Sciences · #Plant Diversity and Evolution #Plant and animal studies #Plant Taxonomy and Phylogenetics

paper · doi:10.1111/jse.70005

openalex publication_date 2025/07/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Understanding the origins and evolutionary trajectories of biodiversity, along with their spatiotemporal patterns, is a long-standing theme in biology (Kapli et al., 2020). Phylogenetics, reconstructing evolutionary relationships among organisms based on molecular, morphological, and other biological data, acts as a cornerstone of biodiversity studies (Hill et al., 2025). The last two decades have witnessed significant advances in phylogenetics and have kept revolutionizing our understanding of organism evolution (Guo et al., 2023). One of the most notable trends is that molecular data available for phylogenetic and evolutionary analyses experienced a rapid growth (Sayers et al., 2025). Over the past two decades, public sequence repositories have undergone truly transformative growth. For example, the NCBI GenBank and WGS database has expanded from roughly 96 billion base pairs in June 2005 to nearly 44 trillion base pairs by June 2025—representing a more than 450-fold increase. Due to advancements in high-throughput sequencing technologies, the DNA sequencing cost has dropped significantly; the most cost-efficient platform may generate 1 Mb of data for less than 0.002 USD, while the fastest DNA sequencer, BGI DNBSEQ-T1+, can generate more than 100 Gb of data in 2 hours, as of June 2025. Another notable trend is the abundant types of molecular markers available for plant molecular phylogenetic and evolutionary analyses, which were developed based on high-throughput sequencing technology (Guo et al., 2023). Early plant phylogenetics relied on a handful of plastid loci (e.g., rbcL, matK, trnL-F, and trnH-psbA) and the nuclear ITS region, yielding alignments of a few kilobases and dozens to hundreds of informative sites. In the 2000s, single/low-copy nuclear genes such as LEAFY and NEEDLY augmented resolution at deeper nodes (Sayou et al., 2014). Today, genome-scale approaches—from reduced-representation methods (RAD-seq, ddRAD, SLAF-seq) and targeted Hyb-Seq probe sets (e.g., Angiosperm353, now sampled across >7900 genera) to transcriptome sequencing (RNA-seq) and whole-genome resequencing—offer tens to hundreds of thousands of loci per sample (Zuntini et al., 2024). In addition to nuclear genomes, there is also a mushroom growth of plastid genomes and, most recently, mitochondrial genomes, which are available for plant phylogenomic studies. A search of “plastid complete genome” in NCBI GenBank Nucleotide returned 16399 accessions (June 7, 2025), and a search of “plastid phylogenomics” in NCBI PubMed returned 5610 results (June 7, 2025). During the multiple-locus era, methods like Maximum Likelihood (ML), Maximum Parsimony (MP), and Bayesian Inference (BI) were commonly used for phylogenetic analyses (Yang and Rannala, 2012). With the surge in genomic data, two major strategies for species tree inference have emerged: concatenation-based (supermatrix) and coalescent-based methods (Kapli et al., 2020). The concatenation method combines all gene data into a single supermatrix, but it assumes a shared evolutionary history across genes, which can lead to misinterpretations due to gene-specific evolutionary paths. In contrast, coalescent-based methods build individual gene trees and then combine them into a species tree, reducing errors from incomplete lineage sorting and providing a more accurate reflection of evolutionary history (Edwards et al., 2016). Additionally, phylogenetic network methods can capture more complex histories, including hybridization and gene flow, which are critical when reticulate evolution is involved (Than et al., 2008; Solís-Lemus et al., 2017). Rapid growth of molecular data and methodological innovations in the postgenomic era have propelled the fast development of plant phylogenetics and evolution studies, as demonstrated in this issue of the Journal of Systematics and Evolution. Nine case studies and one methodological study have focused on: (i) plastid phylogenomics and biogeographic history, (ii) nuclear phylogenomics and biogeographic history, (iii) data-mining of published DNA sequences, (iv) evolution of innovative traits, and (v) adaptive evolution at the population level. We also noted a handful of growth points in the field of phylogenetics and evolution. Plastid genomes (plastomes) have long been widely used in plant phylogenetics due to their uniparental inheritance, low mutation rates, and relatively conserved structure across species (Jansen and Ruhlman, 2012). Plastomes typically feature a quadripartite structure consisting of two inverted repeat (IR) regions, a large single-copy region, and a small single-copy region (Zhu et al., 2016). With a size range of 120–160 kb across species and a low recombination rate, plastid genomes are ideal for resolving evolutionary relationships within plant lineages (Shaw et al., 2014). Historically, plastome was the first molecular data source used in plant systematics, with early researchers focusing on short plastid DNA regions, such as rbcL, matK, and trnL-F (Shaw et al., 2005). With the advent of next-generation sequencing technologies, the sequencing of whole plastid genomes has become increasingly feasible, offering much higher resolution and accuracy for phylogenetic reconstructions (Li et al., 2019). The use of complete plastomes allows researchers to overcome the limit of phylogenetically informative sites, enabling more robust reconstructions of both deep evolutionary divergences and more recent radiations (Gitzendanner et al., 2018). Currently, the plastome remains an important source of genetic data when inferring phylogenetic relationships. Oulo et al. (2025) conducted a phylogenomic and biogeographic survey of Trichoneura (Poaceae), linking its evolutionary history to the formation of the African arid corridor around 5.78 million years ago. In the aid of a robust phylogeny based on plastomes, the study offers insights into how climatic shifts in the past shape the distribution of plant taxa. Yan et al. (2025) analyzed 67 plastomes across four subfamilies of the Primulaceae, a robust phylogeny was reconstructed, yet conflicts between nuclear and plastid data, particularly regarding Stimpsonia were uncovered. Further analyses indicated that rapid diversification events occurred during the Eocene and Oligocene epochs, which temporally coincide with major geological and climatic changes. The above studies suggest that a robust plastome phylogeny lays the foundation for subsequent inferences of diversification dynamics, biogeographic history, as well as cytonuclear discordance and taxonomic revision in plants. In contrast to the uniparental inheritance and structural conservatism of plastid genomes, genome-wide nuclear genes offer biparental inheritance, higher evolutionary rates, and a broader range of loci, providing a finer-scale signal for phylogenetic inference (Zeng et al., 2014; Zuntini et al., 2024). With advances in high-throughput sequencing and decreasing costs, acquiring large numbers of single-copy nuclear genes (SCGs) has become increasingly feasible, transforming phylogenetic research studies (Guo et al., 2023). Genome-scale data, with their extensive informative sites, allow for the integration of evolutionary histories across most genes, facilitating the construction of more accurate species trees (McKain et al., 2018; Burki et al., 2020). These approaches have been successfully applied across both plant and animal lineages (Li et al., 2022; Jurdzinski et al., 2023; Zhang et al., 2023; Zuntini et al., 2024; Ma et al., 2025). Moreover, phylogenomics enables the exploration of complex evolutionary processes like hybridization, introgression, and whole-genome duplication (Guo et al., 2020; Wang et al., 2022; Bessa et al., 2024), offering deeper insights into evolutionary histories. However, the influx of genomic data has also highlighted the issue of phylogenetic conflict, arising from gene tree estimation errors, incomplete lineage sorting (ILS), and gene flow (Cai et al., 2021; Steenwyk et al., 2023). By leveraging genome-scale data, phylogenomics not only reconstructs species relationships but also dissects these conflicts, providing a clearer understanding of evolutionary trajectories (Zhang et al., 2024). In this issue, Wu et al. (2025) obtained 1991 SCGs employing the target-capture technique and reconstructed a phylogeny of Cupressus, revealing reticulate evolution of this gymnosperm genus. They further quantified the sources of phylogenetic conflict, attributing 61.6% to gene tree estimation error, 14.9% to gene flow, and 7.5% to ILS. Similarly, Xia et al. (2025) reconstructed a phylogeny of Dennstaedtiaceae using 859 orthologous genes derived from transcriptome sequencing data, revealing at least one round of whole genome duplication that is shared by all Dennstaedtiaceae species, as well as complex reticulate evolution within Hypolepidoideae. Their findings confirmed that both ILS and introgression played significant but different roles in shaping the evolutionary history of this fern family. The above studies also discussed cytonuclear discordance. While the plastid genome is often treated as a single locus due to its low recombination frequency compared to nuclear loci, the complex interactions between plastid and nuclear genomes should also be considered. Thus, in addition to ILS and introgression, these complexities should be thoroughly evaluated when interpreting cytonuclear discordance (Larson et al., 2024). The exponential growth of molecular sequence databases, particularly the expansion of NCBI GenBank, has fundamentally transformed phylogenetic research studies by enabling large-scale, data-rich analyses of evolutionary relationships (Li et al., 2019). This rapid accumulation of genomic data facilitates denser and more representative taxon sampling, allowing researchers to access and compile molecular sequences across a broader array of lineages (Smith and Brown, 2018; Sayers et al., 2025). Tools such as PyNCBIminer (Cheng et al., 2025) have been developed to efficiently harness the explosion of molecular sequence data by automating the retrieval and construction of supermatrices directly from NCBI GenBank. PyNCBIminer employs an iterative BLAST algorithm and a hit-extension strategy to improve sequence completeness and recovery rates while selecting representative sequences based on consensus scores. By eliminating the need for local BLAST installation and database setup, and by integrating customizable filtering, trimming, and assembly steps, PyNCBIminer significantly lowers the technical barrier for constructing large-scale phylogenetic datasets, making it accessible to researchers across diverse backgrounds (Cheng et al., 2025). The availability of large-scale molecular datasets provides a robust framework for exploring the drivers of biodiversity patterns. Li et al. (2025a) combined genome-scale phylogeny, fossil records, and distribution data to investigate species richness anomalies in Fraxinus across the Northern Hemisphere. Generalized linear regression combined with pathway-model analyses indicated that historical global cooling was responsible for extinctions and the resulting low richness of extant lineages in Europe, while evolutionary divergence gave rise to young, rapidly diversifying lineages in East Asia alongside long-persisting lineages in North America. Their work further demonstrated that the observed diversity anomaly between these regions stems from environmental heterogeneity coupled with evolutionary history. This study highlights how integrating phylogenomics with paleogeographic and ecological modeling can clarify the evolutionary processes underlying regional diversity heterogeneity. Advances in phylogenomics have enabled researchers to move beyond tree reconstruction toward uncovering the molecular and evolutionary mechanisms underlying trait innovation (Clark, 2023; Miller et al., 2023; Zhang et al., 2023). A robust phylogenetic framework provides a powerful tool for tracing the origin, transitions, and convergences of complex traits of organisms (Yang et al., 2023). Recent studies continue to highlight the profound insights that phylogenomics and genomics offer, particularly in understanding how traits, such as flowers (Shan et al., 2019), evolve and diversify across lineages. In the case of Asteraceae, Niu et al. (2025) employed phylogenetic analysis to investigate the genetic basis of the shape-color association in radiate capitula. Through a detailed comparison of 11 species, their findings revealed that the CCD4a gene, rarely missing, interacted with CYC2g or other regulators, played a pivotal role in regulating this association. Notably, their phylogenetic analysis traced the evolution of CCD4a across the family, demonstrating that its co-option occurred early in the evolutionary history of Asteraceae, prior to the divergence of the tribes Anthemideae and Astereae. This underscores the power of an integrative approach, which combines evidence from phylogenetics, phenotype, expression pattern, and molecular interactions, and reveals not only the origin of specific traits but also the conserved genetic mechanisms underlying trait evolution across diverse lineages. In a similar vein, Li et al. (2025b) used phylogenomic and transcriptomic data to explore the molecular mechanisms behind floral dimorphism in Viola. By comparing the gene expression profiles of Viola philippica and V. cornuta, they identified the key genetic pathways that contribute to the development of chasmogamous and cleistogamous flowers in response to photoperiod. Their phylogenetic analysis highlighted how different flower types evolved independently across species, providing a deeper understanding of the role that environmental factors, such as light, can have in shaping plant reproductive strategies. The combination of phylogenetic data and gene expression profiles illuminated how specific regulatory networks, such as those involving ARF genes and MADS-box genes, contribute to floral plasticity and reproductive success. These studies exemplify how integrating phylogenetics and genomics with trait evolution provides a powerful framework for understanding the genetic, ecological, and environmental factors that drive the diversification of plant traits. By tracing the evolutionary history of key genes and their role in trait formation, phylogenomics not only informs us about the genetic architecture of traits but also offers insights into how environmental pressures shape the adaptive landscape of plant species. As these examples illustrate, phylogenetic analysis plays a crucial role in uncovering the underlying genetic mechanisms and adaptive processes that shape plant diversity over time. Adaptive evolution occurs as populations undergo genetic changes in response to environmental pressures, facilitating their survival and reproduction in changing conditions (Savolainen et al., 2013). This process is influenced by genetic variance, which encompasses both the standing genetic variation and new mutations. The availability of standing genetic variation is critical for rapid adaptation, as it provides an immediate pool of genetic material that populations can draw from when responding to new environmental challenges (Schlötterer, 2023). In contrast, adaptation through new mutations generally requires a longer time frame for beneficial traits to emerge and spread throughout the population. Research by Han et al. (2025) demonstrates that evolutionary history plays a pivotal role in shaping genetic variance, which in turn affects how populations adapt to new environments. Populations with divergent evolutionary histories may exhibit different patterns of genetic variance, and they exhibit enhanced performance and fast adaptation when returning to a previously experienced environment, most likely due to the reuse of standing genetic variation. These findings emphasize the importance of considering evolutionary history in models of adaptation as well as ecological restoration. On the other hand, Tang et al. (2025) provide insights into how local adaptations of tree species from arid conditions evolved, and how they will respond to environmental challenges in the future. The genetic adaptation of these populations to their local environments is also tied to their capacity for adaptive responses at the population level. They identified a series of climate-associated genetic variations, which are related to the adaptation of Juniperus przewalskii to environments of the Northeastern Qinghai–Tibet Plateau, such as genes that are related to low temperature, drought, and ultraviolet radiation. Based on these genetic variations and by employing landscape genomic approaches, they forecast that most populations of the Qilian Juniper may experience a weak disturbance of local adaptation and low genomic vulnerability in response to future climate change. Together, these findings underscore the importance of both genetic variance and evolutionary history in determining adaptive capacity and inform conservation efforts aimed at preserving evolutionary potential under environmental change. Recent advances in multi-omics technologies and analytical methodologies drive transformative breakthroughs in plant phylogenetics and evolution studies. Here, we outlined five growth points in the field that are experiencing fast expansions. There are two cytoplasmic genomes, in addition to nuclear genomes, in green plants. Mitochondrial genomes, with their low mutation rate and high structural rearrangement frequency, offer new insights into plant evolution (Palmer and Herbon, 1988; Christensen, 2021; Wang et al., 2024). Yet, there are far fewer phylogenomic studies that adopt mitochondrial genomes compared to plastid genomes. As sequencing technologies improve, mitochondrial genomes are experiencing rapid growth and becoming important evidence in phylogenetic and evolution studies (Wang et al., 2024; Krawczyk et al., 2025). While nuclear genomes are preferred for inferring species relationships, cytoplasmic genomes, including mitochondria, are powerful tools for studying dispersal histories and biogeographic patterns. Furthermore, mitochondrial genomics helps address evolutionary questions, such as why mitogenomes mutate slowly in sequence while exhibiting rapid structural variation in angiosperms. These insights are key to understanding the evolutionary processes shaping plant diversity. As whole genome sequencing and assembly methods advance, chromosomal changes such as fission, fusion, and translocation can be identified more reliably and offer valuable signals for phylogenetic analysis, in addition to DNA sequence variation (Qin et al., 2021; Sun et al., 2024). Using the WGDI software (Sun et al., 2022) for ancestral karyotype evolutionary analysis, researchers identify shared chromosomal fusion breakpoints in the karyotype evolution of Camelinodae and Brassicodae, reconstruct evolutionary relationships through these chromosome structures, and demonstrate the stable inheritance of chromosome fusions (Jiang et al., 2025). Several recent chromosome-level assemblies have illustrated how large-scale chromosomal rearrangements correlate with phylogenetic divergence and trait innovation, offering new opportunities to investigate the evolutionary relevance of chromosomal architecture and to dissect how structural variation contributes to phylogenetic diversification and the emergence of lineage-specific traits (Yang et al., 2024; Cui et al., 2025). Spatial and temporal information on gene expression and cellular interactions is crucial when understanding the development of innovative traits. Spatial genomics, particularly spatial transcriptomics, allows researchers to map trait-associated cells with a high resolution, capture the spatial of identify traits within and gene expression patterns over time and This technology provides insights into the of gene linking molecular pathways to variation et al., 2025). Furthermore, it enables the of or gene offering a deeper understanding of how plant traits are and These the foundation to study the molecular mechanisms behind plant trait innovations and adaptation et al., 2023). are genomic as DNA sequence changes in et al., These and traits by gene structure and The rise of sequencing technologies has the accurate and of at a genome-wide analysis as a tool for plant adaptation and evolutionary et al., 2021; et al., 2024). This such as chromosomal related to local adaptation et al., 2019), inferring introgression at the population et al., and exploring structural variation during et al., 2019). In addition to their contribute to genetic divergence among lineages and providing phylogenetic signals beyond et al., 2025). genome sequencing with high provides the most molecular yet this can be when sample are In contrast, has such as cost large sample and data sequencing is remains for genomic and will an increasingly role in studies of genetic evolutionary dynamics, and adaptive particularly for species with large genomes et al., Moreover, as a valuable data source for phylogenetic reconstruction et al., Zhang et al., et al., 2025). We are to and Wang for offering valuable This work was by the of and the local for and technology development

Citations

Discussions

Related