vix.ing · top · new · best · stats · spec

Evolutionary implications of a power-law distribution of protein family sizes

1999/08/17 by Joel S. Bader, Bader, Joel S.
Biochemistry, Genetics and Molecular Biology · Physics and Astronomy · #Biological Physics (physics.bio-ph) #FOS: Biological sciences #FOS: Physical sciences #Fractal and DNA sequence analysis #Genomics and Phylogenetic Studies #Machine Learning in Bioinformatics #Quantitative Biology (q-bio) #physics.bio-ph #q-bio

paper · pdf · doi:10.48550/arxiv.physics/9908032

11 pages, 3 figures

arxiv created 1999/08/17 · openalex publication_date 1999/08/17 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Current-day genomes bear the mark of the evolutionary processes. One of the strongest indications is the sequence homology among families of proteins that perform similar biological functions in different species. The number of proteins in a family can grow over time as genetic information is duplicated through evolution. We explore how evolution directs the size distribution of these families. Theoretical predictions for family sizes are obtained from two models, one in which individual genes duplicate and a second in which the entire genome duplicates. Predictions from these models are compared with the family size distributions for several organisms whose complete genome sequence is known. We find that protein family size distributions in nature follow a power-law distribution. Comparing these results to the model systems, we conclude that genome duplication is the dominant mechanism leading to increased genetic material in the species considered.

Related