2002/07/15 by Vo Anh, V. V. Anh, K. S. Lau +2 · 1 citation
Biochemistry, Genetics and Molecular Biology · Physics and Astronomy · #Fractal and DNA sequence analysis #Genomics and Phylogenetic Studies #Machine Learning in Bioinformatics #physics.bio-ph #q-bio
paper · pdf · doi:10.1103/physreve.66.031910
published as Phys. Rev. E, vol. 66, (2002) 031910. · 9 pages,5 figures, Accepted for publication by Phys. Rev. E
arxiv created 2002/07/15 · openalex publication_date 2002/09/24 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper considers the problem of matching a fragment to an organism using its complete genome. Our method is based on the probability measure representation of a genome. We first demonstrate that these probability measures can be modeled as recurrent iterated function systems (RIFS) consisting of four contractive similarities. Our hypothesis is that the multifractal characteristics of the probability measure of a complete genome, as captured by the RIFS, is preserved in its reasonably long fragments. We compute the RIFS of fragments of various lengths and random starting points, and compare with that of the original sequence for recognition using the Euclidean distance. A demonstration on five randomly selected organisms supports the above hypothesis.