2013/08/13 by Boris Brimkov, Brimkov, Boris, Valentin E. Brimkov +1
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Algorithms and Data Compression #Combinatorics (math.CO) #FOS: Biological sciences #FOS: Mathematics #Fractal and DNA sequence analysis #Genomics and Phylogenetic Studies #Machine Learning in Bioinformatics #Quantitative Methods (q-bio.QM) #RNA and protein synthesis mechanisms #math.CO #q-bio.QM
paper · pdf · doi:10.48550/arxiv.1308.2885
16 pages, 2 figures (the first with 2 subfigures and the second with 8 subfigures)
arxiv created 2013/08/13 · openalex publication_date 2013/08/13 · arxiv updated 2013/08/14 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
Tools that effectively analyze and compare sequences are of great importance in various areas of applied computational research, especially in the framework of molecular biology. In the present paper, we introduce simple geometric criteria based on the notion of string linearity and use them to compare DNA sequences of various organisms, as well as to distinguish them from random sequences. Our experiments reveal a significant difference between biosequences and random sequences - the former having much higher deviation from linearity than the latter - as well as a general trend of increasing deviation from linearity between primitive and biologically complex organisms.