vix.ing · top · new · best · stats · spec

Di-nucleotide Entropy as a Measure of Genomic Sequence Functionality

2006/11/17 by Dmitri Parkhomchuk, Parkhomchuk, Dmitri
Biochemistry, Genetics and Molecular Biology · #FOS: Biological sciences #Genomics (q-bio.GN) #Genomics and Phylogenetic Studies #Machine Learning in Bioinformatics #RNA and protein synthesis mechanisms #q-bio.GN

paper · pdf · doi:10.48550/arxiv.q-bio/0611059

10 pages, 7 figures, grammatical revision

openalex publication_date 2006/11/17 · arxiv created 2006/12/19 · arxiv updated 2009/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Considering vast amounts of genomic sequences of mostly unknown functionality, in-silico prediction of functional regions is an important enterprise. Many genomic browsers employ GC content, which was observed to be elevated in gene-rich functional regions. This report shows that the entropy of di- and tri-nucleotides distributions provides a superior measure of genomic sequence functionality, and proposes an explanation on why the GC content must be elevated (closer to 50%) in functional regions. Regions with high entropy strongly co-localize with exons and provide genome-wide evidences of purifying selection acting on non-coding regions, such as decreased SNPs density. The observations suggest that functional non-coding regions are optimised for mutation load in a way, that transition mutations have less impact on functionality than transversions, leading to the decrease in transversions to transitions ratio in functional regions.

Citations

Related