2020/09/22 by Vincenzo Bonnici, Bonnici, Vincenzo, Giuditta Franco +3
Biochemistry, Genetics and Molecular Biology · #FOS: Biological sciences #FOS: Computer and information sciences #Fractal and DNA sequence analysis #Genomics (q-bio.GN) #Genomics and Phylogenetic Studies #Information Theory (cs.IT) #RNA and protein synthesis mechanisms
paper · pdf · doi:10.48550/arxiv.2009.10449
openalex publication_date 2020/09/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Genomes may be analyzed from an information viewpoint as very long strings, containing functional elements of variable length, which have been assembled by evolution. In this work an innovative information theory based algorithm is proposed, to extract significant (relatively small) dictionaries of genomic words. Namely, conceptual analyses are here combined with empirical studies, to open up a methodology for the extraction of variable length dictionaries from genomic sequences, based on the information content of some factors. Its application to human chromosomes highlights an original inter-chromosomal similarity in terms of factor distributions.