vix.ing · top · new · best · stats

The Protein Family Classification in Protein Databases via Entropy Measures

2018/06/12 by R. P. Mondaini, Mondaini, R. P., S. C. de Albuquerque Neto +1
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomolecules (q-bio.BM) #Computational Drug Discovery Methods #FOS: Biological sciences #Fractal and DNA sequence analysis #Machine Learning in Bioinformatics #q-bio.BM

paper · pdf · doi:10.48550/arxiv.1806.05172

arxiv created 2018/06/12 · openalex publication_date 2018/06/12 · arxiv updated 2018/06/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In the present work, we review the fundamental methods which have been developed in the last few years for classifying into families and clans the distribution of amino acids in protein databases. This is done through functions of random variables, the Entropy Measures of probabilities of occurrence of the amino acids. An intensive study of the Pfam databases is presented with restrictions to families which could be represented by rectangular arrays of amino acids with m rows (protein domains) and n columns (amino acids). This work is also an invitation to scientific research groups worldwide to undertake the statistical analysis with different numbers of rows and columns since we believe in the mathematical characterization of the distribution of amino acids as a fundamental insight on the determination of protein structure and evolution.

Related