2024/10/21 by Bahar Ali, Ali, Bahar, Anwar Shah +8
Biochemistry, Genetics and Molecular Biology · #FOS: Biological sciences #FOS: Computer and information sciences #Gene expression and cancer classification #Genetics, Bioinformatics, and Biomedical Research #Machine Learning (cs.LG) #Machine Learning in Bioinformatics #Quantitative Methods (q-bio.QM)
paper · pdf · doi:10.48550/arxiv.2410.17293
openalex publication_date 2024/10/21 · openalex created_date 2024/11/13 · openalex updated_date 2026/07/28
Advanced automated AI techniques allow us to classify protein sequences and discern their biological families and functions. Conventional approaches for classifying these protein families often focus on extracting N-Gram features from the sequences while overlooking crucial motif information and the interplay between motifs and neighboring amino acids. Recently, convolutional neural networks have been applied to amino acid and motif data, even with a limited dataset of well-characterized proteins, resulting in improved performance. This study presents a model for classifying protein families using the fusion of 1D-CNN, BiLSTM, and an attention mechanism, which combines spatial feature extraction, long-term dependencies, and context-aware representations. The proposed model (ProFamNet) achieved superior model efficiency with 450,953 parameters and a compact size of 1.72 MB, outperforming the state-of-the-art model with 4,578,911 parameters and a size of 17.47 MB. Further, we achieved a higher F1 score (98.30% vs. 97.67%) with more instances (271,160 vs. 55,077) in fewer training epochs (25 vs. 30).