2019/11/19 by Cole Casey, Christopher Ott, Cole, Casey A +5
Biochemistry, Genetics and Molecular Biology · Computer Science · #Machine Learning in Bioinformatics #Algorithms and Data Compression #Bioinformatics and Genomic Networks
paper · pdf · doi:10.48550/arxiv.1911.08614
Large scale initiatives such as the Human Genome Project, Structural\nGenomics, and individual research teams have provided large deposits of genomic\nand proteomic data. The transfer of data to knowledge has become one of the\nexisting challenges, which is a consequence of capturing data in databases that\nare optimally designed for archiving and not mining. In this research, we have\ntargeted the Protein Databank (PDB) and demonstrated a transformation of its\ncontent, named PDBMine, that reduces storage space by an order of magnitude,\nand allows for powerful mining in relation to the topic of protein structure\ndetermination. We have demonstrated the utility of PDBMine in exploring the\nprevalence of dimeric and trimeric amino acid sequences and provided a\nmechanism of predicting protein structure.\n