2016/05/23 by Laura Deming, Sasha Targ, Deming, Laura +8 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Genomics and Phylogenetic Studies #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Bioinformatics #Neural and Evolutionary Computing (cs.NE) #RNA and protein synthesis mechanisms #cs.AI #cs.LG #cs.NE #stat.ML
paper · pdf · doi:10.48550/arxiv.1605.07156
10 pages, 4 figures
arxiv created 2016/05/23 · openalex publication_date 2016/05/23 · arxiv updated 2016/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Each human genome is a 3 billion base pair set of encoding instructions. Decoding the genome using deep learning fundamentally differs from most tasks, as we do not know the full structure of the data and therefore cannot design architectures to suit it. As such, architectures that fit the structure of genomics should be learned not prescribed. Here, we develop a novel search algorithm, applicable across domains, that discovers an optimal architecture which simultaneously learns general genomic patterns and identifies the most important sequence motifs in predicting functional genomic outcomes. The architectures we find using this algorithm succeed at using only RNA expression data to predict gene regulatory structure, learn human-interpretable visualizations of key sequence motifs, and surpass state-of-the-art results on benchmark genomics challenges.