2016/05/23 by Laura Deming, Sasha Targ, Deming, Laura +7 · 1 citation
Biochemistry, Genetics and Molecular Biology · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Genomics and Phylogenetic Studies #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Bioinformatics #Neural and Evolutionary Computing (cs.NE) #RNA and protein synthesis mechanisms
paper · pdf · doi:10.48550/arxiv.1605.07156
openalex publication_date 2016/05/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Each human genome is a 3 billion base pair set of encoding instructions.\nDecoding the genome using deep learning fundamentally differs from most tasks,\nas we do not know the full structure of the data and therefore cannot design\narchitectures to suit it. As such, architectures that fit the structure of\ngenomics should be learned not prescribed. Here, we develop a novel search\nalgorithm, applicable across domains, that discovers an optimal architecture\nwhich simultaneously learns general genomic patterns and identifies the most\nimportant sequence motifs in predicting functional genomic outcomes. The\narchitectures we find using this algorithm succeed at using only RNA expression\ndata to predict gene regulatory structure, learn human-interpretable\nvisualizations of key sequence motifs, and surpass state-of-the-art results on\nbenchmark genomics challenges.\n