2017/12/18 by Adam J. Riesselman, Riesselman, Adam J., John Ingraham +3 · 2 citations
Biochemistry, Genetics and Molecular Biology · #Biological Physics (physics.bio-ph) #Disordered Systems and Neural Networks (cond-mat.dis-nn) #FOS: Biological sciences #FOS: Computer and information sciences #FOS: Physical sciences #Genetics, Bioinformatics, and Biomedical Research #Genomics and Phylogenetic Studies #Machine Learning (stat.ML) #Quantitative Methods (q-bio.QM) #RNA and protein synthesis mechanisms
paper · pdf · doi:10.48550/arxiv.1712.06527
openalex publication_date 2017/12/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The functions of proteins and RNAs are determined by a myriad of interactions between their constituent residues, but most quantitative models of how molecular phenotype depends on genotype must approximate this by simple additive effects. While recent models have relaxed this constraint to also account for pairwise interactions, these approaches do not provide a tractable path towards modeling higher-order dependencies. Here, we show how latent variable models with nonlinear dependencies can be applied to capture beyond-pairwise constraints in biomolecules. We present a new probabilistic model for sequence families, DeepSequence, that can predict the effects of mutations across a variety of deep mutational scanning experiments significantly better than site independent or pairwise models that are based on the same evolutionary data. The model, learned in an unsupervised manner solely from sequence information, is grounded with biologically motivated priors, reveals latent organization of sequence families, and can be used to extrapolate to new parts of sequence space