vix.ing · top · new · best · stats

Learning inverse folding from millions of predicted structures

2022/04/10 by Chloe Hsu, Robert Verkuil, Jason Liu +5 · 1 voice · 74 citations
Biochemistry, Genetics and Molecular Biology · Materials Science · #Enzyme Structure and Function #Protein Structure and Dynamics #RNA and protein synthesis mechanisms

paper · pdf · doi:10.1101/2022.04.10.487779

openalex publication_date 2022/04/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30

Abstract

Abstract We consider the problem of predicting a protein sequence from its backbone atom coordinates. Machine learning approaches to this problem to date have been limited by the number of available experimentally determined protein structures. We augment training data by nearly three orders of magnitude by predicting structures for 12M protein sequences using AlphaFold2. Trained with this additional data, a sequence-to-sequence transformer with invariant geometric input processing layers achieves 51% native sequence recovery on structurally held-out backbones with 72% recovery for buried residues, an overall improvement of almost 10 percentage points over existing methods. The model generalizes to a variety of more complex tasks including design of protein complexes, partially masked structures, binding interfaces, and multiple states.

Citations

Cited by

Discussions

Related