2024/12/17 by Filippo Stocco, Maria Artigues-Lleixa, Stocco, Filippo +10 · 1 voice · 3 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Topic Modeling #Machine Learning in Bioinformatics #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2412.12979
Protein language models (pLMs) have demonstrated success at generating functional proteins across vast sequence spaces but lack the ability to design high-fitness variants on demand. Here, we iteratively guide pLMs toward user-defined objectives by applying reinforcement learning (RL). We demonstrate that RL can steer pLMs toward various protein properties, such as topologies or binding affinities, in a few iterations through long evolutionary trajectories. We apply our framework to the design of epidermal growth factor receptor (EGFR) binders, achieving a 26-fold increase in binding affinity in two iterations.