vix.ing · top · new · best · stats

Protein Language Models: Is Scaling Necessary?

2024/09/23 by Quentin Fournier, Robert M. Vernon, Almer M. van der Sloot +3 · 1 voice · 4 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Machine Learning in Bioinformatics #Natural Language Processing Techniques #Topic Modeling

paper · doi:10.1101/2024.09.23.614603

openalex publication_date 2024/09/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/14

Abstract

Public protein sequence databases contain samples from the fitness landscape explored by nature. Protein language models (pLMs) pre-trained on these sequences aim to capture this landscape for tasks like property prediction and protein design. Following the same trend as in natural language processing, pLMs have continuously been scaled up. However, the premise that scale leads to better performance assumes that source databases provide an accurate representation of the underlying fitness landscape, which is likely false. By developing an efficient codebase, designing a modern architecture, and addressing data quality concerns such as sample bias, we introduce AMPLIFY, a best-in-class pLM that is orders of magnitude less expensive to train and deploy than previous models. Furthermore, to support the scientific community and democratize the training of pLMs, we have open-sourced AMPLIFY’s pre-training codebase, data, and model checkpoints.

Citations

Cited by

Discussions

Related