Evolutionary-scale prediction of atomic-level protein structure with a language model
2023/03/16 by Zeming Lin, Halil Akin, Roshan Rao +12 · 269 citations
Biochemistry, Genetics and Molecular Biology · #Genomics and Phylogenetic Studies #Machine Learning in Bioinformatics #RNA and protein synthesis mechanisms
paper · doi:10.1126/science.ade2574
openalex publication_date 2023/03/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Recent advances in machine learning have leveraged evolutionary information in multiple sequence alignments to predict protein structure. We demonstrate direct inference of full atomic-level protein structure from primary sequence using a large language model. As language models of protein sequences are scaled up to 15 billion parameters, an atomic-resolution picture of protein structure emerges in the learned representations. This results in an order-of-magnitude acceleration of high-resolution structure prediction, which enables large-scale structural characterization of metagenomic proteins. We apply this capability to construct the ESM Metagenomic Atlas by predicting structures for >617 million metagenomic protein sequences, including >225 million that are predicted with high confidence, which gives a view into the vast breadth and diversity of natural proteins.
Citations
Cited by
- Artificial intelligence in bioinformatics: a survey
- Self-driving labs for biotechnology
- ProGen2: Exploring the boundaries of protein language models
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- Generalizable and scalable protein stability prediction with rewired protein generative models
- Harnessing AlphaFold to reveal hERG channel conformational state secrets
- RAG-ESM: Improving pretrained protein language models via sequence retrieval
- Generative AI as a tool to accelerate the field of ecology
- Evolving strategies for virus discovery
- AI‐Physics‐Experiment Trinity for Integrated Protein Dynamics Modeling
- Integrative modeling of seasonal influenza evolution via AI-powered antigenic cartography
- Recent advances in the inference of deep viral evolutionary history
- Generative AI for synthetic biology: Designing biological parts, circuits, and genomes
- FrustrAI-Seq: Scaling Local Energetic Frustration to the Protein Sequence Space
- Machine and deep learning to predict viral fusion peptides
- Training data composition determines machine learning generalization and biological rule discovery
- Learning the language of plant immunity: opportunities and challenges for AI-assisted modelling of fungal effector x host protein complexes
- Identification of cognate recombination directionality factors for large serine recombinases by virtual pulldown
- Artificial intelligence-driven computational methods for antibody design and optimization
- Variant effect predictor correlation with functional assays is reflective of clinical classification performance
- De novo design of protein structure and function with RFdiffusion
- ProteinCLIP: enhancing protein language models with natural language
- ProSTAGE: Predicting Effects of Mutations on Protein Stability by Using Protein Embeddings and Graph Convolutional Networks
- Active non-redundancy and viral orchestration sustain diel microbial successions in the coastal ocean
- IMMREP25: Unseen Peptides
- Emerging technologies for the discovery of biosynthetic genes in plants
- InstructPLM-mu: 1-Hour Fine-Tuning of ESM2 Beats ESM3 in Protein Mutation Predictions
- Accurate structure prediction of biomolecular interactions with AlphaFold 3
- Fast and accurate protein structure search with Foldseek
- Computational strategies for cross-species knowledge transfer
- Limitations of the refolding pipeline for de novo protein design
- Acquiring Improved Protein Variants With Probabilistic Preferential Learning
- Leveraging large language models to predict antibody biological activity against influenza A hemagglutinin
- Mind the gap: the challenges and opportunities for genomics-driven harnessing of plant metabolic diversity for therapeutic applications
- A century of vitamin E research: The innovative journey from basic biology to synthetic bio‐manufacturing
- AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
- Advancing Antibody Engineering through Synthetic Evolution and Machine Learning
- Machine-guided dual-objective protein engineering for deimmunization and therapeutic functions
- Multi-pathway feature-level interpretability in MHC-I antigen presentation via concept-based modeling
- HapScoreDB: a database of protein language model functional scores for haplotype-resolved protein sequences
- ConforNets: Latents-Based Conformational Control in OpenFold3
- DefensePredictor: A machine learning model to discover prokaryotic immune systems
- MDCompress: better, faster compression of molecular dynamics simulation trajectories
- Protein Language Model Embeddings Improve Generalization of Implicit Transfer Operators
- Contrastive Geometric Learning Unlocks Unified Structure- and Ligand-Based Drug Design
- Attentive multilayer fusion for vision transformers
- Guiding questions to avoid data leakage in biological machine learning applications
- USHER: Guiding Foundation Model Representations through Distribution Shifts
- EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation
- Connecting the dots: deep learning-based automated model building methods in cryo-EM
- Evolution of the Tri-PDZ Domain in PSD95 (DLG-4 Gene)
- Taxonomic expansion and reorganization of Flaviviridae
- Protein Language Model Fitness Is a Matter of Preference
- Learning the shape of protein microenvironments with a holographic convolutional neural network
- Marine-derived products as pharmaceutical treasure troves: a focus on recent research techniques and potential bioactive activities of marine peptides
- The Infant Gut Virome: Knowns, Unknowns, and Avenues for Future Studies
- Opportunities for artificial intelligence and synthetic biology in designing living drug delivery systems
- ExplainBind: Explainable Physicochemical Determinants of Protein–Ligand Binding via Non-Covalent Interactions
- Relaxed Sequence Sampling for Diverse Protein Design
- Protein language models reveal evolutionary constraints on synonymous codon choice
- Learning the PTM Code through a Coarse-to-Fine, Mechanism-Aware Framework
- GBNL: Graded Betti Number Learning of Complex Biological Data
- Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMs
- QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
- ODesign: A World Model for Biomolecular Interaction Design
- CAP: Commutative Algebra Prediction of Protein-Nucleic Acid Binding Affinities
- A Multimodal Human Protein Embeddings Database: DeepDrug Protein Embeddings Bank (DPEB)
- Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing
- Meta-Learning for Cross-Task Generalization in Protein Mutation Property Prediction
- g-DPO: Scalable Preference Optimization for Protein Language Models
- Accurate proteome-wide missense variant effect prediction with AlphaMissense
- Protein generation with embedding learning for motif diversification
- SSEmb: A joint embedding of protein sequence and structure enables robust variant effect predictions
- Inference-Time Compute Scaling For Flow Matching
- Accelerating Genetic Sensor Development, Scale-up, and Deployment Using Synthetic Biology
- AlphaFold 2, but not AlphaFold 3, predicts confident but unrealistic β-solenoid structures for repeat proteins
- Are protein language models the new universal key?
- High‐Throughput Evaluation of Natural Diversity of F‐Type <scp>ATP</scp> Synthase Rotor Ring Stoichiometries
- Simulating 500 million years of evolution with a language model
- Pedoman Peri-operatif Renal Assessment
- Domainator, a flexible software suite for domain-based annotation and neighborhood analysis, identifies proteins involved in antiviral systems
- Protein Language Models: Is Scaling Necessary?
- Prediction-Augmented Trees for Reliable Statistical Inference
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- Adaptive Individual Uncertainty under Out-Of-Distribution Shift with Expert-Routed Conformal Prediction
- Kernel-Based Evaluation of Conditional Biological Sequence Models
- Representation-Based Exploration for Language Models: From Test-Time to Post-Training
- Protein as a Second Language for LLMs
- Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity
- PRISM: Enhancing Protein Inverse Folding through Fine-Grained Retrieval on Structure-Sequence Multimodal Representations
- Fast and Interpretable Protein Substructure Alignment via Optimal Transport
- ProteinAE: Protein Diffusion Autoencoders for Structure Encoding
- Calibrating Generative Models
- Proximity analysis of native proteomes reveals phenotypic modifiers in a mouse model of autism and related neurodevelopmental conditions
- Linear motifs regulating protein secretion, sorting and autophagy in Leishmania parasites are diverged with respect to their host equivalents
- Open Biomedical Network Benchmark A Python Toolkit for Benchmarking Datasets with Biomedical Networks
- Lignocellulose-mediated selection of potential halophilic PET-degrading enzymes from mangrove soil
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- Large Language Models Meet Virtual Cell: A Survey
- BioBlobs: Unsupervised Discovery of Functional Substructures for Protein Function Prediction
- Evolutionary Profiles for Protein Fitness Prediction
- Native Hybrid Attention for Efficient Sequence Modeling
- Need for speed: advances in the era of high throughput interaction proteomics
- Evolutionary transfer learning enables organism-wide inference of mammalian enhancer landscapes
- BOTANIC-0: a series of foundation models for plant genomic data
- GeneCAD: Plant Genome Annotation with a DNA Foundation Model
- UniOTalign: A Global Matching Framework for Protein Alignment via Optimal Transport
- Adaptive Protein Design Protocols and Middleware
- Physicochemically Informed Dual-Conditioned Generative Model of T-Cell Receptor Variable Regions for Cellular Therapy
- PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- TCR-EML: Explainable Model Layers for TCR-pMHC Prediction
- Self-Speculative Masked Diffusions
- Distilled Protein Backbone Generation
- NS-Pep: De novo Peptide Design with Non-Standard Amino Acids
- GeoGraph: Geometric and Graph-based Ensemble Descriptors for Intrinsically Disordered Proteins
- Predicting adaptive immune receptor specificities by machine learning is a data generation problem
- BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning
- AReUReDi: Annealed Rectified Updates for Refining Discrete Flows with Multi-Objective Guidance
- Commutative algebra neural network reveals genetic origins of diseases
- Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization
- Let Physics Guide Your Protein Flows: Topology-aware Unfolding and Generation
- Towards Understanding the Shape of Representations in Protein Language Models
- HyperHELM: Hyperbolic Hierarchy Encoding for mRNA Language Modeling
- Advancing Protein Ensemble Predictions Across the Order–Disorder Continuum
- TR2-D2: Tree Search Guided Trajectory-Aware Fine-Tuning for Discrete Diffusion
- Beyond the canonical: The role of post-transcriptional regulation in drug-target interaction prediction
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
- Navigating prokaryotic viral genome analysis from metagenomic data
- Completion of partial structures using Patterson maps with the <i>CrysFormer</i> machine-learning model
- VN-EGNN: E(3)- and SE(3)-Equivariant Graph Neural Networks with Virtual Nodes Enhance Protein Binding Site Identification
- <scp>AI</scp> Methods for Antimicrobial Peptides: Progress and Challenges
- Protein language models trained on biophysical dynamics inform mutation effects
- Active learning-guided optimization of cell-free biosensors for lead testing in drinking water
- Perspective on structure predictions of disorder
- Bacterial proteome foundation model enhances functional prediction from enzymes to ecological interactions
- Discontinuous Epitope Fragments as Sufficient Target Templates for Efficient Binder Design
- The cellular dogma
- Planner Aware Path Learning in Diffusion Language Models Training
- Foundation models in plant molecular biology: advances, challenges, and future directions
- ET-Pfam: Ensemble transfer learning for protein family prediction
- From text to traits: exploring the role of large language models in plant breeding
- From data to discovery: leveraging big data in plant natural products biosynthesis research
- Autonomous laboratories in China: an embodied intelligence-driven platform to accelerate chemical discovery
- Sequence modeling and design from molecular to genome scale with Evo
- Fine-tuning protein language models to understand the functional impact of missense variants
- ProtFun: A Protein Function Prediction Model Using Graph Attention Networks with a Protein Large Language Model
- BFVD—a large repository of predicted viral protein structures
- Comprehensive prediction and analysis of human protein essentiality based on a pretrained large language model
- Rational Design Principles for <i>De Novo</i> α-Helical Peptide Barrels with Dynamic Conductive Channels
- Themis and Grb2 form a constitutive structural hub in T cell receptor signalling
- Foundation models in bioinformatics
- ABConformer: Physics-inspired Sliding Attention for Antibody-Antigen Interface Prediction
- Engineering surface electrostatics affords control over morphological preference, synergy, and activity in polymer degrading enzymes
- Investigating the ability of deep learning-based structure prediction to extrapolate and/or enrich the set of antibody CDR canonical forms
- Proteomic analysis of the Aggregation Factor from the sponge <i>Clathria (Microciona) prolifera</i> suggests an ancient protein domain toolkit for allorecognition in animals
- Design of intrinsically disordered protein variants with diverse structural properties
- ProtHyena: A fast and efficient foundation protein language model at single amino acid Resolution
- The social and structural architecture of the yeast protein interactome
- Context-aware geometric deep learning for protein sequence design
- Can ChatGPT pass Glycobiology?
- GRAM-DTI: adaptive multimodal representation learning for drug target interaction prediction
- Alphafold2 refinement improves designability of large de novo proteins
- Twin Peaks: Dual-Head Architecture for Structure-Free Prediction of Protein-Protein Binding Affinity and Mutation Effects
- SpecMER: Fast Protein Generation with K-mer Guided Speculative Decoding
- Learning to Align Molecules and Proteins: A Geometry-Aware Approach to Binding Affinity
- Direct prediction of intrinsically disordered protein conformational properties from sequence
- The future of machine learning for small-molecule drug discovery will be driven by data
- <scp> S <sup>2</sup> </scp> ‐ <scp>PepAnalyst</scp> : A Web Tool for Predicting Plant Small Signalling Peptides
- Computational enzyme design by catalytic motif scaffolding
- Leveraging ancestral sequence reconstruction for protein representation learning
- VESTIGE: A Knowledge-Guided Masking Strategy for Corruption-Aware Fine-Tuning of Genomic Transformers, Validated on Ancient DNA Reconstruction
- Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
- From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
- Protein design drives synthetic biology research of plant natural products
- Genome modeling and design across all domains of life with Evo 2
- Benchmarking gene embeddings from sequence, expression, network, and text models for functional prediction tasks
- Sequence-aware Prediction of Point Mutation-induced Effects on Protein-Protein Binding Affinity using Deep Learning
- Squidly: Enzyme Catalytic Residue Prediction Harnessing a Biology-Informed Contrastive Learning Framework
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language Models
- DomDiff: protein family and domain annotation via diffusion model and ESM2 embedding
- SpheronizaTor: Spherical Voxelization for Interpretable Protein Microenvironment Modeling
- The State of Peptide Detectability in Computational Proteomics and Guidelines for AI Applications
- Integrative Learning of Disentangled Representations from Single-Cell RNA-Sequencing Datasets
- Unraveling circadian rhythms—computational insights into molecular mechanisms
- PLM-interact: extending protein language models to predict protein-protein interactions
- DrugForm-DTA: Towards real-world drug-target binding affinity model
- Training Compute-Optimal Protein Language Models
- Leveraging learned representations and multitask learning for lysine methylation site discovery
- Evaluating variant effect prediction across viruses
- A foundation model of transcription across human cell types
- Beyond protein functions: evaluating completeness, coherence, and consistency of genome-scale function annotations
- EZpred: improving deep learning-based enzyme function prediction using unlabeled sequence homologs
- AI-assisted protein design to rapidly convert antibody sequences to intrabodies targeting diverse peptides and histone modifications
- SSAlign: Ultrafast and Sensitive Protein Structure Search at Scale
- Unlocking Genetic Diversity in Colombian Cassava Landraces for Accelerated Breeding
- AlphaFold2 and ESMFold: A large-scale pairwise model comparison of human enzymes upon Pfam functional annotation
- Descriptive power and predictive limits of a discrete Hasimoto--DNLS model of protein backbone structure
- Using artificial intelligence to document the hidden RNA virosphere
- Adverse reactions to the use of large language models in social interactions
- Evaluation of predictions of disordered binding regions in the CAID2 experiment
- Detection of circular permutations by Protein Language Models
- TransHLA: a Hybrid Transformer model for HLA-presented epitope detection
- Missense mutation knowledge can decrease prediction inaccuracies on protein secondary structure
- Functionally Important Residues from Graph Analysis of Coevolved Dynamic couplings
- <i>Slice'N'Dice</i> : maximizing the value of predicted models for structural biologists
- Structural characterization of the ACDC domain from ApiAP2 proteins, a potential molecular target against apicomplexan parasites
- Enhancing predictions of protein stability changes induced by single mutations using MSA-based language models
- Bacterial sensor evolved by decreasing complexity
- Annotation Vocabulary (Might Be) All You Need
- Rapid simulation of glycoprotein structures by grafting and steric exclusion of glycan conformer libraries
- High-throughput algorithm predicts F-Type ATP synthase rotor ring stoichiometries of 8 to 27 protomers
- A Perspective on the Prospective Use of AI in Protein Structure Prediction
- Recent Advances and Challenges in Protein Structure Prediction
- Contextual protein and antibody encodings from equivariant graph transformers
- Understanding language model scaling for protein fitness prediction
- Optimizing protein tokenization: reduced amino acid alphabets for efficient and accurate protein language models
- Evaluating completeness, coherence, and consistency of genome-scale function annotations
- The structural chemistry and biosynthesis of chlorophylls
- Adversarial Sequence Mutations in AlphaFold and ESMFold Reveal Nonphysical Structural Invariance, Confidence Failures, and Concerns for Protein Design
- Tools For Building Artificial Biological Nanostructures
- Deciphering small sequence differences in T cell receptor-antigen pairing
- iMFP-LG: Identify Novel Multi-functional Peptides Using Protein Language Models and Graph-based Deep Learning
- Evaluating zero‐shot prediction of monomeric protein design success by <scp>AlphaFold</scp> , <scp>ESMFold</scp> , and <scp>ProteinMPNN</scp>
- Optimizing enzyme thermostability by combining multiple mutations using protein language model
- Protein engineering in the deep learning era
- LUCID: An Integrative Approach for Target Discovery and dsRNA Design in Plant Fungal Pathogens
- RemoteFoldSet: Benchmarking Structural Awareness of Protein Language Models
- Quantum Autoencoder: An efficient approach to quantum feature map generation
- Structure-guided discovery of ancestral CRISPR-Cas13 ribonucleases
- Unveiling the ghost: machine learning’s impact on the landscape of virology
- Securing the Language of Life: Inheritable Watermarks from DNA Language Models to Proteins
- Monte Carlo Tree Diffusion with Multiple Experts for Protein Design
- ShortListing Model: A Streamlined SimplexDiffusion for Discrete Variable Generation
- Property-Isometric Variational Autoencoders for Sequence Modeling and Design
- Multimodal Regression for Enzyme Turnover Rates Prediction
- Tokenizing Loops of Antibodies
- Topotein: Topological Deep Learning for Protein Representation Learning
- AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences
- SOLD: SELFIES-based Objective-driven Latent Diffusion
- SafeProtein: Red-Teaming Framework and Benchmark for Protein Foundation Models
- Categorizing prediction modes within low-pLDDT regions of <i>AlphaFold</i> 2 structures: near-predictive, pseudostructure and barbed wire
- Morphology-Specific Peptide Discovery via Masked Conditional Generative Modeling
- Computational design of a soluble mimic of the outer membrane <scp>LPS</scp> transport protein <scp>LptD</scp> suitable for screening of antibiotics
- Intelligent Spectrum Management in Satellite Communications
- Supervised learning of protein variant effects across large‐scale mutagenesis datasets
- Tracking World States with Language Models: State-Based Evaluation Using Chess
- CrystalICL: Enabling In-Context Learning for Crystal Generation
- TopoBind: Multi-Modal Prediction of Antibody-Antigen Binding Free Energy via Sequence Embeddings and Structural Topology
- Energy-Based Flow Matching for Generating 3D Molecular Structure
- ProtSAE: Disentangling and Interpreting Protein Language Models via Semantically-Guided Sparse Autoencoders
- Predicting Drug-Drug Interactions Using Heterogeneous Graph Neural Networks: HGNN-DDI
- Sparse Autoencoders for Low-N Protein Function Prediction and Design
- Learning Protein-Ligand Binding in Hyperbolic Space
- Equi-mRNA: Protein Translation Equivariant Encoding for mRNA Language Models
- Protein Set Transformer: a protein-based genome language model to power high-diversity viromics
- Deep Learning Model for Amyloidogenicity Prediction using a Pre-trained Protein LLM
- <scp>CASP16</scp> Protein Monomer Structure Prediction Assessment
- BConformeR: A Conformer Based on Mutual Sampling for Unified Prediction of Continuous and Discontinuous Antibody Binding Sites
- A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design
- Highlights of Model Quality Assessment in <scp>CASP16</scp>
- Probing the Dark Energy in the Functional Protein Universe
- DeepFold-PLM: accelerating protein structure prediction via efficient homology search using protein language models
- Multiple protein structure alignment at scale with FoldMason
- Learning the language of protein-protein interactions
- FrustraMPNN: An ultra-fast deep learning tool for proteome-scale analysis of deep mutational single-residue local energetic frustration in proteins
- Designing de novo TIM Barrels: Insights into Stabilization, Diversification, and Functionalization Strategies
- FlowBack-Adjoint: Physics-Aware and Energy-Guided Conditional Flow-Matching for All-Atom Protein Backmapping
- Scaling and Data Saturation in Protein Language Models
- Conversations over Clicks: Impact of Chatbots on Information Search in Interdisciplinary Learning
- Topological Learning Prediction of Virus-like Particle Stoichiometry and Stability
- Zero-Shot Learning with Subsequence Reordering Pretraining for Compound-Protein Interaction
Related