Toward the Explainability of Protein Language Models
2025/06/24 by Hunklinger, Andrea, Ferruz, Noelia · 2 voices · 1 citation
#Biomolecules (q-bio.BM) #FOS: Biological sciences
paper · doi:10.48550/arxiv.2506.19532
Abstract
Protein language models (pLMs) excel in a variety of tasks that range from structure prediction to the design of functional enzymes. However, these models operate as black boxes, and their underlying working principles remain unclear. Here, we survey emerging applications of explainable artificial intelligence (XAI) to pLMs and describe the potential of XAI in protein research. We divide the workflow of protein AI modeling into four information contexts: (i) training sequences, (ii) input prompt, (iii) model architecture, and (iv) input-output pairs. For each, we describe existing methods and applications of XAI. Additionally, from published studies we distil five (potential) roles that XAI can play in protein research: Evaluator, Multitasker, Engineer, Coach, and Teacher, with the Evaluator role being the only one widely adopted so far. These roles aim to help both protein scientists and model developers understand the possibilities and limitations of implementing XAI for predictive and generative tasks. While our analysis focuses on pLMs, both this categorization and roles are broadly applicable to any other model architectures. We conclude by highlighting critical areas of application for the future, including risks related to security, trustworthiness, and bias, and we call for community benchmarks, open-source tooling, domain-specific visualizations, and wet-lab characterization to advance the interpretability of protein AI.
Citations
- What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models
- Training a Scientific Reasoning Model for Chemistry
- BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
- Towards Interpretable Protein Structure Prediction with Sparse Autoencoders
- Genome modeling and design across all domains of life with Evo 2
- Are protein language models the new universal key?
- InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders
- Concept Bottleneck Language Models For protein design
- Towards evaluations-based safety cases for AI scheming
- Publishing neural networks in drug discovery might compromise training data privacy
- F-Fidelity: A Robust Framework for Faithfulness Evaluation of Explainable AI
- Protein Language Model Fitness Is a Matter of Preference
- Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks
- DataComp-LM: In search of the next generation of training sets for language models
- Accurate structure prediction of biomolecular interactions with AlphaFold 3
- LM Transparency Tool: Interactive Tool for Analyzing Transformer Language Models
- Gradient based Feature Attribution in Explainable AI: A Technical Review
- ProtChatGPT: Towards Understanding Proteins with Large Language Models
- Model Compression and Efficient Inference for Large Language Models: A Survey
- AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
- Universal Neurons in GPT2 Language Models
- Visual Analytics for Generative Transformer Models
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Insights into the inner workings of transformer models for protein function prediction
- A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
- A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations
- Kronfluence: Influence Functions with Eigenvalue-corrected Kronecker-Factored Approximate Curvature
- De novo design of protein structure and function with RFdiffusion
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Sequential Integrated Gradients: a simple but effective method for explaining language models
- LLM-Pruner: On the Structural Pruning of Large Language Models
- AttentionViz: A Global View of Transformer Attention
- ExplainableFold: Understanding AlphaFold Prediction with Explainable AI
- Large language models generate functional protein sequences across diverse families
- Evolutionary-scale prediction of atomic-level protein structure with a language model
- How Much Does Attention Actually Attend? Questioning the Importance of Attention in Pretrained Transformers
- Towards Faithful Model Explanation in NLP: A Survey
- Toy Models of Superposition
- Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks
- OpenXAI: Towards a Transparent Evaluation of Model Explanations
- Exploring evolution-aware & -free protein language models as protein function predictors
- How to Dissect a Muppet: The Structure of Transformer Embedding Spaces
- Visualizing and Explaining Language Models
- XAI for Transformers: Better Explanations through Conservative Propagation
- From Kepler to Newton: Explainable AI for Science
- Acquisition of Chess Knowledge in AlphaZero
- HMD-AMP: Protein Language-Powered Hierarchical Multi-label Deep Forest for Annotating Antimicrobial Peptides
- Protein Folding Neural Networks Are Not Robust
- Discretized Integrated Gradients for Explaining Language Models
- T3-Vis: a visual analytic framework for Training and fine-Tuning Transformers in NLP
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- Highly accurate protein structure prediction with AlphaFold
- Synthetic Benchmarks for Scientific Research in Explainable Machine Learning
- Dodrio: Exploring Transformer Models with Interactive Visualization
- Interpretable Deep Learning: Interpretation, Interpretability, Trustworthiness, and Beyond
- Interpretable deep learning: interpretation, interpretability, trustworthiness, and beyond
- Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models
- Generating Plausible Counterfactual Explanations for Deep Transformers in Financial Text Classification
- Analyzing the Source and Target Contributions to Predictions in Neural Machine Translation
- Analyzing Individual Neurons in Pre-trained Language Models
- Captum: A unified and generic model interpretability library for PyTorch
- The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models
- ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Deep Learning and High Performance Computing
- ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning
- BERTology Meets Biology: Interpreting Attention in Protein Language Models
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- InteractionNet: Modeling and Explaining of Noncovalent Protein-Ligand Interactions with Noncovalent Graph Neural Network and Layer-Wise Relevance Propagation
- ProGen: Language Modeling for Protein Generation
- Explainable Deep Relational Networks for Predicting Compound-Protein Affinities and Contacts
- Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI
- Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
- One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques
- A Multiscale Visualization of Attention in the Transformer Model
- What Does BERT Look At? An Analysis of BERT's Attention
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- Attention is not Explanation
- Sanity Checks for Saliency Maps
- Towards Robust Interpretability with Self-Explaining Neural Networks
- Tell Me Where to Look: Guided Attention Inference Network
- Attention Is All You Need
- A Unified Approach to Interpreting Model Predictions
- Understanding Black-box Predictions via Influence Functions
- Axiomatic Attribution for Deep Networks
- Technical Report on the CleverHans v2.1.0 Adversarial Examples Library
- "Why Should I Trust You?": Explaining the Predictions of Any Classifier
- Explaining and Harnessing Adversarial Examples
- Intriguing properties of neural networks
Cited by
Discussions
Related