Related Experiment Video
Updated: Sep 22, 2026

Protein WISDOM: A Workbench for In silico De novo Design of BioMolecules
Published on: July 25, 2013
Mechanistic Interpretability of Fine-Tuned Protein Language Models for Nanobody Thermostability Prediction
Taihei Murakami1,2, Yuki Hashidate1, Yasuhiro Matsunaga1,3
1Graduate School of Science and Engineering, Saitama University, Saitama 338-8570, Japan.
Motivation:
While Protein Language Models (PLMs) fine-tuned on biophysical data achieve high predictive accuracy, the physical principles underlying their predictions remain obscure. Deciphering these representations offers a unique opportunity to not only interpret model decisions but also to discover novel biophysical insights governing protein properties. Here, we present a framework using Sparse Autoencoders (SAEs) to extract mechanistic knowledge from PLMs fine-tuned for nanobody thermostability.
Results:
We fine-tuned the ESM-2 model on the nanobody thermostability dataset, achieving superior performance compared to significantly larger state-of-the-art models. SAE analysis successfully decomposed the model's dense embeddings into sparse, interpretable features without loss of predictive accuracy. We characterized these features through both global and local analyses. Global analysis provided an aggregate map of position-dependent feature contributions, whereas local analysis identified specific residue-level patterns, including known determinants such as the VHH-tetrad and critical disulfide bonds, as well as candidate stabilizing residues. Free Energy Perturbation calculations supported the structural plausibility of selected residue-level hypotheses. These results show that SAE-based interpretation can generate testable, structurally grounded hypotheses for rational protein engineering.
Availability:
The data and source code of the proposed method are available at GitHub (https://github.com/matsunagalab/paper_nanobody-thermostability-sae) and Zenodo (DOI: 10.5281/zenodo.18012027).
Supplementary Information:
Supplementary data are available at Bioinformatics online.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Cooperative Allosteric Transitions
Covalently Linked Protein Regulators
These groups modify specific amino acids in a protein.
Improving Translational Accuracy
Termination of Translation
Leaky Scanning

