Related Experiment Video
Updated: May 27, 2025

09:47
Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
940
From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders
Etowah Adams1, Liam Bai2, Minji Lee1
1Department of Systems Biology, Columbia University.
Biorxiv : the Preprint Server for Biology
|February 20, 2025
Summary
Protein language models (pLMs) represent proteins using generic and family-specific features. Analyzing these features with sparse autoencoders (SAEs) can reveal novel biological mechanisms and improve pLM understanding.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning in Biology
Background:
- Protein language models (pLMs) are advanced tools for predicting protein structure and function, trained on vast sequence datasets.
- The internal features or representations learned by pLMs are not fully understood, limiting insights into their predictive power and potential for biological discovery.
- Understanding pLM features could illuminate how these models capture biological information and potentially uncover new protein biology.
Purpose of the Study:
- To investigate the specific features learned by protein language models (pLMs).
- To characterize the nature of these features (generic vs. specific) and their relationship to protein properties.
- To explore the utility of these features for hypothesis generation regarding unknown biological mechanisms.
Main Methods:
- Training sparse autoencoders (SAEs) on the residual stream outputs of the ESM-2 pLM.
- Characterizing the learned SAE features to understand their representation of protein sequence information.
- Employing linear probing on SAE features to identify sequence determinants of protein properties like thermostability and localization.
- Developing visualization tools for interpreting predictive SAE features.
Main Results:
- Demonstrated that pLMs utilize a mix of generic and family-specific features to represent proteins.
- Successfully identified known sequence determinants for protein thermostability and subcellular localization using linear probing of SAE features.
- Highlighted predictive features lacking clear functional associations, suggesting potential roles in undiscovered biological mechanisms.
Conclusions:
- pLM features offer a blend of general and specific biological information.
- SAE analysis provides a method to interpret pLM representations and link them to protein functions.
- This approach enhances understanding of pLM limitations and facilitates the generation of new biological hypotheses.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
10.7K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.7K
Protein Organization
6.2K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.2K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
38
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
38
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K

