Related Experiment Video
Updated: Jul 4, 2025

Super-resolution Imaging of Neuronal Dense-core Vesicles
Published on: July 2, 2014
Interpretable feature extraction and dimensionality reduction in ESM2 for protein localization prediction
Zeyu Luo1, Rui Wang1, Yawen Sun1
1Chongqing Key Laboratory of Vector Insects, Chongqing Key Laboratory of Animal Biology, College of Life Science, Chongqing Normal University, Chongqing 401331, China.
Large language models (LLMs) are revolutionizing biological predictions. This study explores how different feature extraction strategies from ESM2 models can improve understanding of protein subcellular localization.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning in Biology
Background:
- Large language models (LLMs) are increasingly applied to biological predictions, particularly for amino acid sequence feature representation.
- Previous research has focused on LLM architecture and fine-tuning, with less attention paid to the nature of extracted features.
- Subcellular localization prediction is a key downstream task where LLM features show promise.
Purpose of the Study:
- To investigate diverse feature extraction strategies from the ESM2 model for biological sequence analysis.
- To understand the relationship between extracted features and specific subcellular localization.
- To evaluate the impact of feature inputs on prediction performance and interpretability.
Main Methods:
- Proposed and evaluated different ESM2 representation extraction strategies, considering sequence character type and position.
- Employed dimensionality reduction, predictive analysis, and interpretability techniques to analyze feature associations.
- Assessed prediction performance and interpretability robustness using Random Forest and Deep Neural Networks with varied feature inputs.
Main Results:
- Identified associations between specific feature types and subcellular localizations, such as N-terminal preference for Mitochondrion and Golgi apparatus.
- Demonstrated that phosphorylation site-based features can reflect phosphorylation properties.
- Showcased varied prediction performance and interpretability based on different feature inputs.
Conclusions:
- Novel insights into maximizing the utility of LLMs for biological domain knowledge extraction.
- Enhanced understanding of the mechanisms underlying LLM-based biological predictions.
- Provided accessible code and APIs for feature extraction and further research.
More Related Videos
11:06Multi-color Localization Microscopy of Single Membrane Proteins in Organelles of Live Mammalian Cells
Published on: June 30, 2018
14:58Identification of Protein Complexes in Escherichia coli using Sequential Peptide Affinity Purification in Combination with Tandem Mass Spectrometry
Published on: November 12, 2012