Related Experiment Video
Updated: Jun 21, 2025

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Fine-tuning protein embeddings for functional similarity evaluation.
Andrew Dickson1, Mohammad R K Mofrad1
1Departments of Bioengineering and Mechanical Engineering, Molecular Cell Biomechanics Laboratory, University of California, Berkeley, CA 94720, United States.
Fine-tuning protein language models improves their embeddings for functional annotation and protein family discovery. These enhanced embeddings maintain interpretability and outperform standard methods in key bioinformatics tasks.
Area of Science:
- Bioinformatics
- Computational Biology
- Protein Science
Background:
- Proteins with unknown functions are often compared to known proteins using sequence similarity or learned embedding spaces.
- Protein sequence embeddings aid in annotation, clustering, and discovering protein families.
- The potential for deliberately designing embeddings to enhance downstream tasks remains underexplored.
Purpose of the Study:
- To investigate whether fine-tuning protein language models can improve the quality and utility of protein embeddings for downstream tasks.
- To assess the impact of fine-tuning on functional annotation and protein family discovery.
Main Methods:
- Fine-tuning pre-trained protein language models using a classification loss on Gene Ontology (GO) terms.
- Evaluating the performance of fine-tuned embeddings in K-nearest neighbor classifiers for GO annotation.
- Assessing the quality of fine-tuned embeddings for protein family discovery using clustering.
Main Results:
- Direct fine-tuning of language models significantly enhances protein embedding quality for functional annotation.
- Fine-tuned embeddings achieve superior performance in GO annotation compared to standard methods, even outperforming directly fine-tuned classifiers.
- The improved embeddings maintain interpretability through protein similarity comparisons and perform well in clustering tasks for protein family rediscovery.
Conclusions:
- Fine-tuning protein language models is an effective strategy to improve protein embedding quality.
- Enhanced embeddings facilitate more accurate functional annotation and protein family discovery.
- This approach offers a powerful tool for bioinformatics research, combining interpretability with improved performance.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Improving Translational Accuracy
Protein Folding Quality Check in the RER
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...

