Related Experiment Video
Updated: Mar 24, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Evolutionary profile enhancement improves protein function annotation for remote homologs
Shitong Dai1, Jiaqi Luo1, Yunan Luo1
1School of Computational Science and Engineering, Georgia Institute of Technology.
Abstract:
Accurate annotation of protein function is essential for understanding biological processes, yet this remains challenging for proteins lacking characterized homologs or belonging to underrepresented functional classes. Although machine learning approaches have become the gold standard for automated function prediction, they often perform poorly on out-of-distribution samples with low sequence identity to training proteins with known annotations. We propose EPERep, an evolutionary input enhancement strategy that leverages the vast space of unannotated protein sequences to improve the prediction of the functions of underrepresented proteins. Our key insight is that, even if a query protein has insufficient similarity to annotated proteins for direct annotation transfer, a wider range of similar unannotated sequences can be identified to facilitate better representation learning. Inspired by profile-based sequence search methods, EPERep incorporates homologous sequences as contextual input to refine the representations of individual proteins from pre-trained protein language models, effectively constructing a pLM-based profile for each query protein. Across four major annotation benchmarks on EC numbers, structural domains, Pfam families, and Gene Ontology predictions, EPERep consistently outperforms strong ML and sequence-alignment baselines. Gains are most pronounced for proteins from rare functional classes, with few or no labeled homologs, and for sequences exhibiting remote homology to the training distribution. These results demonstrate that evolutionary input enhancement provides a principled and scalable strategy for improving protein function prediction, particularly in long-tail and low-identity regimes.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Evolutionary Relationships through Genome Comparisons
Conservation of Protein Domains
Ribosome Profiling
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Gene Evolution - Fast or Slow?

