Evolutionary profiles for protein fitness prediction
Xiaoran Jiao1, Shengdong Lin2, Jigang Fan3
1Computer Science and Technology, Zhejiang University, Hangzhou, 310058, China.
Bioinformatics (Oxford, England)
|July 20, 2026
Summary
We introduce EvoIF, a novel protein language model that predicts mutation fitness by integrating evolutionary profiles and inverse folding signals. This efficient model achieves competitive performance on large-scale benchmarks, advancing protein engineering.
Area of Science:
- Computational Biology
- Protein Engineering
- Bioinformatics
Background:
- Predicting mutation fitness is crucial for protein engineering but limited by assay availability.
- Protein language models (pLMs) show promise in zero-shot fitness prediction.
- Natural evolution can be viewed as reward maximization, with MLM as inverse reinforcement learning.
Purpose of the Study:
- To develop a computationally efficient model for predicting mutation fitness.
- To integrate evolutionary profiles and inverse folding signals for improved prediction accuracy.
- To interpret pLM fitness prediction through the lens of inverse reinforcement learning.
Main Methods:
- Introduced EvoIF, a lightweight model integrating evolutionary profiles and inverse folding (IF) profiles.
- Fused sequence-structure representations with profiles using a compact transition block.
- Utilized calibrated probabilities for log-odds scoring.
Main Results:
- EvoIF achieved competitive performance on the ProteinGym benchmark (>2.5M mutants).
- The model used significantly less training data (0.15%) and fewer parameters than large models.
- Complementary evolutionary and IF profiles enhanced robustness across diverse conditions.
Conclusions:
- EvoIF offers an efficient and effective approach to predicting mutation fitness.
- Integrating evolutionary and inverse folding signals improves model robustness.
- The inverse reinforcement learning framework provides an interpretive lens for pLM fitness prediction.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...


