Related Experiment Video
Updated: Jun 23, 2026

Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
Fitness translocation: improving variant effect prediction with biologically-grounded data augmentation.
Adrien Mialland1, Shuzo Fukunaga2, Riku Katsuki3
1Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology (AIST), Tokyo, Japan.
Fitness translocation enhances protein variant effect prediction by creating synthetic data from homologous proteins. This data augmentation strategy improves model accuracy, especially with limited training data.
Area of Science:
- Biochemistry and Molecular Biology
- Computational Biology
- Protein Engineering
Background:
- Protein fitness landscape characterization and variant effect prediction are hindered by data scarcity.
- Existing methods struggle with limited datasets, impacting protein engineering advancements.
Purpose of the Study:
- To introduce fitness translocation, a novel data augmentation strategy for protein variant effect prediction.
- To address data scarcity challenges in characterizing protein fitness landscapes.
Main Methods:
- Fitness translocation leverages variant fitness data from homologous proteins to generate synthetic variants for a target protein.
- Protein language model embeddings are used to compute variant differences and create synthetic variants in embedding space.
- The method involves applying fitness offsets from homolog variants to the target wild-type embedding.
Main Results:
- Fitness translocation consistently improves predictive performance across different protein families (IGPS, GFP, SARS-CoV-2 spike proteins).
- The strategy is particularly effective in low-data regimes, enhancing variant effect prediction accuracy.
- The method demonstrates efficacy even when augmenting with remote homologs (as low as 35% sequence identity).
Conclusions:
- Fitness translocation expands and diversifies protein fitness landscapes through biologically grounded data augmentation.
- This approach supports more data-efficient protein engineering by improving variant effect prediction models.
- The study highlights the potential of leveraging homologous protein data for enhanced predictive modeling.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Overview of Transposition and Recombination
Transduction
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
Mutation, Gene Flow, and Genetic Drift
