Related Experiment Video
Updated: Jan 15, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Data-efficient protein mutational effect prediction with weak supervision by molecular simulation and protein
Teppei Deguchi1,2, Nur Syatila Ab Ghani3, Yoichi Kurumida3
1Graduate School of Frontier Sciences, The University of Tokyo, 5-1-5, Kashiwanoha, Kashiwa, Chiba 277-0882, Japan.
This study introduces a novel data augmentation method for machine learning models predicting protein mutations. It enhances prediction accuracy in protein engineering and pathogenicity analysis, especially with limited experimental data.
Area of Science:
- Computational Biology
- Biophysics
- Machine Learning
Background:
- Protein mutational effect prediction is vital for protein engineering and pathogenicity assessment.
- Current methods face challenges due to limited experimental data and high costs.
- Previous data augmentation relied on molecular simulations, limited to thermostability.
Purpose of the Study:
- To develop a new data augmentation technique for protein mutational effect prediction.
- To extend the applicability of computational predictions to diverse protein properties.
- To improve prediction accuracy in low-data regimes.
Main Methods:
- Combined molecular simulation with zero-shot prediction from protein language models for data augmentation.
- Utilized computational estimates as 'weak' training data.
- Dynamically adjusted the weight and inclusion of weak data based on experimental data availability.
Main Results:
- The new method enhances prediction accuracy, particularly when experimental data is scarce.
- Successfully extended applicability to protein binding affinity and enzymatic activity prediction.
- Demonstrated improved performance in benchmark tests for small data scenarios.
Conclusions:
- The proposed method effectively supplements experimental data for machine learning models.
- Offers a powerful approach for protein engineering and pathogenicity prediction with limited data.
- Advances the field of protein mutational effect prediction in data-scarce environments.
More Related Videos
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
11:36A Protocol for Functional Assessment of Whole-Protein Saturation Mutagenesis Libraries Utilizing High-Throughput Sequencing
Published on: July 3, 2016
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Covalently Linked Protein Regulators
These groups modify specific amino acids in a protein....
Improving Translational Accuracy