Related Experiment Video
Updated: Aug 15, 2025

Implementation of In Vitro Drug Resistance Assays: Maximizing the Potential for Uncovering Clinically Relevant Resistance Mechanisms
Published on: December 9, 2015
Using machine learning to predict the effects and consequences of mutations in proteins
Daniel J Diaz1, Anastasiya V Kulikova2, Andrew D Ellington3
1Department of Chemistry, The University of Texas at Austin, 105 E 24TH St., Austin, 78712, Texas, USA; Department of Molecular Biosciences, The University of Texas at Austin, 100 East 24th St., Stop A5000, Austin, 78712, Texas, USA. Electronic address: https://twitter.com/aiproteins.
Machine and deep learning models predict protein variants with improved fitness using large datasets. Data quality and availability are crucial, but pre-training methods offer solutions for limited experimental data, driving protein science advancements.
Area of Science:
- Biochemistry
- Computational Biology
- Protein Engineering
Background:
- Massive datasets of protein sequences, structures, and mutational effects are increasingly available.
- Machine and deep learning (ML) models can leverage these datasets for variant prediction.
- Systematic benchmarking highlights data availability and quality as key constraints for ML model performance.
Purpose of the Study:
- To review the application of machine and deep learning in predicting protein variants with improved fitness.
- To discuss the impact of data availability and quality on predictive model performance.
- To explore strategies for leveraging limited experimental data in protein ML models.
Main Methods:
- Review of current machine and deep learning approaches for protein variant prediction.
- Analysis of benchmarking studies on ML model performance.
- Discussion of unsupervised, self-supervised, hybrid, and transfer learning strategies.
Main Results:
- ML and deep learning can predict protein variants with enhanced fitness.
- Data availability and quality are more critical than specific ML algorithm choices.
- Unsupervised/self-supervised pre-training on generic datasets can be effective with subsequent refinement.
- Recent progress in ML for protein science has been substantial.
Conclusions:
- Machine learning approaches are poised to significantly impact future breakthroughs in protein biochemistry and engineering.
- Effective strategies exist for utilizing limited experimental data through pre-training and transfer learning.
- Continued development and application of ML will accelerate protein design and understanding.
Related Concept Videos
Mutations
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Mismatch Repair
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
Mutations in Microorganisms
Protein-protein Interfaces

