Protein stability models fail to capture epistatic interactions of double point mutations
Henry Dieckhaus1,2, Brian Kuhlman1,3,4
1Department of Biochemistry and Biophysics, University of North Carolina School of Medicine, Chapel Hill, North Carolina, USA.
Protein Science : a Publication of the Protein Society
|December 20, 2024
Summary
Predicting protein stability changes from double mutations is challenging. Additive models perform surprisingly well, but current AI and physics-based models struggle with epistatic interactions.
Area of Science:
- Biochemistry and Molecular Biology
- Computational Biology
- Protein Engineering
Background:
- Accurate prediction of protein stability changes from mutations is crucial for therapeutics and understanding disease.
- Recent advances focus on single-point mutations, with less attention on double point mutations.
- Understanding double mutations is vital as they can significantly alter protein function and stability.
Purpose of the Study:
- To analyze the largest available dataset of double point mutation stability.
- To benchmark existing protein stability prediction models on double mutant data.
- To investigate the performance of additive versus non-additive models, including AI and physics-based approaches, in capturing epistatic interactions.
Main Methods:
- Analysis of the largest available double point mutation stability dataset.
- Benchmarking of multiple protein stability prediction models (AI-based, physics-based, additive, non-additive).
- Development of an extension to the ThermoMPNN framework and a novel data augmentation scheme.
Main Results:
- Additive models demonstrate surprisingly strong performance, comparable to non-additive models for predicting double mutant stability.
- Current AI and physics-based models do not consistently capture epistatic interactions between mutations.
- Epistasis-aware models show marginal improvement over additive models for predicting stabilizing double mutations.
Conclusions:
- Current protein stability models have limitations in predicting effects of concurrent mutations due to dataset constraints and model sensitivity.
- Additive models offer a robust baseline for double mutant stability prediction.
- Further development is needed to accurately model complex epistatic interactions in protein engineering and disease research.
Related Concept Videos
Epistasis Analysis
4.9K
Although Mendel chose seven unrelated traits in peas to study gene segregation, most traits involve multiple gene interactions that create a spectrum of phenotypes. When the interaction of various genes or alleles at different locations influences a phenotype, this is called epistasis. Epistasis often involves one gene masking or interfering with the expression of another (antagonistic epistasis). Epistasis often occurs when different genes are part of the same biochemical pathway. The...
4.9K
Epistasis
45.8K
In addition to multiple alleles at the same locus influencing traits, numerous genes or alleles at different locations may interact and influence phenotypes in a phenomenon called epistasis. For example, rabbit fur can be black or brown depending on whether the animal is homozygous dominant or heterozygous at a TYRP1 locus. However, if the rabbit is also homozygous recessive at a locus on the tyrosinase gene (TYR), it will have an unshaded coat that appears white, regardless of its TYRP1...
45.8K
Mutations
79.7K
Overview
79.7K
Mismatch Repair
4.8K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
4.8K
Genome Copying Errors
4.1K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.1K
Covalently Linked Protein Regulators
6.8K
Proteins can undergo many types of post-translational modifications, often in response to changes in their environment. These modifications play an important role in the function and stability of these proteins. Covalently linked molecules include functional groups, such as methyl, acetyl, and phosphate groups, and also small proteins, such as ubiquitin. There are around 200 different types of covalent regulators that have been identified.
These groups modify specific amino acids in a protein....
These groups modify specific amino acids in a protein....
6.8K


