Data set and fitting dependencies when estimating protein mutant stability: Toward simple, balanced, and
Kristoffer T Baek1, Kasper P Kepp1
1DTU Chemistry, Technical University of Denmark, Lyngby, Denmark.
Abstract:
Accurate prediction of protein stability changes upon mutation (ΔΔG) is increasingly important to evolution studies, protein engineering, and screening of disease-causing gene variants but is challenged by biases in training data. We investigated 45 linear regression models trained on data sets that account systematically for destabilization bias and mutation-type bias BM . The models were externally validated on three test data sets probing different pathologies and for internal consistency (symmetry and neutrality). Model structure and performance substantially depended on training data and even fitting method. We developed two final models: SimBa-IB for typical natural mutations and SimBa-SYM for situations where stabilizing and destabilizing mutations occur to a similar extent. SimBa-SYM, despite is simplicity, is essentially non-biased (vs. the Ssym data set) while still performing well for all data sets (R ~ 0.46-0.54, MAE = 1.16-1.24 kcal/mol). The simple models provide advantage in terms of interpretability, use and future improvement, and are freely available on GitHub.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Physiological Pharmacokinetic Models: Assumption with Protein Binding
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
The Equilibrium Binding Constant and Binding Strength
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Covalently Linked Protein Regulators
These groups modify specific amino acids in a protein....


