Predicting protein thermostability changes from sequence upon multiple mutations
Ludovica Montanucci1, Piero Fariselli, Pier Luigi Martelli
1Department of Biology, University of Bologna, via Irnerio 42, 40126 Bologna, Italy.
Bioinformatics (Oxford, England)
|July 1, 2008
Summary
This study introduces a computational method to predict if protein mutations enhance thermostability. The new support vector machine predictor achieves high accuracy in identifying thermostable protein mutants.
Area of Science:
- Protein Science
- Biotechnology
- Computational Biology
Background:
- Understanding how mutations affect protein thermostability is crucial for protein engineering.
- Existing experimental methods for assessing thermostability changes are often serendipitous.
- A computational approach is needed to predict enhanced protein thermostability.
Purpose of the Study:
- To develop a computational method for predicting whether mutations can enhance protein thermostability.
- To provide a tool for identifying protein mutants with improved thermal stability.
Main Methods:
- A support vector machine (SVM) based computational method was developed.
- The method was trained and tested on a redundancy-reduced dataset of protein sequences and mutations.
- The predictor handles various mutation types, including insertions and deletions.
Main Results:
- The predictor achieved 88% accuracy and a correlation coefficient of 0.75 on a test dataset.
- It correctly classified 12 out of 14 experimentally verified protein mutants with enhanced thermostability.
- The predictor successfully identified all 11 mutated proteins with a stability temperature increase greater than 10 degrees C.
Conclusions:
- The developed SVM-based method is effective in predicting enhanced protein thermostability.
- This computational tool can aid in the engineering of thermostable enzymes.
- The predictor demonstrates high accuracy in identifying beneficial mutations for protein thermal stability.
Related Concept Videos
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
Protein Denaturation
The function of proteins depends on their native three-dimensional structure, which is dictated by the amino acid sequence of the specific protein. Folding of the polypeptide chain takes place under specific conditions that energetically favor the folded conformation. In contrast, protein denaturation occurs spontaneously under unfavorable conditions that disrupt the integrity of the folded conformation. Thus, the chemical and physical environment of a protein, such as significant changes in pH...
Mutations
Overview
Mutations
Mutations are changes in the sequence of DNA. These changes can occur spontaneously or they can be induced by exposure to environmental factors. Mutations can be characterized in a number of different ways: whether and how they alter the amino acid sequence of the protein, whether they occur over a small or large area of DNA, and whether they occur in somatic cells or germline cells.
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
Chromosomal Alterations Are Large-Scale Mutations
While point mutations are changes in a single nucleotide in...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein Folding
Overview

