Related Experiment Videos
How good is prediction of protein structural class by the component-coupled method?
1National Laboratory of Biomacromolecules, Institute of Biophysics, Academia Sinica, Beijing, Peoples Republic of China. zxwang@sun5.ibp.ac.cn
Proteins
|February 3, 2000
Summary
Predicting protein structural class from amino acid composition has conflicting accuracy reports. This study introduces a new statistical method, revealing that 60% is the maximum achievable accuracy for classifying protein structures.
Area of Science:
- Structural biology
- Bioinformatics
- Computational biology
Background:
- Proteins are classified into four structural classes: all-alpha, all-beta, alpha+beta, and alpha/beta.
- Existing methods for predicting protein structural class from amino acid composition show contradictory success rates, ranging from 60% to nearly 100%.
Purpose of the Study:
- To resolve the paradox of contradictory success rates in protein structural class prediction.
- To determine the upper limit of prediction accuracy for protein structural classes based solely on amino acid composition.
Main Methods:
- A new statistical method based on the normality assumption and Bayes decision rule for minimum error was developed.
- The proposed method was evaluated using a non-redundant dataset of 1,189 protein domains.
Main Results:
- The new method provides optimum predictive results in a statistical sense, assuming normal distributions for the four protein folding classes.
- The study demonstrates that 60% correctness represents the upper limit for predicting the four protein structural classes from amino acid composition alone.
- Previous high accuracy rates (over 90%) were attributed to unrepresentative preselected test sets.
Conclusions:
- The inherent limitation in predicting protein structural class from amino acid composition alone is approximately 60% accuracy.
- The apparent high accuracies reported previously were likely due to biased selection of training and testing datasets.