Related Experiment Videos
Complete fold annotation of the human proteome using a novel structural feature space
Sarah A Middleton1, Joseph Illuminati2, Junhyong Kim1,3
1Genomics and Computational Biology Program, University of Pennsylvania, Philadelphia, PA 19104, USA.
Scientific Reports
|April 14, 2017
Summary
This study introduces a novel machine learning method for protein structural fold recognition, achieving over 94% accuracy. The approach aids in predicting protein functions and identifying novel protein fold families.
Area of Science:
- Computational Biology
- Structural Bioinformatics
- Machine Learning
Background:
- Protein structural fold recognition is crucial for predicting protein structure and function.
- Current methods face challenges with computational demands and identifying novel folds, leaving many proteins unclassified.
Purpose of the Study:
- To develop a new machine learning approach for accurate protein fold recognition.
- To enable the inference of unknown and novel protein folds.
- To apply the method for large-scale human protein domain fold prediction and functional insights.
Main Methods:
- Utilized a novel feature space for machine learning-based protein fold recognition.
- Trained and validated the model on known protein folds, including those with limited training data.
- Applied the method to predict folds for 34,330 human protein domains.
Main Results:
- Achieved over 94% accuracy in recognizing known protein folds.
- Demonstrated high accuracy even with single training examples per fold.
- Predicted human protein domain folds, revealing potential biological functions like RNA-binding ability.
Conclusions:
- The developed machine learning method offers accurate and efficient protein fold recognition.
- This approach can be applied to de novo proteome-wide fold prediction and the discovery of novel fold families.
- Predicted folds provide valuable insights into protein function and biological roles.