Related Experiment Videos
Estimating the probability for a protein to have a new fold: A statistical computational model
1Department of Biological Chemistry, Institute of Life Sciences, The Hebrew University, Jerusalem 91904, Israel.
Summary
This study introduces a method to predict proteins with novel folds using sequence and structure classifications. It prioritizes proteins for structural determination, improving the efficiency of structural genomics.
Area of Science:
- Structural biology
- Bioinformatics
- Computational biology
Background:
- Structural genomics aims to determine protein structures but faces cost limitations.
- Identifying proteins with novel folds is crucial for comprehensive structural coverage.
- Existing methods require efficient strategies to prioritize targets for structure determination.
Purpose of the Study:
- To develop a computational method for predicting the probability of a protein having an unsolved fold.
- To identify a prioritized list of proteins for experimental structure determination.
Main Methods:
- Utilized protomap (sequence-based) and scop (structure-based) classifications.
- Constructed a graph representing protein space based on protomap clusters.
- Developed a statistical model using distances between known and predicted protein folds.
Main Results:
- A significantly different distribution of distances was observed between solved and unsolved protein folds.
- Bayes' rule was applied to estimate the probability of a protein having an undetermined fold.
- Predicted probabilities correlated well with newly discovered folds in recent structural data.
Conclusions:
- The developed method effectively predicts proteins likely to possess novel folds.
- This approach enhances the efficiency of structural genomics by prioritizing targets.
- The findings contribute to a more systematic exploration of the protein structure universe.