Related Experiment Video
Updated: Apr 19, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Quality assessment of modeled protein structure using physicochemical properties
Prashant Singh Rana1, Harish Sharma, Mahua Bhattacharya
1Department of Information Communication and Technology, ABV-Indian Institute of Information Technology and Management, Gwalior MP-474015, India.
Physicochemical properties predict protein structure quality using machine learning. Random Forest models accurately estimate Root Mean Square Deviation (RMSD), Template Modeling (TM-score), and Global Distance Test (GDT_TS-score) for modeled proteins.
Area of Science:
- Computational Biology
- Structural Bioinformatics
- Machine Learning
Background:
- Protein structure quality assessment is crucial for distinguishing native structures from predicted ones.
- Physicochemical properties offer valuable insights into protein structure characteristics.
Purpose of the Study:
- To develop a fast and cost-effective method for predicting protein structure quality metrics.
- To evaluate the efficacy of various machine learning methods and physicochemical properties for predicting Root Mean Square Deviation (RMSD), Template Modeling (TM-score), and Global Distance Test (GDT_TS-score).
Main Methods:
- Utilized six physicochemical properties: total surface area, Euclidean distance (ED), total empirical energy, secondary structure penalty (SS), sequence length (SL), and pair number (PN).
- Employed nine machine learning algorithms, with feature importance determined by a Self-adaptive Differential Evolution (SaDE) algorithm.
- Validated model robustness using K-fold cross-validation on a dataset of 95,091 modeled structures from 4896 native targets.
Main Results:
- The Random Forest (RF) model demonstrated superior performance compared to other machine learning methods.
- Achieved high prediction accuracy for RMSD (78.82%), TM-score (86.56%), and GDT_TS-score (87.37%) on the testing dataset.
- Reported low Root Mean Square Error (RMSE) values of 1.20 for RMSD, 0.06 for TM-score, and 0.06 for GDT_TS-score, with strong correlation scores (0.96, 0.92, 0.91 respectively).
Conclusions:
- Physicochemical properties combined with machine learning, particularly Random Forest, provide an efficient approach for protein structure quality assessment.
- This method accelerates and reduces the cost of predicting key protein structure quality indicators.
- The developed models offer reliable predictions for RMSD, TM-score, and GDT_TS-score, aiding in the evaluation of modeled protein structures.
Related Concept Videos
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein Folding
Protein Folding
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Molecular Models
Protein Folding Quality Check in the RER

