Related Experiment Videos
DDBASE2.0: updated domain database with improved identification of structural domains.
A Vinayagam1, J Shi, G Pugalenthi
1National Centre for Biological Sciences, Tata Institute of Fundamental Research, UAS-GKVK campus, Bellary Road, Bangalore, Karnataka 560 065, India.
Bioinformatics (Oxford, England)
|September 27, 2003
Summary
This study introduces an improved, automated method for identifying protein structural domains, achieving high agreement with existing databases. The enhanced approach updates a structural domain database with 5409 newly identified domains from 4592 protein chains.
Area of Science:
- Structural biology
- Bioinformatics
- Computational biology
Background:
- Accurate protein domain identification often requires manual intervention.
- Protein domain definitions are crucial for various analyses, including sequence analysis, structure comparison, and understanding protein folding.
Purpose of the Study:
- To develop an improved, automated method for identifying protein structural domains.
- To reduce the number of discontinuous segments in identified protein domains.
- To update the structural domain database with new domain definitions.
Main Methods:
- Modified domain identification method based on clustering secondary structural elements.
- Comparison of automated domain identification results with crystallographers' definitions and established databases (SCOP, CATH, DALI, 3Dee, PDP).
Main Results:
- The automated method achieved 88% agreement in the number of domains compared to crystallographers' definitions on a test set of 55 proteins.
- Achieved a 98% overlap score in defining domain boundaries with other resources for the same 55 proteins.
- Examined 4592 non-redundant protein chains, identifying 5409 domains and updating the structural domain database.
Conclusions:
- The improved automated method provides accurate protein domain identification.
- The updated structural domain database offers valuable data for further research.
- The method enhances the efficiency and accuracy of protein domain analysis.