Related Experiment Videos
The predictive power of the CluSTr database
Robert Petryszak1, Ernst Kretschmann, Daniela Wieser
1EMBL Outstation Hinxton, The European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridgeshire CB10 1SD, UK.
Bioinformatics (Oxford, England)
|June 18, 2005
Summary
The CluSTr database uses a novel clustering method for protein sequences, significantly improving biological annotation accuracy when combined with existing tools like InterPro. This enhances protein classification and data mining capabilities.
Area of Science:
- Bioinformatics
- Computational Biology
- Protein Science
Background:
- The CluSTr database provides a resource for protein sequence clustering.
- Protein sequence similarity is crucial for understanding protein function and evolution.
Purpose of the Study:
- To evaluate the predictive power and biological relevance of the CluSTr database.
- To assess the utility of CluSTr in automated protein annotation experiments.
Main Methods:
- Utilized a single-linkage hierarchical clustering method based on a protein similarity matrix.
- Computed similarity matrix using all-against-all Smith-Waterman algorithm comparisons.
- Assessed statistical significance of similarity scores via Monte Carlo analysis to derive Z-values.
Main Results:
- Automated annotation experiments demonstrated the predictive power of CluSTr data.
- Combining CluSTr with InterPro in a UniProt data-mining framework significantly increased prediction precision.
- CluSTr data alone showed biological relevance in annotation predictions.
Conclusions:
- The CluSTr clustering approach offers a valuable addition to traditional protein classification methods.
- Integration of CluSTr enhances the precision of automated protein annotation and data mining.