Related Experiment Videos
The CATH extended protein-family database: providing structural annotations for genome sequences.
Frances M G Pearl1, David Lee, James E Bray
1Department of Biochemistry and Molecular Biology, University College London, University of London, London WC1E 6BT, UK. frances@biochem.ucl.ac.uk
Protein Science : a Publication of the Protein Society
|January 16, 2002
Summary
A new protocol, DomainFinder, automatically integrates gene sequences into structural families within the CATH database. This significantly expands the protein family database, aiding in structural and evolutionary relationship analysis.
Area of Science:
- Bioinformatics
- Structural Biology
- Genomics
Background:
- Integrating vast gene sequence data into structural families is crucial for understanding protein function and evolution.
- Existing methods may lack the automation and reliability needed for comprehensive database expansion.
Purpose of the Study:
- To develop an automated protocol (DomainFinder) for reliable integration of gene sequences into the CATH domain database.
- To expand the CATH protein family database (CATH-PFDB) with new sequences and identify putative homologous relationships.
Main Methods:
- Utilized PSI-BLAST and IMPALA with conservative thresholds for sequence analysis.
- Developed DomainFinder to assign gene sequences to CATH homologous superfamilies based on PSI-BLAST relationships.
- Generated IMPALA/PSI-BLAST profiles for sequence families and created a web server for new sequence scanning.
Main Results:
- Expanded the CATH-PFDB from 19,563 to 176,597 domain sequences.
- Identified an additional 50,000 putative homologous relationships using less stringent cut-offs.
- Analysis revealed only 15% of sequence families are close enough to known structures for reliable homology modeling.
Conclusions:
- DomainFinder provides a reliable method for expanding protein domain databases.
- The expanded CATH-PFDB and associated profiles facilitate the assignment of new sequences to structural families and superfamilies.
- The findings highlight the need for further structural data to improve homology modeling for a larger proportion of protein families.