Related Experiment Videos
NCBI's Conserved Domain Database and Tools for Protein Domain Analysis.
Mingzhang Yang1, Myra K Derbyshire1, Roxanne A Yamashita1
1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland.
Current Protocols in Bioinformatics
|December 19, 2019
Summary
The Conserved Domain Database (CDD) provides over 57,000 protein domain models for sequence annotation. This resource aids researchers in identifying protein footprints, functional sites, and motifs for biological discovery.
Area of Science:
- Bioinformatics
- Computational Biology
- Molecular Biology
Background:
- The Conserved Domain Database (CDD) is a vital resource for annotating protein sequences.
- It houses curated and imported protein domain and family models.
Purpose of the Study:
- To describe the Conserved Domain Database (CDD) and its latest release (v3.17).
- To detail methods for accessing and utilizing CDD annotations for protein sequence analysis.
Main Methods:
- Utilizing the Conserved Domain Database (CDD) for protein domain footprint annotation.
- Employing CD-Search, Batch CD-Search, and standalone RPS-BLAST with rpsbproc for sequence analysis.
Main Results:
- The latest CDD release (v3.17) contains over 57,000 domain models, with nearly 15,000 curated by CDD staff.
- CDD curation enhances coverage and classification of protein domain families.
- Live search and archived annotations are available for NCBI Entrez protein sequences.
Conclusions:
- The CDD is a comprehensive, freely available resource for protein sequence annotation.
- Multiple protocols facilitate efficient retrieval and computation of domain annotations for research.