Related Experiment Video
Updated: Jun 6, 2026

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
CDD: a Conserved Domain Database for the functional annotation of proteins
Aron Marchler-Bauer1, Shennan Lu, John B Anderson
1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bldg 38 A, Room 8N805, 8600 Rockville Pike, Bethesda, MD 20894, USA. bauer@ncbi.nlm.nih.gov
The Conserved Domain Database (CDD) annotates protein sequences with functional sites using curated models and 3D structures. It simplifies protein analysis by clustering domains into superfamilies for efficient annotation.
Area of Science:
- Bioinformatics
- Structural Biology
- Computational Biology
Background:
- Protein sequence annotation is crucial for understanding biological function.
- Conserved Domain Database (CDD) at NCBI provides annotations for protein sequences.
- Existing methods require efficient ways to manage and utilize domain information.
Purpose of the Study:
- To describe the NCBI's Conserved Domain Database (CDD) resource.
- To highlight CDD's capabilities in annotating protein sequences with conserved domain footprints and functional sites.
- To introduce the Batch CD-Search tool for large-scale annotation.
Main Methods:
- Utilizes manually curated domain models incorporating protein 3D structure.
- Hierarchically organizes related domain families.
- Clusters redundant and homologous models into superfamilies for simplified annotation.
Main Results:
- CDD provides accurate protein annotation by integrating sequence, structure, and function information.
- Superfamily clustering simplifies the assignment of domain footprints.
- Batch CD-Search enables efficient, large-scale annotation of protein datasets.
Conclusions:
- The Conserved Domain Database is a valuable resource for protein sequence annotation.
- Integration of 3D structure refines domain models and enhances functional inference.
- CDD and its Batch CD-Search tool streamline protein analysis workflows.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Genome Annotation and Assembly
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...

