Related Experiment Videos
The COG database: a tool for genome-scale analysis of protein functions and evolution
R L Tatusov1, M Y Galperin, D A Natale
1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA.
Nucleic Acids Research
|December 11, 1999
Summary
The Clusters of Orthologous Groups (COGs) database provides a phylogenetic classification for proteins from 21 genomes. This resource aids in understanding protein function and evolution across diverse organisms.
Area of Science:
- Genomics
- Bioinformatics
- Evolutionary Biology
Background:
- Genome sequencing generates vast amounts of protein data.
- Functional and evolutionary interpretation of this data requires systematic classification.
- Existing classification methods may not fully capture orthologous relationships across diverse species.
Purpose of the Study:
- To develop a phylogenetic classification system for proteins encoded in complete genomes.
- To create a comprehensive database of Clusters of Orthologous Groups (COGs).
- To provide a tool for functional and phylogenetic annotation of newly sequenced genomes.
Main Methods:
- Exhaustive comparison of all protein sequences from 21 complete bacterial, archaeal, and eukaryotic genomes.
- Application of the criterion of consistency of genome-specific best hits for COG construction.
- Development of the COGNITOR program for protein classification and annotation.
Main Results:
- Construction of a database comprising 2091 COGs.
- Inclusion of 56-83% of gene products from bacterial and archaeal genomes.
- Inclusion of approximately 35% of gene products from the yeast Saccharomyces cerevisiae genome.
Conclusions:
- The COGs database offers a robust framework for the phylogenetic classification of proteins.
- This classification is essential for maximizing the utility of genome sequences for functional and evolutionary studies.
- The COGNITOR program facilitates the annotation of new proteins and genomes within the COG system.