Related Experiment Video
Updated: May 6, 2026

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Current status and new features of the Consensus Coding Sequence database.
Catherine M Farrell1, Nuala A O'Leary, Rachel A Harte
1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Building 38A, 8600 Rockville Pike, Bethesda, MD 20894, USA, Center for Biomolecular Science and Engineering, University of California Santa Cruz (UCSC), Santa Cruz, CA 95064, USA, Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge, CB10 1SA, UK and Howard Hughes Medical Institute, University of California Santa Cruz, Santa Cruz, CA 95064, USA.
The Consensus Coding Sequence (CCDS) project provides a high-quality, stable dataset of human and mouse protein-coding regions. Recent updates enhance data accessibility and reporting for improved genomic research.
Area of Science:
- Genomics
- Bioinformatics
Background:
- The Consensus Coding Sequence (CCDS) project is a collaborative effort to maintain a consistent dataset of protein-coding regions.
- It ensures identical annotations between human and mouse reference genomes from NCBI and Ensembl pipelines.
- CCDS IDs provide stable identifiers for these high-quality, quality-assured annotations.
Purpose of the Study:
- To describe the current status and growth of the CCDS dataset.
- To report recent enhancements to the CCDS web and FTP sites.
- To outline future curation targets for the CCDS project.
Main Methods:
- Collaborative review by NCBI, Wellcome Trust Sanger Institute, and UC Santa Cruz.
- Quality assurance tests for identical annotations.
- Development of new search, display, and reporting features on CCDS web and FTP sites.
Main Results:
- The CCDS dataset has experienced recent growth.
- Web and FTP sites now offer more explicit reporting on compared annotation releases.
- New features include improved search/display options and biologically descriptive information.
Conclusions:
- The CCDS project ensures high-quality, consistent genomic annotations.
- Recent updates enhance data usability and transparency for researchers.
- Ongoing curation efforts aim to further improve the CCDS dataset.
More Related Videos
07:38Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
Published on: April 11, 2019
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Cis-regulatory Sequences
Cis-regulatory Sequences
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Genome Annotation and Assembly
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....