Related Experiment Video
Updated: Aug 8, 2026

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens
Published on: June 28, 2018
Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences
1Burnham Institute for Medical Research La Jolla, CA 92037, USA. liwz@sdsc.edu
We present new ultrafast sequence clustering tools: cd-hit-2d, cd-hit-est, and cd-hit-est-2d. These programs efficiently handle massive protein and nucleotide datasets, significantly outperforming existing methods like BLAST.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
- Proteomics
Background:
- The cd-hit program, introduced in 2001-2002, provides ultrafast protein sequence clustering for large databases.
- The underlying algorithm's efficiency extends beyond protein clustering applications.
- Existing sequence comparison tools like BLAST can be computationally intensive for massive datasets.
Purpose of the Study:
- To introduce novel sequence analysis tools based on the cd-hit algorithm.
- To extend the utility of the cd-hit algorithm to comparative analysis of two datasets and nucleotide sequences.
- To provide highly efficient solutions for clustering and comparing large-scale biological sequence data.
Main Methods:
- Development of new programs: cd-hit-2d for comparing two protein datasets, cd-hit-est for nucleotide sequence clustering, and cd-hit-est-2d for comparing two nucleotide datasets.
- Leveraging the established ultrafast algorithm from the original cd-hit program.
- Designed to handle datasets containing millions of sequences.
Main Results:
- The new cd-hit suite (cd-hit-2d, cd-hit-est, cd-hit-est-2d) demonstrates high efficiency in handling large protein and nucleotide sequence datasets.
- These tools offer significant speed advantages, being hundreds of times faster than traditional methods like BLAST for similar tasks.
- The programs successfully perform comparative analysis between two datasets and clustering of large sequence databases.
Conclusions:
- The cd-hit algorithm's underlying principles can be effectively applied to develop versatile and high-performance tools for biological sequence analysis.
- The newly developed cd-hit programs offer substantial computational savings for researchers working with massive sequence data.
- These tools represent a significant advancement in efficient database searching and sequence comparison within bioinformatics.
More Related Videos
12:23In Vitro Selection of Aptamers to Differentiate Infectious from Non-Infectious Viruses
Published on: September 7, 2022
09:06High-throughput Identification of Gene Regulatory Sequences Using Next-generation Sequencing of Circular Chromosome Conformation Capture (4C-seq)
Published on: October 5, 2018
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...