Related Experiment Video
Updated: Apr 20, 2026

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches
Baris E Suzek1, Yuqi Wang2, Hongzhan Huang2
1Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, Department of Computer Engineering, Muğla Sıtkı Koçman University, Muğla 48000, Turkey, Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland Protein Information Resource, Georgetown University Medical Center, Washington, DC 20007, USA, Department of Computer Engineering, Muğla Sıtkı Koçman University, Muğla 48000, Turkey, Center for Bioinformatics and Computational Biology and Protein Information Resource, University of Delaware, Newark, DE 19711, USA, European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK and Swiss Institute of Bioinformatics, Centre Medical Universitaire, 1 rue Michel Servet, 1211 Geneva 4, Switzerland.
UniRef clusters, enhanced for non-redundancy, improve protein similarity searches. These clusters offer faster, more sensitive similarity detection and reliable functional annotation, outperforming traditional sequence databases.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- UniRef databases cluster UniProtKB sequences for functional annotation.
- Recent improvements include a sequence length overlap threshold to enhance non-redundancy and homogeneity.
- These updates aim to boost similarity search speed, sensitivity, and annotation consistency.
Purpose of the Study:
- To evaluate the impact of UniRef's recent improvements on similarity searches and functional annotation.
- To assess the speed, sensitivity, and coverage of similarity searches using UniRef clusters.
- To validate the consistency of functional annotation within UniRef clusters.
Main Methods:
- Analysis of Gene Ontology terms to assess intra-cluster molecular function consistency.
- Comparison of BLASTP searches against UniRef50 versus UniProtKB sequences.
- Evaluation of hit list size, search speed, and detection sensitivity for remote similarities.
Main Results:
- UniRef clusters demonstrate high intra-cluster molecular function consistency (>97%).
- BLASTP searches against UniRef50 yield more concise hit lists (∼7x shorter).
- Searches against UniRef50 are faster (∼6x) and more sensitive for remote similarities (>96% recall).
Conclusions:
- UniRef clusters are a reliable and scalable alternative to native sequence databases for similarity searches.
- The enhanced UniRef databases improve the speed and sensitivity of similarity searches.
- UniRef clusters enhance the consistency and reliability of functional annotation.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Conservative Site-specific Recombination and Phase Variation
The recognition sites for Cre recombinase called LoxP...
Evolutionary Relationships through Genome Comparisons

