Related Experiment Video
Updated: Jul 16, 2026

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
UniRef: comprehensive and non-redundant UniProt reference clusters
Baris E Suzek1, Hongzhan Huang, Peter McGarvey
1Protein Information Resource, Department of Biochemistry and Molecular & Cellular Biology, Georgetown University Medical Center, Washington, DC 20007, USA. bes23@georgetown.edu
UniRef clusters protein sequences, reducing redundancy to speed up similarity searches and improve biological discovery. This organization of protein sequence data enhances analysis across various research fields.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
- Proteomics
Background:
- Redundant protein sequences in databases complicate similarity searches and result interpretation.
- Clustering protein sequences by similarity organizes data, reducing bias and overrepresentation.
Purpose of the Study:
- To present UniRef (UniProt Reference Clusters) as a solution for managing redundant protein sequences.
- To describe the UniRef database structure and its benefits for sequence similarity searches.
Main Methods:
- Clustering of protein sequences from UniProt Knowledgebase (UniProtKB) and UniProt Archive.
- Creation of UniRef100 by merging identical sequences and subfragments.
- Generation of UniRef90 and UniRef50 by clustering UniRef100 at 90% and 50% identity levels, respectively.
Main Results:
- UniRef provides clustered sequence sets at multiple resolutions (UniRef100, UniRef90, UniRef50).
- Database size reduction of 10%, 40%, and 70% for UniRef100, UniRef90, and UniRef50, respectively.
- Improved speed and detection of distant relationships in similarity searches.
- UniRef entries include cluster information, member counts, taxonomy, and links to UniProtKB annotations.
Conclusions:
- UniRef effectively organizes protein sequence space, enhancing search efficiency and biological discovery.
- The database is valuable for applications in genome annotation and proteomics data analysis.
- UniRef is regularly updated and accessible online and via download.
More Related Videos
10:40Comprehensive Workflow for the Genome-wide Identification and Expression Meta-analysis of the ATL E3 Ubiquitin Ligase Gene Family in Grapevine
Published on: December 22, 2017
07:49Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
Related Concept Videos
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Genome Annotation and Assembly
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order to...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order to...