Related Experiment Video
Updated: Mar 11, 2026

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
Uniclust databases of clustered and deeply annotated protein sequences and alignments
Milot Mirdita1, Lars von den Driesch1,2, Clovis Galiez1
1Quantitative and Computational Biology Group, Max Planck Institute for Biophysical Chemistry, Göttingen, Germany.
We introduce Uniclust and Uniboost, novel protein sequence databases and multiple sequence alignment (MSA) resources. These databases offer improved functional annotation consistency and comprehensive domain coverage for protein analysis.
Area of Science:
- Bioinformatics
- Computational Biology
- Structural Biology
Background:
- Protein sequence databases are crucial for analyzing protein function and evolution.
- Existing databases like UniRef have limitations in functional annotation consistency and domain coverage.
- Efficient and sensitive clustering and homology detection tools are needed for large-scale protein sequence analysis.
Purpose of the Study:
- To present the Uniclust and Uniboost databases as a comprehensive resource for protein sequence analysis, function prediction, and sequence searches.
- To improve the consistency of functional annotation compared to existing databases.
- To provide sensitive homology detection and multiple sequence alignment generation.
Main Methods:
- Clustering of UniProtKB sequences at 90%, 50%, and 30% pairwise sequence identity using the MMseqs2 software.
- Annotation of Uniclust sequences with Pfam, SCOP domains, and PDB proteins using the HHblits homology detection tool.
- Construction of Uniboost multiple sequence alignment (MSA) databases by enriching Uniclust30 MSAs with sensitive local sequence matches.
Main Results:
- Uniclust databases (Uniclust90, Uniclust50, Uniclust30) provide improved functional annotation consistency over UniRef90 and UniRef50.
- Uniclust databases contain 17% more Pfam domain annotations than UniProt due to high sensitivity.
- Uniboost databases (Uniboost10, Uniboost20, Uniboost30) offer diverse MSAs for advanced protein analysis.
- The Uniclust server allows keyword searches, exploration of MSAs, taxonomic representation, and annotations.
Conclusions:
- Uniclust and Uniboost represent a significant advancement in protein sequence databases and MSA resources.
- These databases enhance protein analysis, function prediction, and sequence search capabilities.
- The optimized clustering and sensitive homology detection provide a more comprehensive view of protein sequence space and function.
Related Concept Videos
Genome Annotation and Assembly
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks
Protein Organization
The primary structure of a protein is its amino acid sequence....
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...

