Related Experiment Video
Updated: Jun 4, 2026

07:49
Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
Protein sequence redundancy reduction: comparison of various method
Bioinformation
|March 3, 2011
Summary
Creating non-redundant protein datasets is crucial in bioinformatics. This study compares popular tools, finding moderate similarity between different programs, advising the use of multiple applications for robust results.
Area of Science:
- Bioinformatics
- Computational Biology
- Proteomics
Background:
- Non-redundant protein datasets are essential for various bioinformatics applications, requiring the removal of highly similar sequences.
- Several computational tools exist for generating non-redundant protein datasets, including 'Decrease redundancy', 'cd-hit', 'Pisces', 'BlastClust', and 'SkipRedundant'.
- The degree of similarity between non-redundant datasets produced by different tools remains an important consideration.
Purpose of the Study:
- To systematically compare the outputs of different non-redundant protein dataset generation programs.
- To assess the extent of similarity between non-redundant datasets created by various bioinformatics tools using identical input data and similarity thresholds.
Main Methods:
- Utilized subsets of the UniProt database as input for multiple redundancy reduction programs.
- Performed a systematic comparison of the features and outputs generated by 'Decrease redundancy', 'cd-hit', 'Pisces', 'BlastClust', and 'SkipRedundant'.
- Analyzed the overlap and similarity levels between the resulting non-redundant datasets based on varying identity thresholds.
Main Results:
- High overlap was observed between non-redundant datasets generated by the same program with different identity thresholds.
- Moderate similarity was found between non-redundant datasets produced by different programs using the same identity threshold.
- The choice of program significantly influences the composition of the final non-redundant protein dataset.
Conclusions:
- Non-redundant protein dataset generation tools exhibit variability in their output.
- Users should be aware of potential differences in datasets produced by distinct software.
- Employing multiple computational applications is recommended to ensure comprehensive and reliable non-redundant protein datasets.
More Related Videos
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Multi-species Conserved Sequences
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Gene Duplication and Divergence
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are characterized.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are characterized.
The Central Dogma
The central dogma explains the flow of genetic information from DNA nucleotides to the amino acid sequence of proteins.
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...

