Related Experiment Video
Updated: Apr 6, 2026

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
AlignBucket: a tool to speed up 'all-against-all' protein sequence alignments optimizing length constraints
Giuseppe Profiti1, Piero Fariselli2, Rita Casadio3
1Department of Computer Science and Engineering, via Mura Anteo Zamboni 7, Bologna, Bologna Biocomputing group, via S. Giacomo 9/2, Bologna and Health Sciences and Technologies ICIR, via Tolara di Sopra 41/E, Ozzano dell'Emilia, Italy.
This study introduces an optimized algorithm for partitioning large biological sequence databases. This method significantly speeds up sequence comparison, achieving a 5-fold improvement in real-world applications.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Next-generation sequencing necessitates efficient biological sequence annotation.
- All-against-all sequence comparison is crucial for annotation but computationally intensive.
- Large sequence datasets require pre-processing to reduce comparison time.
Purpose of the Study:
- To develop a method for optimizing the partitioning of large sequence databases.
- To reduce the computational time required for all-against-all sequence comparisons.
- To improve the efficiency of biological sequence annotation.
Main Methods:
- An algorithm was developed to partition sequences into sets based on constrained length values.
- Sequences are optimally grouped by length for targeted all-against-all alignment.
- A mathematically optimal partitioning solution is described.
Main Results:
- The proposed algorithm optimizes sequence partitioning for efficient comparison.
- The method achieves a 5-fold speed-up in real-world sequence comparison tasks.
- This optimization aids in faster annotation of biological sequences.
Conclusions:
- The developed partitioning algorithm offers a significant speed-up for sequence comparison.
- This approach enhances the efficiency of annotating large biological sequence datasets.
- The software is publicly available for use in bioinformatics research.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Conservation of Protein Domains
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Evolutionary Relationships through Genome Comparisons

