Related Experiment Videos
SABmark--a benchmark for sequence alignment that covers the entire known fold space
Ivo Van Walle1, Ignace Lasters, Lode Wyns
1Department of Ultrastructure, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussel, Belgium. ivwalle@vub.ac.be
Bioinformatics (Oxford, England)
|August 31, 2004
Summary
The Sequence Alignment Benchmark (SABmark) offers diverse multiple alignment challenges from the SCOP database. These benchmark sets evaluate sequence alignment tools across varying similarity levels and include tricky, unalignable sequences.
Area of Science:
- Bioinformatics
- Computational Biology
- Structural Biology
Background:
- Accurate protein sequence alignment is crucial for understanding protein function and evolution.
- Existing benchmarks may not adequately represent the full spectrum of sequence similarity encountered in real-world datasets.
- The Structural Classification of Proteins (SCOP) database provides a hierarchical classification of protein folds.
Purpose of the Study:
- To introduce the Sequence Alignment Benchmark (SABmark) as a comprehensive resource for evaluating sequence alignment algorithms.
- To provide benchmark sets covering a wide range of sequence similarities, including challenging low-similarity regions.
- To assess the performance of alignment tools on datasets containing potentially misleading, unalignable sequences.
Main Methods:
- Generation of multiple sequence alignment problems derived from the SCOP classification.
- Creation of two main benchmark sets: Twilight Zone (very low to low similarity) and Superfamilies (low to intermediate similarity).
- Development of alternate versions of these sets by incorporating unalignable but superficially similar sequences.
Main Results:
- SABmark encompasses the entire known protein fold space.
- The benchmark sets offer varying degrees of sequence similarity, from very low to intermediate.
- Alternate sets introduce complexity by including sequences that are difficult to align correctly.
Conclusions:
- SABmark provides a robust and diverse platform for benchmarking protein sequence alignment methods.
- The inclusion of challenging sequences enhances the evaluation of alignment tool sensitivity and specificity.
- This benchmark facilitates the development and refinement of more accurate sequence alignment algorithms.