Related Experiment Video
Updated: Apr 18, 2026

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
15.9K
Shifted Hamming distance: a fast and accurate SIMD-friendly filter to accelerate alignment verification in read
Hongyi Xin1, John Greth1, John Emmons1
1Computer Science Department, Department of Electrical and Computer Engineering, Computational Biology Department, Carnegie Mellon University, Pittsburgh, PA 15213, USA and Department of Computer Engineering, Bilkent University, Bilkent, Ankara 06800, Turkey.
Bioinformatics (Oxford, England)
|January 12, 2015
Summary
A new algorithm, Shifted Hamming Distance (SHD), efficiently filters out error-abundant DNA sequence pairs. This speeds up read mapping by removing computationally intensive comparisons for sequences with too many differences.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Seed-and-extend mappers are crucial for comparing billions of DNA sequences.
- Accurate alignment requires small edit-distances, but many pairs have excessive errors, wasting computational resources.
- Efficiently filtering error-abundant pairs is essential for improving read mapper performance.
Purpose of the Study:
- To develop a fast and accurate algorithm for filtering error-abundant DNA sequence pairs.
- To accelerate the alignment verification process in read mapping.
Main Methods:
- Developed the Shifted Hamming Distance (SHD) algorithm.
- Utilized bit-parallel and SIMD-parallel operations for efficient computation.
- Implemented a user-defined threshold for comprehensive filtering.
Main Results:
- SHD achieves a 3-fold speedup compared to Gene Myers's bit-vector algorithm.
- The algorithm maintains high accuracy for error thresholds up to 5% of sequence length.
- SHD is compatible with existing sequence alignment verification methods.
Conclusions:
- SHD provides a significant performance improvement for read mapping by rapidly filtering irrelevant sequence pairs.
- The algorithm offers a practical solution for handling large-scale DNA sequence comparisons with high error rates.

