Related Experiment Video
Updated: May 21, 2026

07:50
Chromatin Immunoprecipitation of Murine Brown Adipose Tissue
Published on: November 21, 2018
How do alignment programs perform on sequencing data with varying qualities and from repetitive regions?
Xiaoqing Yu1, Kishore Guda, Joseph Willis
1Department of Epidemiology and Biostatistics, Case Western Reserve University, Cleveland, OH, 44106, USA. ssun5211@yahoo.com.
Biodata Mining
|June 20, 2012
Summary
Evaluating next-generation sequencing alignment algorithms reveals that data quality significantly impacts performance. Trimming low-quality reads improves alignment accuracy and consistency across tools like SOAP2, Bowtie, BWA, and Novoalign, especially for repetitive genomic regions.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Next-generation sequencing (NGS) generates numerous short reads for biological research.
- Sequencing reads often exhibit low quality at the 3' end and originate from repetitive genomic regions.
- Performance variations of alignment algorithms under these conditions are not well understood.
Purpose of the Study:
- To systematically evaluate the performance of four common alignment algorithms: SOAP2, Bowtie, BWA, and Novoalign.
- To investigate the impact of read quality and repetitive regions on alignment accuracy.
- To compare algorithm concordance using real and simulated sequencing data.
Main Methods:
- Utilized real and simulated next-generation sequencing data with varying quality and from repetitive regions.
- Assessed alignment algorithm performance based on concordance between aligners (real data) and alignment accuracy (simulated data).
- Evaluated the effect of trimming low-quality bases on alignment outcomes.
Main Results:
- All four alignment programs (SOAP2, Bowtie, BWA, Novoalign) performed similarly on high-quality or trimmed data.
- Trimming low-quality ends significantly increased aligned read counts and inter-aligner consistency, particularly for low-quality data.
- Novoalign showed increased sensitivity to data quality improvements; trimming enhanced its concordance with other aligners.
- Simulated data indicated that reads from repetitive regions are prone to incorrect alignment, and filtering multi-hit reads improves accuracy.
Conclusions:
- This study offers a comprehensive comparison of popular alignment algorithms for sequencing data with quality variations and from repetitive regions.
- The evaluation methodology is adaptable for diverse sequencing datasets and platforms.
- The approach can be extended to assess the performance of other alignment programs.
Related Concept Videos
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Sanger Sequencing
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...

