Related Experiment Video
Updated: Mar 24, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.7K
Illumina error profiles: resolving fine-scale variation in metagenomic sequencing data
Melanie Schirmer1,2,3, Rosalinda D'Amore4, Umer Z Ijaz5
1The Broad Institute of MIT and Harvard, 415 Main Street, Cambridge, MA 02142, USA. melanie@broadinstitute.org.
BMC Bioinformatics
|March 13, 2016
Summary
Illumina sequencing data contains errors, particularly motif biases ending in "GG", which current quality-based removal methods do not fully address. Understanding these systematic errors is crucial for accurate bacterial genome analysis.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Illumina sequencing platforms are widely used globally, offering high throughput and low costs.
- Rapid advancements in sequencing technology have outpaced the understanding of associated data errors and biases.
- Noise in sequencing data necessitates a deeper investigation into platform-specific error profiles.
Purpose of the Study:
- To systematically investigate errors and biases in Illumina sequencing data.
- To evaluate different Illumina sequencing platforms (Genome Analyzer II, HiSeq, MiSeq) and library preparation methods.
- To identify sequence motifs and nucleotide incorporation biases specific to the sequencing-by-synthesis process.
Main Methods:
- Analysis of the largest collection of in vitro metagenomic datasets to date.
- Position- and nucleotide-specific analysis of sequencing errors.
- Evaluation of quality-score-based error removal strategies.
Main Results:
- A significant bias was identified in motifs (3mers) preceding errors, particularly those ending in 'GG', accounting for approximately 16% of substitution errors.
- Preferential incorporation of ddGTPs was observed, likely due to the engineered polymerase and ddNTPs in sequencing-by-synthesis.
- Quality-score-based error removal eliminated about 69% of substitution errors, but the identified motif bias persisted.
Conclusions:
- Detecting single-nucleotide polymorphism changes in bacterial genomes is vital for understanding phenotypes like antibiotic resistance and virulence.
- Existing error removal techniques are insufficient for addressing Illumina-specific biases, potentially impacting downstream analyses and conclusions.
- Distinguishing systematic sequencing errors from true genetic variation is essential for developing accurate diagnostic and therapeutic strategies.

