Sources of erroneous sequences and artifact chimeric reads in next generation sequencing of genomic DNA from

Simon Haile1, Richard D Corbett1, Steve Bilobram1

  • 1Canada's Michael Smith Genome Sciences Centre, BC Cancer, Vancouver, British Columbia, Canada.

Nucleic Acids Research
|November 13, 2018
PubMed

Insights

Formalin-fixed, paraffin-embedded (FFPE) samples can cause strand-split artifact reads (SSARs) during next-generation sequencing (NGS). A novel method using S1 nuclease effectively reduces these artifacts, improving sequencing data quality.

Area of Science:

  • Pathology
  • Genomics
  • Molecular Biology

Background:

  • Formalin-fixed, paraffin-embedded (FFPE) tissues are standard in pathology labs but pose challenges for next-generation sequencing (NGS).
  • Nucleic acid damage during FFPE processing can lead to sequencing artifacts, including chimeric reads.
  • Strand-split artifact reads (SSARs), which align to both DNA strands, are a significant artifact class in FFPE-derived NGS data.

Purpose of the Study:

  • To elucidate the mechanistic basis of SSARs and other chimeric artifacts in FFPE NGS data.
  • To develop and validate methods for reducing SSARs and associated sequencing artifacts.
  • To establish an analytical approach for quantifying SSARs in NGS datasets.

Main Methods:

  • Investigated the alignment patterns of chimeric reads in NGS data from FFPE samples.
  • Developed a conceptual framework for SSAR genesis.
  • Applied S1 nuclease treatment to remove single-stranded DNA fragments and overhangs.
  • Evaluated the impact of S1 nuclease treatment on sequencing bias, error rates, and variant detection.
  • Created an analytical method for SSAR quantification.

Main Results:

  • A substantial proportion of chimeric reads in FFPE NGS data align to both Watson and Crick strands, identified as SSARs.
  • S1 nuclease treatment significantly reduced SSAR levels.
  • S1 nuclease treatment also decreased sequence bias, base error rates, and false positive variant calls (copy number and single nucleotide variants).

Conclusions:

  • SSARs are a prevalent artifact in FFPE NGS, arising from specific molecular mechanisms.
  • S1 nuclease treatment is an effective strategy to mitigate SSARs and improve overall NGS data quality from FFPE samples.
  • The developed analytical approach enables accurate quantification of SSARs, facilitating artifact assessment in FFPE-derived NGS studies.

Related Concept Videos

Next-generation Sequencing03:00

Next-generation Sequencing

The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
98.5K
Cis-regulatory Sequences02:02

Cis-regulatory Sequences

Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.8K
Sequences01:29

Sequences

Sequences are fundamental mathematical objects consisting of ordered lists of numbers that follow a specific rule or pattern. Sequences are critical in various mathematical concepts, including calculus, series, and number theory. They can model real-world phenomena such as population growth, financial investments, and physical processes like the diminishing height of a bouncing ball.Each number in a sequence is referred to as a term. Typically, the terms are denoted as a1, a2, a3,…, where...
277
Sanger Sequencing01:57

Sanger Sequencing

DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
774.5K
Arithmetic Sequences01:30

Arithmetic Sequences

An arithmetic sequence is a structured arrangement of numbers where each term is derived by adding a constant value, known as the common difference, to the previous term. This consistent pattern allows for the efficient computation of any term within the sequence as well as the cumulative sum of multiple terms. The formula for finding the nth term of an arithmetic sequence is:Here, aₙ represents the nth term of the sequence, a is the first term, d is the common difference, and n is the...
239
Genomics02:02

Genomics

Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
40.7K