Related Experiment Video
Updated: Mar 16, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.7K
Statistical modeling for sensitive detection of low-frequency single nucleotide variants.
Yangyang Hao1,2, Pengyue Zhang2,3, Xiaoling Xuei4,5
1Department of Medical and Molecular Genetics, Indiana University School of Medicine, Indianapolis, IN, 46202, USA.
BMC Genomics
|August 25, 2016
Summary
This study introduces a new method for accurately detecting low-frequency single nucleotide variants (SNVs) down to 0.5%. The developed approach enhances precision in cancer genetics and population studies by modeling sequencing errors.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Low-frequency single nucleotide variants (SNVs) are crucial in cancer genetics and population studies.
- Challenges include tumor heterogeneity, low circulating tumor DNA in liquid biopsies, and signal dilution in pooled sequencing.
- Existing methods lack sensitivity and fail to account for sequencing artifacts, with detection limits around 2-5%.
Purpose of the Study:
- To develop a method for sensitive detection of low-frequency SNVs, pushing the detection limit close to sequencing error rates.
- To model erroneous read counts based on genomic sequence contexts.
- To evaluate different statistical distributions for count data modeling in generalized linear models.
Main Methods:
- Modeled observed erroneous read counts using genomic sequence contexts.
- Characterized four count data distributions (generalized linear models) for goodness-of-fit and performance on real sequencing data.
- Tested the method on two sequencing technologies (Ion Proton and Illumina MiSeq) with different chemistries to identify systematic errors.
Main Results:
- The zero-inflated negative binomial distribution generalized linear model demonstrated superior performance, especially for variants between 0.5% and 1%.
- Achieved high recall and precision across different sequencing platforms: 95.3% recall and 79.9% precision for Ion Proton, and 95.6% recall and 97.0% precision for Illumina MiSeq at >=1% frequency.
- The method's detection limit was found to be around 0.5% SNVs and is generalizable to various sequencing technologies.
Conclusions:
- The developed method enables sensitive detection of low-frequency SNVs across diverse sequencing platforms.
- Facilitates advancements in pooled sequencing, early cancer detection, prognostic assessment, and monitoring of metastasis, relapse, or acquired resistance.
Related Concept Videos
Comparing Copy Number Variations and SNPs
19.1K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.1K
Single Nucleotide Polymorphisms-SNPs
19.5K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
19.5K

