Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

RNA-seq03:21

RNA-seq

9.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.4K
Next-generation Sequencing03:00

Next-generation Sequencing

87.9K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
87.9K
Multi-species Conserved Sequences02:51

Multi-species Conserved Sequences

3.3K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale  studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The mutational landscape and functional effects of noncoding ultraconserved elements in human cancers.

Science advances·2025
Same author

Extend the benchmarking indel set by manual review using the individual cell line sequencing data from the Sequencing Quality Control 2 (SEQC2) project.

Scientific reports·2024
Same author

Cross-oncopanel study reveals high sensitivity and accuracy with overall analytical performance depending on genomic regions.

Genome biology·2021
Same author

Selective AKT kinase inhibitor capivasertib in combination with fulvestrant in PTEN-mutant ER-positive metastatic breast cancer.

NPJ breast cancer·2021
Same author

Evaluating the analytical validity of circulating tumor DNA sequencing assays for precision oncology.

Nature biotechnology·2021
Same author

Genomic, Transcriptomic, and Proteomic Profiling of Metastatic Breast Cancer.

Clinical cancer research : an official journal of the American Association for Cancer Research·2021

Related Experiment Video

Updated: May 4, 2026

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

14.6K

Bias from removing read duplication in ultra-deep sequencing experiments.

Wanding Zhou1, Tenghui Chen1, Hao Zhao1

  • 1Department of Bioinformatics and Computational Biology, Department of Systems Biology, Institute of Personalized Cancer Therapy and Department of Investigational Cancer Therapy, The University of Texas MD Anderson Cancer Center, Houston TX 77030, USA.

Bioinformatics (Oxford, England)
|January 7, 2014
PubMed
Summary

Accurate variant allele fraction estimation requires accounting for sampling-induced read duplication in deep sequencing. This study quantifies overcorrection from duplicate read removal and offers solutions for improved accuracy.

More Related Videos

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome
06:40

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome

Published on: March 22, 2018

5.0K
Ultra-long Read Sequencing for Whole Genomic DNA Analysis
10:34

Ultra-long Read Sequencing for Whole Genomic DNA Analysis

Published on: March 15, 2019

24.9K

Related Experiment Videos

Last Updated: May 4, 2026

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

14.6K
G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome
06:40

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome

Published on: March 22, 2018

5.0K
Ultra-long Read Sequencing for Whole Genomic DNA Analysis
10:34

Ultra-long Read Sequencing for Whole Genomic DNA Analysis

Published on: March 15, 2019

24.9K

Area of Science:

  • Genomics
  • Bioinformatics
  • Computational Biology

Background:

  • Accurate estimation of mutant allele fractions is crucial for identifying subclonal mutations.
  • Duplicate sequencing reads can arise from PCR amplification or DNA fragmentation, impacting variant allele fraction (VAF) and copy number variation (CNV) estimations.
  • Systematic investigation into sampling-induced read duplication in deep sequencing data is lacking.

Purpose of the Study:

  • To investigate the impact of sampling-induced read duplication on VAF and CNV estimations in deep sequencing.
  • To develop a quantitative solution for overcorrection caused by duplicate read removal.
  • To provide guidance for designing deep sequencing platforms that ensure accurate VAF and CNV estimations.

Main Methods:

  • Analysis of deep sequencing data with 500× to 2000× coverage.
  • Investigating sampling coincidence from DNA fragmentation as a source of read duplication.
  • Developing a quantitative method to correct for overcorrection bias.

Main Results:

  • Sampling-induced read duplication is non-negligible in deep sequencing and can lead to systemic biases in VAF and CNV estimations if not properly handled.
  • Minimal overcorrection is achieved when duplicate reads are identified considering mate reads, varied insert lengths, and separate batch sequencing.
  • A quantitative solution for overcorrection bias in deep sequencing data is provided.

Conclusions:

  • Standard duplicate read removal methods can introduce significant biases in VAF and CNV estimations, particularly in deep sequencing.
  • Accounting for sampling-induced read duplication is essential for accurate genomic analyses.
  • The study offers practical guidance for optimizing deep sequencing strategies and provides a computational tool for addressing read duplication biases.