Related Experiment Videos
A computational method for estimating the PCR duplication rate in DNA and RNA-seq experiments.
1Department of Pediatrics, School of Medicine, University of California San Diego, 9500 Gilman Drive, 92093, La JollaCA, USA. vibansal@cs.ucsd.edu.
BMC Bioinformatics
|April 1, 2017
Summary
This study introduces a new computational method to accurately estimate PCR duplication rates in DNA sequencing data. The method distinguishes true PCR duplicates from natural duplicates, improving data quality assessment for DNA-seq and RNA-seq.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Polymerase chain reaction (PCR) amplification is crucial for DNA sequencing library preparation.
- PCR amplification can introduce redundant reads, necessitating accurate estimation of PCR duplication rates.
- Existing methods overestimate PCR duplication by failing to differentiate between PCR duplicates and natural read duplicates.
Purpose of the Study:
- To develop a computational method for accurately estimating average PCR duplication rates in high-throughput sequencing datasets.
- To differentiate PCR duplicates from natural read duplicates in DNA-seq and RNA-seq experiments.
- To improve the assessment of read duplication frequency in sequencing data.
Main Methods:
- Developed a novel computational method leveraging heterozygous variants to distinguish PCR duplicates from natural duplicates.
- Applied the method to simulated data and exome sequence data from the 1000 Genomes project.
- Validated the method on paired-end and single-end read datasets, including those with high proportions of natural duplicates.
Main Results:
- The method accurately estimates PCR duplication rates in both simulated and real-world exome and RNA-seq datasets.
- Analysis revealed that 45-50% of read duplicates in exome data (Nextera preparation) are natural duplicates, attributed to fragmentation bias.
- In RNA-seq data, 70-95% of observed read duplicates are natural duplicates from highly expressed genes, with outlier samples showing doubled PCR duplication rates.
Conclusions:
- The developed method provides a valuable tool for estimating PCR duplication rates and quantifying natural read duplicates in high-throughput sequencing data.
- This approach enhances the reliability of sequence data analysis by accurately accounting for different types of read duplication.
- An implementation of the method is publicly available for researchers.