Related Experiment Video
Updated: Apr 25, 2026

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
Improved variant calling accuracy by merging replicates in whole-exome sequencing studies.
Yanfeng Zhang1, Bingshan Li2, Chun Li3
1Division of Epidemiology, Department of Medicine, Vanderbilt Epidemiology Center, Vanderbilt-Ingram Cancer Center, Vanderbilt University School of Medicine, Nashville, TN 37203, USA.
Merging duplicated whole-exome sequencing (WES) data improves variant calling accuracy. Combining lower and higher depth sequencing data (group M) yielded more accurate single nucleotide polymorphisms (SNPs) than using either depth alone.
Area of Science:
- Genomics
- Population Genetics
- Bioinformatics
Background:
- Large-scale population studies often involve whole-exome sequencing (WES).
- Duplicate sequencing of samples can occur for various reasons, presenting challenges for data utilization.
- Efficiently leveraging duplicated WES data is crucial for maximizing study power and accuracy.
Purpose of the Study:
- To evaluate different variant calling strategies for duplicated whole-exome sequencing data.
- To determine the optimal method for utilizing redundant sequencing information in population studies.
- To assess the impact of merging data on variant calling accuracy.
Main Methods:
- Selected 92 samples from a large population study that underwent whole-exome sequencing twice.
- Divided samples into high-depth (H) and low-depth (L) groups, and created a merged group (M).
- Employed the GATK multisample toolkit to compare variant calling accuracy across the three groups, analyzing metrics like Hete/Homo ratio, Ti/Tv ratio, and overlap with the 1000 Genomes Project.
Main Results:
- Hierarchical clustering confirmed high homogeneity between the two sequencing replicates for each subject.
- Variant calling accuracy, assessed by multiple metrics, was consistently higher in the merged group (M) compared to the high-depth (H) and low-depth (L) groups.
- Single nucleotide polymorphisms (SNPs) detected from merged data demonstrated superior data quality.
Conclusions:
- Merging homogeneous duplicated exome sequencing data significantly improves variant calling accuracy.
- This strategy offers a more efficient way to utilize redundant sequencing data in large-scale population studies.
- Combining data from multiple sequencing runs of the same sample enhances the reliability of variant detection.
Related Concept Videos
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Sanger Sequencing
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...

