Related Experiment Video
Updated: Aug 19, 2025

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
15.3K
Benchmarking challenging small variants with linked and long reads
Justin Wagner1, Nathan D Olson1, Lindsay Harris1
1Material Measurement Laboratory, National Institute of Standards and Technology, 100 Bureau Dr, MS8312, Gaithersburg, MD 20899, USA.
Cell Genomics
|December 1, 2022
Summary
New Genome in a Bottle benchmarks use long reads to improve variant calling accuracy, especially in difficult genomic regions. This enhanced benchmark identifies more false negatives, aiding sequencing method development.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- Genome in a Bottle (GiaB) benchmarks are crucial for validating clinical sequencing and developing variant calling methods.
- Short-read sequencing technologies face challenges in accurately mapping complex genomic regions like segmental duplications.
- Existing benchmarks may not fully cover clinically relevant genes or difficult-to-map areas.
Purpose of the Study:
- To expand existing Genome in a Bottle benchmarks using accurate long and linked reads.
- To incorporate challenging genomic regions, including segmental duplications and difficult-to-map areas.
- To improve the accuracy and comprehensiveness of variant calling benchmarks for clinical sequencing.
Main Methods:
- Utilized accurate long and linked reads to generate expanded benchmarks across seven samples.
- Incorporated difficult-to-map regions and segmental duplications into the benchmark datasets.
- Expanded variant sets to include over 300,000 single nucleotide variants (SNVs) and 50,000 insertions/deletions (indels).
- Increased exonic variant representation by 16%, focusing on clinically relevant genes like PMS2.
Main Results:
- The expanded benchmark covers 92% of the autosomal GRCh38 assembly for HG002, excluding problematic regions.
- Identified eight times more false negatives in a short-read variant call set compared to previous benchmarks.
- Added significant numbers of SNVs and indels, enhancing coverage of challenging exonic regions.
- Demonstrated improved identification of false positives and false negatives across different sequencing technologies.
Conclusions:
- The new long-read-based benchmarks provide a more comprehensive and accurate resource for evaluating sequencing pipelines.
- These benchmarks are essential for identifying limitations in short-read sequencing data and improving variant calling algorithms.
- The enhanced datasets facilitate the development of more reliable sequencing methods for clinical applications, particularly for complex genomic regions and genes.

