Related Experiment Video
Updated: Apr 27, 2026

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
16.5K
Toward better understanding of artifacts in variant calling from high-coverage samples
1Medical Population Genetics Program, Broad Institute of Harvard and MIT, Cambridge, MA 02142, USA.
Bioinformatics (Oxford, England)
|June 30, 2014
Summary
This study identifies erroneous realignment in low-complexity regions and incomplete reference genomes as key error sources in whole-genome sequencing variant calling. Post-filtering significantly reduces genotype call error rates without compromising sensitivity.
Area of Science:
- Genomics
- Bioinformatics
Background:
- Whole-genome sequencing is crucial for personal and cancer genomics.
- Evaluating variant calling methods is challenging due to the lack of unbiased truth sets.
Purpose of the Study:
- To identify major error sources in whole-genome variant calling.
- To assess the impact of filtering on variant call accuracy.
Main Methods:
- Generated variant call sets using two read mappers and five variant callers on haploid and diploid human genomes.
- Analyzed false heterozygous calls in the haploid genome to pinpoint error origins.
Main Results:
- Identified erroneous realignment in low-complexity regions and incomplete reference genomes as primary error sources.
- Estimated raw genotype call error rates at 1 in 10-15 kb.
- Reduced post-filtered call error rates to 1 in 100-200 kb with maintained sensitivity.
Conclusions:
- Erroneous realignment and reference genome incompleteness require further improvement in variant calling pipelines.
- Post-filtering effectively enhances variant call accuracy in whole-genome sequencing.
Related Concept Videos
Comparing Copy Number Variations and SNPs
11.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
11.5K
Sanger Sequencing
800.5K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
800.5K
Single Nucleotide Polymorphisms-SNPs
14.4K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.4K
Next-generation Sequencing
87.7K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
87.7K
DNA Microarrays
16.7K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
16.7K
Histone Variants at the Centromere
4.0K
Histone variants are the histone proteins with structural and sequence variations. These variants may be regarded as “mutant” forms that replace their canonical histone counterparts in the nucleosomes. Specific post-translational modifications on the histone variants enable further chromatin complexity and regulate tissue-specific gene expression. The most common histone variants are from histone H2A, H2B, and linker histone H1 families. However, several variants of histone H3...
4.0K

