Impact of post-alignment processing in variant discovery from whole exome data
Shulan Tian1, Huihuang Yan1, Michael Kalmbach2
1Division of Biomedical Statistics and Informatics, Department of Health Sciences Research, Mayo Clinic, 200 1st St SW, Rochester, MN, 55905, USA.
BMC Bioinformatics
|October 8, 2016
Summary
Post-processing steps like local realignment and base quality score recalibration (BQSR) do not universally improve variant calling accuracy. Their benefit depends on the specific mapper-caller combination, sequencing depth, and genomic divergence.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- GATK Best Practices recommend post-alignment processing (local realignment, BQSR) for large-scale sequencing.
- The benefit of these computationally intensive steps with updated mappers/callers and new sequencing platforms is unclear.
- Impact in highly divergent genomic regions, like the HLA system, remains unknown.
Purpose of the Study:
- To comprehensively assess the impact of post-processing on variant calling accuracy.
- To evaluate the influence of different mappers, callers, sequencing depth, and divergence levels.
- To determine if post-processing offers benefits across various analytical pipelines.
Main Methods:
- Utilized simulated and NA12878 exome data for analysis.
- Tested five to six popular mappers combined with five variant callers.
- Focused on chromosome 6p21.3 (HLA region) for high-divergence analysis.
Main Results:
- Local realignment had minimal impact on SNP calling but improved INDEL calling in specific pipelines.
- Base Quality Score Recalibration (BQSR) showed negligible effects on INDEL calling and often reduced SNP calling sensitivity.
- BQSR's impact on SNP calling varied by caller, coverage, and divergence, with reduced sensitivity in high-divergence regions.
Conclusions:
- The utility of post-processing is not universal and is highly dependent on the specific bioinformatics pipeline.
- Sequencing depth and the level of genomic divergence significantly influence the effectiveness of post-processing.
- Careful consideration of these factors is crucial when deciding whether to implement computationally intensive post-processing steps for exome data.
Keywords:
Base quality score recalibrationHuman leukocyte antigenLocal realignmentVariant callingWhole exome sequencingMore Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
16.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
16.4K
Comparing Copy Number Variations and SNPs
19.0K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.0K
Sanger Sequencing
777.2K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
777.2K
Evolutionary Relationships through Genome Comparisons
7.2K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
7.2K


