Related Experiment Video
Updated: Jan 22, 2026

Hybrid De Novo Genome Assembly for the Generation of Complete Genomes of Urinary Bacteria using Short- and Long-read Sequencing Technologies
Published on: August 20, 2021
Recurrent miscalling of missense variation from short-read genome sequence data.
Matthew A Field1,2, Gaetan Burgio1, Aaron Chuah1
1Department of Immunology and Infectious Disease, The John Curtin School of Medical Research, The Australian National University, Canberra, Australian Capital Territory, Australia.
Short-read sequencing can miscall genetic variants due to alignment issues in repetitive genomic regions. Identifying these recurrent false positives improves genome data quality and interpretation accuracy.
Area of Science:
- Genomics
- Bioinformatics
- Population Genetics
Background:
- Short-read genome resequencing generates extensive genetic variation data.
- Exhaustive validation of numerous variants is often infeasible.
- Undetected variant miscalls can disproportionately impact individual genome interpretation and public variation databases.
Purpose of the Study:
- To investigate the nature and prevalence of sequence variation miscalling in short-read data.
- To identify sequence-intrinsic factors contributing to recurrent false positive variant calls.
- To develop methods for identifying and mitigating miscalled variants.
Main Methods:
- Analysis of short-read sequencing data from human exomes and mouse strains.
- Simulations involving read resampling, realignment, and variant recalling.
- Assessment of variant call sensitivity to sequence read length and genomic distance from reference.
Main Results:
- Recurrent, sequence-intrinsic miscalling of genetic variants is prevalent in short-read data.
- Miscalls are sensitive to read length and arise from alignment difficulties in redundant genomic regions.
- Thousands of recurrent false positive variants were identified per individual (human and mouse), many present in public databases.
- Over two-thirds of false positive variation can be identified using simulation-based approaches.
Conclusions:
- Removing recurrent false positives enhances the quality of individual variant datasets.
- Miscalled variants are sequence-specific and can be characteristic of individuals, pedigrees, or ethnic groups.
- Read length significantly influences variant calling, impacting cohort studies using diverse datasets.
Related Concept Videos
Conservative Site-specific Recombination and Phase Variation
The recognition sites for Cre recombinase called LoxP...
Genomics
What is Variation?
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
Genome Size and the Evolution of New Genes
Cis-regulatory Sequences
Uncertainty in Measurement: Reading Instruments

