Related Experiment Video
Updated: Jun 7, 2025

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
Empirical versus estimated accuracy of imputation: optimising filtering thresholds for sequence imputation
Tuan V Nguyen1, Sunduimijid Bolormaa2, Coralie M Reich2
1Agriculture Victoria, Centre for AgriBiosciences, AgriBio, Bundoora, VIC, 3083, Australia. tuan.nguyen@agriculture.vic.gov.au.
Genotype imputation accuracy varies by software and data density. Customizing imputation accuracy thresholds (Rsqsoft) is crucial for reliable genome-wide association studies (GWAS), especially for the X chromosome's PAR region.
Area of Science:
- Genomics
- Bioinformatics
- Population Genetics
Background:
- Genotype imputation is vital for cost-effective genomic analyses like genome-wide association studies (GWAS).
- Inaccurate imputation can lead to false positives, necessitating data pre-filtering or accuracy assessment.
- This study benchmarks imputation software and evaluates imputation accuracy in cattle.
Purpose of the Study:
- To benchmark Beagle 5.2, Minimac4, and IMPUTE5 for genotype imputation accuracy.
- To compare empirical imputation accuracy with software-estimated accuracy (Rsqsoft).
- To assess imputation accuracy across different cattle chromosomes (autosomal, X), variant types (SNP, INDEL), and genotype densities (low, high).
Main Methods:
- Benchmarking three imputation programs: Beagle 5.2, Minimac4, and IMPUTE5.
- Comparing empirical imputation accuracy against software-estimated Rsqsoft values.
- Evaluating imputation performance for SNPs and INDELs on autosomal and X chromosomes, using low-density and high-density genotypes.
Main Results:
- Higher imputation accuracy was achieved using high-density genotypes compared to low-density.
- All tested software performed well, with minor accuracy differences.
- Empirical imputation accuracy closely correlated with Rsqsoft, but differed between Minimac4 and Beagle 5.2/IMPUTE5.
- Customizing Rsqsoft thresholds per software is essential for merging data and for regions with poor imputation accuracy (e.g., segmental duplications).
- Indel imputation accuracy was ~6% lower than SNP imputation accuracy.
- Non-PAR X chromosome imputation accuracy was comparable to autosomal, but PAR accuracy was substantially lower, especially with low-density genotypes.
Conclusions:
- An empirically derived approach for applying customized, software-specific Rsqsoft thresholds is proposed for downstream analyses like meta-GWAS.
- The Pseudo-Autosomal Region (PAR) on the X chromosome requires high-density genotypes for accurate imputation, particularly when starting from low-density data.
Related Concept Videos
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Detection of Gross Error: The Q Test
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Censoring Survival Data
Improving Translational Accuracy

