How local reference panels improve imputation in French populations
Anthony F Herzig1, Lourdes Velo-Suárez2,3,
1Univ Brest, Inserm, EFS, UMR 1078, GGB, Brest, France. anthony.herzig@inserm.fr.
Scientific Reports
|January 3, 2024
Summary
Study-specific panels (SSPs) improve genome imputation accuracy for French individuals, especially when geographically matched. Combining SSPs with public panels or merging imputation results offers superior accuracy over single methods.
Area of Science:
- Genomics
- Population Genetics
- Bioinformatics
Background:
- Public reference panels offer high imputation precision for European genomes.
- Study-specific panels (SSPs) can enhance imputation, but combining them with public panels is difficult using external servers.
- Geographic proximity between reference and target individuals is crucial for accurate imputation.
Purpose of the Study:
- To compare genome imputation accuracy using a public panel versus a French study-specific panel (SSP).
- To investigate scenarios favoring SSP-based versus server-based imputation.
- To propose a method for combining imputation strategies for improved accuracy.
Main Methods:
- Imputed 550 French individuals using the University of Michigan imputation server (Haplotype Reference Consortium panel) and an in-house French SSP (850 individuals).
- Analyzed imputation accuracy based on geographic proximity of target and SSP individuals.
- Investigated haplotype sharing and population fine-structure in France.
Main Results:
- SSP-based imputation showed benefits for haplotype phasing and rare variant imputation (MAF < 0.01), particularly for geographically matched individuals.
- Imputation accuracy improved to 58.1% for individuals from regions well-covered by the SSP, compared to 42.3% overall.
- A pragmatic approach combining server-based and SSP-based imputation results via posterior genotype probabilities achieved higher accuracy than either method alone.
Conclusions:
- Study-specific panels offer significant advantages for imputing French genomes, especially for rare variants and when geographically relevant.
- Combining imputation strategies is essential for maximizing accuracy, particularly when direct combination of panels is not feasible.
- The findings provide insights into expected imputation accuracy across different strategies for European populations.
Related Concept Videos
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
Improving Translational Accuracy
2.6K
2.6K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Mechanistic Models: Compartment Models in Individual and Population Analysis
43
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
43
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Stratified Sampling Method
12.0K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
12.0K


