Related Experiment Video
Updated: Jan 8, 2026

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
Improving Genotype Imputation in High-Dimensional Pharmacogenomics Using Multiple Imputation: Evaluation with Machine
Innocent G Asiimwe1, Tao You2, Daniel F Carr2
1Department of Health Data Science, Institute of Population Health Sciences, University of Liverpool, Liverpool, UK.
None:
Multiple imputation is well-established for handling missing data, yet its use in high-dimensional genetic datasets remains limited. Using pharmacokinetic tuberculosis simulations and SNP data (1000 Genomes Project), we compared machine learning (ML) and traditional approaches (e.g., mean imputation and complete-case analysis) for imputation and covariate selection. We developed a multiple imputation framework incorporating genotype probabilities, imputation uncertainty (INFO score), and missingness percentages. Dimensionality reduction enabled scalable random forest and penalized regression for covariate selection. In simulations, only multiple imputation achieved adequate coverage (percentage of 95% confidence intervals containing the true value) exceeding a 90% nominal threshold. For example, on the imputation server, coverage improved from 0% with single imputation to up to 94% under 10% missingness. Applied to clinical warfarin datasets (War-PATH, n = 548; IWPC, n = 316) and the UK Biobank (n = 500, 1000), multiple imputation recovered known pharmacogenomic associations (CYP2C9*8/*9/*11; VKORC1 -1639G>A), reduced false-positives, and detected signals missed by single imputation (e.g., genome-wide significant rs4697699, SLC2A9 locus). Computational costs were modest, adding only ~1.25 minutes for 10 imputations to the 22.7 minutes required by single imputation on the Michigan Imputation Server. For SNP selection, penalized regression performed best in the high-effect scenario (F1 = 0.897 ± 0.091), while GWAS followed by random forest performed best in the low-effect scenario (F1 = 0.657 ± 0.110). These findings show that multiple imputation improves reliability and discovery in high-dimensional pharmacogenomics, with ML offering promising but inconsistent benefits during SNP selection. However, generalizability beyond the studied datasets and computational scalability to larger biobank-scale analyses remain important limitations that warrant further investigation.
Related Concept Videos
Analysis of Population Pharmacokinetic Data
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model Approaches for Pharmacokinetic Data: Compartment Models
Two primary types of compartment models are recognized: mammillary and catenary. The more...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Mechanistic Models: Compartment Models in Individual and Population Analysis

