Related Experiment Video
Updated: May 30, 2025

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
High-dimensional, outcome-dependent missing data problems: Models for the human loci
Lars Leonardus Joannes van der Burg1, Hein Putter1, Henning Baldauf2
1Biomedical Data Sciences, LUMC, Leiden, The Netherlands.
Incorporating outcome models into missing data imputation for KIR diplotypes can introduce bias. A baseline expectation-maximization algorithm without outcome modeling often performs better or comparably, especially in high-dimensional biological data.
Area of Science:
- Genetics
- Biostatistics
- Computational Biology
Background:
- Missing data is prevalent in high-dimensional biological datasets.
- Imputation and expectation-maximization (EM) algorithms are used for data reconstruction.
- Integrating regression models into imputation may reduce bias in regression coefficients.
Purpose of the Study:
- To evaluate outcome-based EM algorithms for reconstructing KIR diplotypes with missing data.
- To compare strategies incorporating high-dimensional regression models against a baseline EM algorithm.
Main Methods:
- Extended a previously proposed EM algorithm to include a high-dimensional regression model.
- Evaluated three strategies: allelic predictors only, allelic predictors with haplotype selection, and penalized regression.
- Compared these strategies with a baseline EM algorithm without an outcome model via simulation.
Main Results:
- Outcome-based EM algorithms outperformed the baseline in extreme scenarios of effect sizes and missingness.
- In most cases, the baseline EM algorithm performed superiorly or comparably.
- The inclusion of an outcome model can potentially introduce harmful effects and bias.
Conclusions:
- Outcome-based missing data models in high-dimensional settings require careful application.
- These models may lead to biased results, particularly when reconstructing KIR diplotypes.
- A baseline EM algorithm without outcome modeling is often a more robust approach.
More Related Videos
08:27Large-Scale Multi-Omics Genome-Wide Association Studies Mo-GWAS: Guidelines for Sample Preparation and Normalization
Published on: July 27, 2021
06:52Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills
Published on: September 17, 2019
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Pleiotropy
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
Multiple Allele Traits