Related Experiment Video
Updated: May 31, 2025

Basics of Multivariate Analysis in Neuroimaging Data
Published on: July 24, 2010
High-dimensional multiple imputation for partially observed confounders including natural language processing-derived
Janick Weberpals1, Pamela A Shaw2, Kueiyu Joshua Lin1
1Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, United States.
High-dimensional multiple imputation (MI) using auxiliary covariates (AC) can reduce bias in studies with missing confounders. Combining structured and NLP-derived AC improved efficiency and bias in simulations.
Area of Science:
- Biostatistics
- Health Informatics
- Epidemiology
Background:
- Auxiliary covariates (AC) can enhance multiple imputation (MI) models.
- The performance of AC in high-dimensional data, especially with partially observed confounders, requires further investigation.
- Natural Language Processing (NLP) offers novel ways to derive AC from unstructured data.
Purpose of the Study:
- To develop and compare high-dimensional MI (HDMI) methods using structured and NLP-derived AC.
- To evaluate HDMI performance in simulated cohorts with partially observed confounders, specifically focusing on acute kidney injury studies.
- To assess the impact of omitting key confounders (e.g., atrial fibrillation) on imputation and treatment effect estimation.
Main Methods:
- A plasmode simulation was conducted with 100 cohorts, simulating acute kidney injury outcomes and null treatment effects.
- Creatinine lab values and atrial fibrillation (AFib) were included as confounders, with missingness imposed on creatinine.
- High-dimensional MI (HDMI) covariates were derived from structured (claims) and NLP-derived features, selected using LASSO for MI and propensity score models.
Main Results:
- HDMI using claims data demonstrated the lowest bias (0.072).
- Combining claims data with NLP-derived sentence embeddings improved efficiency (RMSE=0.173) and achieved 94% coverage.
- NLP-derived AC alone did not outperform baseline MI methods.
- Complete case analysis and baseline imputation showed higher bias compared to HDMI methods.
Conclusions:
- High-dimensional multiple imputation (HDMI) with auxiliary covariates can effectively decrease bias in studies where confounder missingness is related to unobserved factors.
- The integration of structured data (claims) and NLP-derived features offers a promising approach for improving the accuracy and efficiency of imputation in complex, high-dimensional datasets.
- Careful selection and derivation of auxiliary covariates are crucial for successful application of HDMI methods in observational health research.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
08:19Simultaneous Data Collection of fMRI and fNIRS Measurements Using a Whole-Head Optode Array and Short-Distance Channels
Published on: October 20, 2023
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Confounding in Epidemiological Studies
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Multi-input and Multi-variable systems
In the absence...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.