Related Experiment Video
Updated: Jun 19, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Incorporating informatively collected laboratory data from EHR in clinical prediction models
Minghui Sun1, Matthew M Engelhard2, Armando D Bedoya3
1Department of Biostatistics and Bioinformatics, Duke University, Durham, NC, USA. ms1008@duke.edu.
Handling informative missing data in Electronic Health Records (EHR) is crucial for accurate clinical prediction models (CPMs). Strategies accounting for Not Missing at Random (NMAR) data improve CPM performance, especially with embedding methods.
Area of Science:
- Health Informatics
- Machine Learning in Healthcare
- Clinical Data Science
Background:
- Electronic Health Records (EHR) are foundational for clinical prediction models (CPMs).
- Informative missing data, particularly Not Missing at Random (NMAR) laboratory values, poses a significant challenge.
- Standard imputation methods are inadequate for NMAR data, necessitating specialized handling strategies.
Purpose of the Study:
- To compare the performance of various missing data handling strategies for clinical prediction models.
- To evaluate the impact of different imputation techniques on models predicting rapid inpatient deterioration.
- To identify optimal methods for managing NMAR data in EHR-derived CPMs.
Main Methods:
- A predictive model for rapid inpatient deterioration was developed using twelve laboratory measures with high missingness rates (50-90%).
- Compared imputation strategies included mean imputation, normal-value imputation, conditional imputation, categorical encoding, and missingness embeddings, some with Last Observation Carried Forward (LOCF).
- Downstream classifiers comprised logistic LASSO regression, Multilayer Perceptron (MLP), and Long Short-Term Memory (LSTM) models, with performance assessed by AUROC and bootstrapping.
Main Results:
- Long Short-Term Memory (LSTM) models generally outperformed other models.
- Embedding approaches and categorical encoding demonstrated superior performance among tested strategies.
- For cross-sectional models, normal-value imputation combined with LOCF yielded the best results.
Conclusions:
- Strategies explicitly addressing Not Missing at Random (NMAR) data significantly enhance clinical prediction model performance.
- Embedding methods offer an advantage by not requiring pre-existing clinical expertise.
- While Last Observation Carried Forward (LOCF) can benefit cross-sectional models, its application may negatively impact LSTM model performance.
Related Concept Videos
Methods of Documentation VII: EMR
Statistical Software for Data Analysis and Clinical Trials
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Steps in Outbreak Investigation
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Purpose of Health Records II

