Related Experiment Video
Updated: Oct 17, 2025

06:55
Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
14.7K
Imputation of missing values for electronic health record laboratory data
Jiang Li1, Xiaowei S Yan2, Durgesh Chaudhary1
1Geisinger Health System, Danville, PA, USA.
NPJ Digital Medicine
|October 12, 2021
Summary
Imputation methods improve prediction models using electronic health records (EHR) laboratory data. Multi-level imputation showed lower error than cross-sectional methods, especially when accounting for patient comorbidity data.
Area of Science:
- Biomedical Informatics
- Health Data Science
- Clinical Prediction Modeling
Background:
- Electronic Health Records (EHR) contain valuable laboratory data for clinical prediction models.
- Missing laboratory data in EHR can introduce estimation bias and degrade model performance.
- Imputation methods are crucial for addressing missingness in EHR data.
Purpose of the Study:
- To demonstrate the utility of imputation methods for EHR laboratory data.
- To characterize missingness patterns and compare imputation algorithms.
- To assess the impact of comorbidity data on imputation performance.
Main Methods:
- Utilized two real-world EHR cohorts: ischemic stroke (Geisinger) and heart failure (Sutter Health).
- Characterized missingness patterns in laboratory variables.
- Simulated missing data under arbitrary and monotone mechanisms.
- Compared cross-sectional and multi-level multivariate imputation algorithms.
- Assessed the incorporation of latent comorbidity information.
Main Results:
- Missingness patterns in EHR laboratory data were non-random and associated with patient comorbidity data.
- Multi-level multivariate imputation algorithms exhibited lower imputation error compared to cross-sectional methods.
- Incorporating latent comorbidity information showed potential for improving imputation performance.
Conclusions:
- Imputation is essential for leveraging EHR laboratory data in prediction models.
- Multi-level imputation methods are more effective than cross-sectional approaches for EHR laboratory data.
- Patient comorbidity data provides valuable latent information that can enhance imputation accuracy.
Related Concept Videos
Methods of Documentation VII: EMR
1.1K
Electronic Medical Records (EMRs) primarily center around electronically documenting patients' health information within a single healthcare organization or practice. They contain essential clinical data related to a patient's medical history, diagnoses, medications, treatment plans, lab results, and other pertinent information relevant to the specific encounter or episode of care. EMRs are designed to streamline documentation and workflow processes within individual healthcare...
1.1K
Documentation of Nursing Diagnosis
1.4K
The nurse documents nursing diagnoses and enters them into the patient record. The identified patient's nursing diagnosis is either written out with a plan of care or entered into the electronic health record.
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
1.4K
Data Reporting and Recording
5.0K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.0K
Detection of Gross Error: The Q Test
6.5K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.5K

