Related Experiment Video
Updated: May 27, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Imputation of missing aggregate EHR audit log data across individual and multiple organizations
Huan Li1, Nate C Apathy2, A Jay Holmgren3
1Department of Emergency Medicine, Yale University School of Medicine, New Haven, CT, United States; Computational Biology and Bioinformatics, Yale Graduate School of Arts and Sciences, New Haven, CT, United States; Biomedical Informatics and Data Science, Yale University School of Medicine, New Haven, CT, United States.
A systematic approach to classifying variables and using tailored imputation methods, like the rolling average time-series, proved effective for missing electronic health record data. This method outperformed machine learning imputation, especially for smaller organizations.
Area of Science:
- Health Informatics
- Data Science
- Electronic Health Records (EHR)
Background:
- Missing data in EHR systems poses challenges for accurate analysis.
- Imputation strategies are crucial for handling missing EHR data.
- Evaluating different imputation methods is essential for reliable healthcare analytics.
Purpose of the Study:
- To compare naive and machine learning imputation strategies for EHR data.
- To explore subgrouping criteria for optimizing imputation.
- To assess the performance and feasibility of in-house imputation implementation.
Main Methods:
- Conducted imputation experiments on EHR audit log data from diverse organizations.
- Classified variables based on coefficient of variation and missing percentage.
- Evaluated imputation performance using R²-values, selecting the most robust model.
Main Results:
- Rolling average time-series imputation showed more consistent R² across organization sizes.
- XGBoost (machine learning) performed slightly better in large organizations.
- Single-site organizations achieved higher R² using their own data for imputation compared to merged data.
Conclusions:
- A systematic variable classification and tailored imputation strategy is effective for EHR data.
- The rolling average time-series method outperformed non-time-series machine learning imputation.
- Organization size did not significantly impact the imputation process; merging diverse site data did not improve imputation.
More Related Videos
Related Concept Videos
Methods of Documentation VII: EMR
Data Reporting and Recording
Guidelines and Strategies for Safe Computer Charting
Maintain Confidentiality and Security:
Types of Reports II: Incident or Occurrence Report
Purposes:
In the healthcare industry, reports play a crucial role in documenting incidents within an agency. The primary objective of these reports is to ensure patient safety, uphold the...
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Legal Guidelines for Documentation

