Related Experiment Video
Updated: Jul 19, 2026

07:31
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Develop and validate a fair machine learning model to identify patients with high data-continuity in electronic
Yao An Lee1, Tiange Tang2, Yu Huang1,3
1Center for Biomedical Informatics, Regenstrief Institute, Indianapolis, IN 46202, United States.
JAMIA Open
|July 14, 2026
Summary
Electronic health record (EHR) data discontinuity can lead to misclassification. A new machine learning model accurately predicts EHR continuity, improving research rigor and fairness across racial and ethnic groups.
Area of Science:
- Health Informatics
- Machine Learning in Healthcare
- Data Quality in Research
Background:
- Electronic health record (EHR) data discontinuity, arising from care outside a single system, can cause significant misclassification of study variables.
- This misclassification poses a challenge to the accuracy and reliability of research conducted using EHR data.
Purpose of the Study:
- To quantify misclassification across varying levels of EHR data discontinuity and establish an optimal continuity threshold.
- To develop a machine learning (ML) model for predicting EHR continuity, with a focus on optimizing fairness across racial and ethnic groups.
- To externally validate the developed ML model using an independent dataset.
Main Methods:
- Linked EHR-Medicaid claims data (OneFlorida+) for development and EHR-LABlue claims data (REACHnet) for validation.
- Utilized a Harmonized Encounter Proportion Score (HEPS) to quantify patient-level EHR data continuity and its impact on 42 clinical variables.
- Trained ML models on routinely available demographic, clinical, and healthcare utilization features from structured EHR data.
Main Results:
- Higher EHR data continuity correlated with reduced misclassification rates.
- An optimal HEPS threshold of approximately 30% was identified for sufficient data continuity.
- ML models achieved strong predictive performance (AUROC = 0.77) for high continuity, with bias against the Hispanic group substantially mitigated.
- External validation confirmed robust and fair model performance.
Conclusions:
- A practical metric (HEPS) for quantifying EHR data continuity in networks has been established.
- The developed ML model accurately identifies patients with high care continuity using routinely collected EHR information.
- A generalizable data-continuity classification tool was created, enhancing the rigor of EHR-based research across diverse systems.