Related Experiment Video
Updated: Feb 24, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Adjusting Covariate Misclassification in Electronic Health Records-Based Machine Learning Prediction Models
Shuang Yang1, Yonghui Wu1, Mei Liu1
1Department of Health Outcomes and Biomedical Informatics, University of Florida, Gainesville, Florida, USA.
This study introduces methods to correct errors in electronic health record data, improving predictive models for lung cancer screening. Adjusted models showed better accuracy than unadjusted ones, reducing bias without manual review.
Area of Science:
- Biomedical Informatics
- Health Services Research
- Machine Learning in Healthcare
Background:
- Electronic health records (EHRs) contain valuable data but are prone to misclassification errors.
- These errors can introduce bias into predictive models, affecting clinical decision-making.
- Accurate prediction of patient outcomes, such as adherence to lung cancer screening, is crucial.
Purpose of the Study:
- To develop and evaluate methods for adjusting misclassification errors in EHR-derived covariates.
- To reduce bias in predictive modeling using group-wise and individualized weights.
- To enhance the accuracy of predictive models for lung cancer screening adherence.
Main Methods:
- Developed methods using group-wise and individualized weights based on sensitivity and specificity to adjust EHR covariates.
- Applied logistic regression, XGBoost, and neural networks to predict lung cancer screening adherence.
- Utilized natural language processing (NLP) for Lung-RADS category extraction and adjusted it using kernel and multinomial regression weights.
- Compared adjusted models against unadjusted (naïve) and true value (oracle) models using Area Under the Receiver Operating Characteristic (AUROC) curves.
Main Results:
- Adjusted models significantly outperformed naïve models across various validation set sizes (10%, 20%, 30%).
- AUROC improvements ranged from 0.3% to 10.4% compared to naïve models.
- Adjusted models reduced the performance gap compared to oracle models to 2.0%-7.5%.
- Individualized weights demonstrated more precise error correction than group-wise weights.
Conclusions:
- The developed framework effectively mitigates misclassification bias in EHR-derived covariates.
- The methods enhance predictive accuracy for lung cancer screening adherence without requiring extensive manual data review.
- This approach offers a scalable solution for improving the reliability of predictive models using real-world clinical data.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Confounding in Epidemiological Studies
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Bias in Epidemiological Studies
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Methods for Analyzing Epidemiological Data