Predicting COVID-19 severity: Challenges in reproducibility and deployment of machine learning methods
Luwei Liu1, Wenyu Song2, Namrata Patil3
1Department of Medicine, Brigham & Women's Hospital, Boston, MA, USA.
International Journal of Medical Informatics
|September 28, 2023
Summary
Electronic health records (EHR) advance AI in medicine. However, inconsistent definitions in COVID-19 severity models hinder clinical application, necessitating standardized approaches for reliable predictions.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Clinical Research Informatics
Background:
- Electronic Health Records (EHR) are increasingly utilized for data-driven medical applications and artificial intelligence (AI) in healthcare.
- EHRs provide longitudinal patient data crucial for developing AI models in medicine.
- Standardization of variables and outcome definitions is vital for the clinical applicability of AI methodologies.
Purpose of the Study:
- To review machine learning (ML) prediction models for Coronavirus Disease 2019 (COVID-19) severity using EHR data.
- To explore a framework for standardizing the development of future disease severity models based on EHR information.
- To identify inconsistencies in COVID-19 severity phenotype definitions and their impact on model generalizability.
Main Methods:
- Systematic review of 2,967 studies published between January 1, 2020, and February 15, 2022.
- Selection of 135 independent studies that developed ML models to predict COVID-19 severity outcomes using EHR data.
- Analysis of 135 studies from 27 countries focusing on severity prediction outcomes.
Main Results:
- Substantial inconsistencies were observed in the definitions of COVID-19 severity phenotypes across the reviewed ML models.
- A notable gap exists between the outcomes predicted by these models and clinically recognized concepts.
- The reviewed studies showed a broad range of severity prediction outcomes but lacked definitional uniformity.
Conclusions:
- Standardized, robust clinical input metrics and unambiguous outcome definitions are recommended to reduce bias and improve model generalizability.
- Addressing inconsistencies in phenotype definitions is crucial for developing reliable and universally applicable COVID-19 severity prediction models.
- The proposed framework for standardization can potentially be extended to other clinical applications beyond COVID-19 severity prediction.
More Related Videos
Related Concept Videos
Steps in Outbreak Investigation
152
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
152
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Improving Translational Accuracy
11.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.5K
Statistical Methods for Analyzing Epidemiological Data
400
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
400
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


