预测COVID-19的严重程度:机器学习方法的复制性和部署的挑战
Luwei Liu1, Wenyu Song2, Namrata Patil3
1Department of Medicine, Brigham & Women's Hospital, Boston, MA, USA.
International journal of medical informatics
|September 28, 2023
概括
电子健康记录 (EHR) 在医学中推进了人工智能. 然而,COVID-19严重程度模型中的不一致定义阻碍了临床应用,需要标准化的方法来进行可靠的预测.
科学领域:
- 医疗信息学 医疗信息学
- 人工智能在医学中的应用
- 临床研究信息学 临床研究信息学
背景情况:
- 电子健康记录 (EHR) 越来越多地用于数据驱动的医疗应用和医疗保健中的人工智能 (AI).
- 电子健康记录提供了对开发医学AI模型至关重要的纵向患者数据.
- 变量和结果定义的标准化对于人工智能方法的临床适用性至关重要.
研究的目的:
- 通过使用EHR数据审查2019年新冠肺炎疾病 (COVID-19) 严重性的机器学习 (ML) 预测模型.
- 探索一个框架来标准化基于EHR信息的未来疾病严重程度模型的开发.
- 识别COVID-19严重性表型定义中的不一致性及其对模型概括性的影响.
主要方法:
- 对2020年1月1日至2022年2月15日期间发表的2967项研究进行了系统审查.
- 选择了135项独立研究,这些研究开发了ML模型,使用EHR数据预测COVID-19严重程度的结果.
- 分析了来自27个国家的135项研究,重点关注严重程度预测结果.
主要成果:
- 在审查的ML模型中,在COVID-19严重性表型的定义中观察到实质性的不一致性.
- 这些模型预测的结果与临床上公认的概念之间存在着显著的差距.
- 审查的研究显示了广泛的严重性预测结果,但缺乏定义统一性.
结论:
- 建议采用标准化,强大的临床输入指标和明确的结果定义,以减少偏差并提高模型通用性.
- 解决表型定义中的不一致性对于开发可靠和普遍适用的COVID-19严重性预测模型至关重要.
- 拟议的标准化框架有可能扩展到其他临床应用,超出COVID-19严重性预测范围.
相关概念视频
Steps in Outbreak Investigation
152
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
152
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Improving Translational Accuracy
11.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.5K
Statistical Methods for Analyzing Epidemiological Data
400
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
400
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


