在西班牙SEMI-COVID-19注册表中使用机器学习改进预测COVID-19死亡率
José-Manuel Casas-Rojo1, Paula Sol Ventura2, Juan Miguel Antón Santos3
1Internal Medicine Department, Infanta Cristina University Hospital, Parla, 28981, Madrid, Spain.
机器学习准确地预测住院患者的COVID-19死亡率. 一个渐变增强决策树模型确定了关键指标,实现了对患者结果的高预测能力.
科学领域:
- 计算生物学是一种计算生物学.
- 医疗信息学医学信息学
- 流行病学 流行病学
背景情况:
- 由于COVID-19造成了显著的死亡率,因此需要改进预测工具.
- 现有的机器学习模型对于预测COVID-19死亡率是不够的.
- 准确预测死亡风险对于患者管理至关重要.
研究的目的:
- 开发和验证用于预测住院COVID-19患者死亡率的机器学习模型.
- 确定与COVID-19死亡率相关的关键临床和实验室指标.
- 使用渐变增强决策树 (GBDT) 进行可靠的死亡率预测.
主要方法:
- 利用西班牙的SEMI-COVID-19注册表 (24,514名患者) 进行模型开发.
- 使用CatBoost和BorutaShap分类器进行特征选择和模型生成.
- 使用时间分割验证了GBDT模型 (培训:2020年2月至2020年12月,测试:2021年1月至2021年11月).
主要成果:
- 分析了23,983名患者的临床和实验室数据.
- 在测试组中,CatBoost GBDT 模型实现了 84.76 (SD 0.45) 的 AUC.
- 16个特征模型显示了对COVID-19医院死亡率的高预测能力.
结论:
- 基于GBDT的机器学习模型有效预测住院COVID-19患者的死亡率.
- 该模型确定了重要的预测因素,为风险分层提供了有价值的见解.
- 这种方法为增强患者护理和资源配置提供了强大的工具.
更多相关视频
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
相关概念视频
Steps in Outbreak Investigation
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Statistical Methods for Analyzing Epidemiological Data
Improving Translational Accuracy
Cancer Survival Analysis
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
