用机器学习预测COVID-19结果的本地和聚合数据培训策略的多中心比较分析
Carine Savalli1,2, Roberta Moreira Wichmann3, Fabiano Barcellos Filho2
1Federal University of São Paulo, Department of Public Politics and Public Health, Santos, Brazil.
预测COVID-19结果的本地机器学习 (ML) 模型,如重症监护室 (ICU) 入院和机械通风 (MV) 使用,在巴西大多数医院中比汇总数据策略更有效. 较小的医院可能会从聚合数据方法中受益.
科学领域:
- 临床信息学 临床信息学
- 人工智能在医学中的应用
- 流行病学 流行病学
背景情况:
- 机器学习 (ML) 显示出临床决策支持的潜力,特别是在资源有限的环境中.
- 从多家医院汇总数据可以增加样本大小,但可能会掩盖本地预测准确度.
- 数据聚合策略对不同临床环境中的ML模型性能的影响仍然不清楚.
研究的目的:
- 为了比较预测COVID-19结果的ML模型的本地与聚合数据训练策略的有效性.
- 评估巴西医院重症监护室 (ICU) 住院和机械通风 (MV) 使用的预测.
- 确定影响最佳数据聚合策略的因素.
主要方法:
- 在巴西14家医院的6,046名COVID-19患者的数据上训练了极端梯度增强,轻GBM和catboost模型.
- 将七个区域数据聚合策略与使用接收器操作特征曲线 (AUROC) 下面的区域的当地培训策略进行了比较.
- 利用变量重要性和元特征分析的SHAP值来探索战略选择驱动因素.
主要成果:
- 当地培训在79%的医院 (11/14) 优于ICU入院的综合策略,在71%的医院 (10/14) 优于MV使用的综合策略.
- 具有较小本地样本规模的医院倾向于使用聚合数据策略实现更好的表现.
- 超特征分析显示,医院的特征影响了选择最有效的数据策略.
结论:
- 当地培训策略在巴西医院预测COVID-19结果方面通常优越.
- 聚合数据方法可能对具有有限本地数据的较小医院有利.
- 当在各种医疗保健环境中实施预测性ML模型时,仔细考虑数据聚合至关重要.
更多相关视频
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
相关概念视频
Steps in Outbreak Investigation
Comparing the Survival Analysis of Two or More Groups
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Cancer Survival Analysis
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Statistical Software for Data Analysis and Clinical Trials
