十二个机器学习模型在预测儿童COVID-19死亡风险方面的比较性能:巴西的一项基于人口的回顾性队列研究
Adriano Lages Dos Santos1,2, Maria Christina L Oliveira2, Enrico A Colosimo3
1Engineering and Informatics, Federal Institute of Science and Technology of Minas Gerais, Belo Horizonte, Minas Gerais, Brazil.
PeerJ. Computer science
|June 26, 2025
概括
机器学习模型可以预测儿童COVID-19患者的死亡率. 后勤回归,梯度提升和AdaBoost显示出高准确度,确定氧和度降低,并发病症和年龄较大是关键风险因素.
科学领域:
- 公共卫生 公共卫生
- 生物医学信息学 生物医学信息学
- 儿科医学 儿科医学
背景情况:
- 随着COVID-19的爆发,人们越来越需要预测成年人死亡率的工具.
- 有限的研究存在于机器学习 (ML) 用于预测儿童COVID-19病例的结果.
- 这项研究解决了预测SARS-CoV-2感染的住院儿童和青少年死亡率的差距.
研究的目的:
- 评估多种机器学习模型在儿童COVID-19患者死亡率预测方面的性能.
- 确定与这一人口群中的死亡率相关的关键临床变量.
- 评估ML在儿童COVID-19护理临床决策中的实用性.
主要方法:
- 利用来自巴西的SIVEP-Gripe数据集,追踪严重急性呼吸系统综合征 (SARS).
- 在分区数据子集上开发和训练了12个机器学习算法,用于培训和测试.
- 使用准确度,精度,灵敏度,回忆和AUC评估模型性能,通过千平方测试确定显著的死亡率预测因素.
主要成果:
- 后勤回归 (LR) 以92.5%的准确度和80.1%的AUC显示了最高的性能.
- 渐变增强分类器 (GBC) 和AdaBoost (ADA) 也显示出强大的预测能力.
- 死亡率的关键预测因素包括儿科患者基线氧度降低,并发症和年龄较大.
结论:
- 机器学习模型,特别是LR,GBC和ADA,是预测儿童COVID-19死亡率的有效工具.
- 这些模型可以帮助临床决策和改善患者管理策略,以获得更好的结果.
- 识别低氧和和伴随疾病等高风险因素对于儿童COVID-19病例的及时干预至关重要.
相关概念视频
Relative Risk
366
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
366
Cancer Survival Analysis
458
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
458
Comparing the Survival Analysis of Two or More Groups
302
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
302
Bias in Epidemiological Studies
705
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
705
Steps in Outbreak Investigation
215
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
215
Residuals and Least-Squares Property
7.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.9K


