对预测孟加拉国登革热疫情的多种机器学习方法进行比较评估
Bowen Liu1,2, Md Farhad Hossain3,4, Shaheed Hossain5
1Division of Computing, Analytics, and Mathematics, Department of Mathematics and Statistics, School of Science and Engineering, University of Missouri - Kansas City, Kansas, MO, 64110, USA.
Scientific reports
|October 14, 2025
概括
机器学习模型准确地预测了孟加拉国的登革热发病率. XGBoost在预测每月登革热病例方面表现出卓越的表现,有助于公共卫生规划.
科学领域:
- 流行病学和公共卫生.
- 数据科学和机器学习
- 传染病建模传染病建模
背景情况:
- 登革热在孟加拉国是一个重大的公共卫生挑战.
- 准确预测登革热发病率对于有效的资源分配和干预策略至关重要.
- 在资源有限的环境中,有限的监控数据需要强大的预测建模方法.
研究的目的:
- 预测孟加拉国五个地区每月的登革热发病率.
- 将各种机器学习技术的性能与传统时间序列模型进行比较.
- 利用现有的监测数据,确定登革热预测最有效的模型.
主要方法:
- 利用了2022年1月至2023年12月的登革热监测数据.
- 应用并比较了季节性自回归集成移动平均 (SARIMA),多层感知器 (MLP),XGBoost和支向量回归 (SVR) 模型.
- 使用根平均平方误差 (RMSE),平均绝对误差 (MAE) 和平均绝对百分比误差 (MAPE) 评估模型性能.
主要成果:
- XGBoost成为最准确的预测模型,显示了最低的整体错误指标.
- 对于达卡部门,XGBoost实现了109%的RMSE,127%的MAE和12.9%的MAPE.
- 萨里马模型显示了合理的准确性,但被机器学习方法所超越;SVR不适合这个数据集.
结论:
- 先进的机器学习模型,特别是XGBoost,对于预测登革热等传染病非常有效,即使数据有限.
- 开发的预测模型可以显著支持孟加拉国的公共卫生规划和干预工作.
- 这项研究强调了数据驱动方法在资源有限的环境中管理传染病爆发的潜力.
相关概念视频
Steps in Outbreak Investigation
488
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
488
Statistical Methods for Analyzing Epidemiological Data
896
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
896
