重复的时间序列交叉验证:一种新的方法来提高马来西亚的COVID-19预测准确性
Azlan Abdul Aziz1,2, Marina Yusoff3,4,5, Wan Fairos Wan Yaacob6
1College of Computing, Informatics and Mathematics, Universiti Teknologi MARA (UiTM) Cawangan Perlis, Arau 02600, Perlis, Malaysia.
MethodsX
|November 19, 2024
概括
准确的COVID-19病例预测对马来西亚至关重要. 这项研究引入了重复时间序列交叉验证,达到90.2%的平均8天预测准确度,以改善公共卫生决策.
科学领域:
- 流行病学 流行病学
- 数据科学数据科学数据科学
- 公共卫生 公共卫生
背景情况:
- 准确的COVID-19预测对于马来西亚有效的公共卫生政策和决策至关重要.
- 不准确的预测可能导致资源分配和控制策略不足于最佳.
- 可靠的预测支持政府机构和金融机构制定及时干预措施.
研究的目的:
- 提出一种新的数据分割策略,即重复时间序列交叉验证,用于识别优越的COVID-19预测模型.
- 为了提高马来西亚每日COVID-19病例预测的准确性.
- 使用已确定的指标,尽量减少预测错误.
主要方法:
- 对21种预测模型进行了全面评估.
- 用重复时间序列交叉验证策略对数据进行分区.
- 用四种错误指标 (RMSE,MAE,MAPE,MASE) 来评估模型的性能.
主要成果:
- 最好的模型在RMSE,MAE,MAPE和MASE中显示了最低的值.
- 平均8天预报准确率达到了90.2%.
- 个别预测准确度在83.7%至98.7%之间.
结论:
- 拟议的重复时间序列交叉验证方法有效地识别了高精度的COVID-19预测模型.
- 准确的预测是可以实现的,对马来西亚的流行病反应至关重要.
- 这种方法有助于当局做出明智的决定,以控制病毒的传播.
相关概念视频
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Steps in Outbreak Investigation
107
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
107
Sensitivity, Specificity, and Predicted Value
194
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
194
Statistical Methods for Analyzing Epidemiological Data
308
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
308
Improving Translational Accuracy
9.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.1K


