全人口抑郁症发病率预测比较自回归集成移动平均线和向量自回归集成移动平均线与时间融合变压器:纵向观测研究
Deliang Yang1, Yiyi Tang1,2, Vivien Kin Yi Chan3
1Department of Medicine, School of Clinical Medicine, Li Ka Shing Faculty of Medicine, The University of Hong Kong, Hong Kong, China (Hong Kong).
Journal of medical Internet research
|May 12, 2025
概括
时间融合变压器 (TFT) 在预测抑郁症发病率方面优于ARIMA和VARIMA模型,特别是在公共卫生危机期间. TFT表现出对突然变化的稳定性,为心理健康负担预测提供了更高的准确性.
科学领域:
- 流行病学 流行病学
- 时间序列分析时间序列分析
- 心理健康研究 心理健康研究
背景情况:
- 准确预测全民抑郁症发病率对于公共心理健康管理至关重要.
- 社会经济因素和流行病等突发事件在时间序列数据中造成了复杂的结构性断裂,挑战了预测准确性.
- 了解各种结构性破坏场景中的模型性能对于可靠的公共卫生预测至关重要.
研究的目的:
- 开发和比较使用自回归集成移动平均线 (ARIMA),矢量-ARIMA (VARIMA) 和时间融合变压器 (TFT) 的抑郁症发病率预测模型.
- 评估这些模型在各种结构性崩场景下的表现.
- 为在公共卫生中选择合适的预测工具提供见解.
主要方法:
- 来自香港 (2002-2022) 的每月抑郁症发病率数据使用滑动窗进行分析,以创建72个十年子样本.
- 预测模型 (ARIMA,VARIMA,单变量TFT,多变量TFT) 在每个子样本上进行了训练,验证和测试.
- 在测试组上使用对称平均绝对百分比误差 (SMAPE) 测量模型准确性.
主要成果:
- 在稳定的时期,多变量TFT显著优于单变量TFT,VARIMA和ARIMA (平均SMAPE11.6%对13.2%,16.4%,14.8%).
- 纳入失业率数据比VARIMA更多地改善了TFT业绩.
- 在疫情爆发期间,TFT对急剧中断的强度更高,而VARIMA和ARIMA在持续的发病激增期间表现更好.
结论:
- TFT模型为预测抑郁症发病率提供了更合适的方法,特别是在负担稳定或突然中断的时期.
- 该研究为选择基于数据特征和社会破坏性质的预测模型提供了一个框架.
- 研究结果支持使用像TFT这样的先进模型来更准确地预测公共卫生负担,特别是在公共卫生危机期间.
相关概念视频
Estimating Population Standard Deviation
3.0K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.0K
Distributions to Estimate Population Parameter
4.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.0K
Steps in Outbreak Investigation
95
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
95
Statistical Methods for Analyzing Epidemiological Data
240
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
240
Estimating Population Mean with Unknown Standard Deviation
7.5K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
7.5K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K


