基于经验模式分解的时间序列预测中的信息泄露研究
Xinyi Yang1,2, Jingyi Li1, Xuchu Jiang3,4
1School of Statistics and Mathematics, Zhongnan University of Economics and Law, Wuhan, 430073, China.
Scientific reports
|November 16, 2024
概括
本研究引入了三种新的分解方法,以防止时间序列预测中的信息泄露,显著提高了水质预测的准确性. 新方法通过解决当前分解技术中常见的"过度装配"问题来增强预测模型.
科学领域:
- 环境科学 环境科学
- 数据科学数据科学数据科学
- 工程 工程师 工程师 工程师
背景情况:
- 时间序列分析对于预测各种领域的未来趋势至关重要,包括金融,经济和环境监测.
- 传统的分解-整合-预测策略经常遭受"信息泄漏",测试集数据影响分解,导致虚幻的准确性.
- 现有的方法在模型训练和测试期间与非静止序列和过拟合作斗争.
研究的目的:
- 提出和评估三个改进的分解策略:滑动窗体实证模式分解 (SW-EMD),单次训练和多次分解 (STMP-EMD) 和多次训练和多次分解 (MTMP-EMD).
- 为解决与当前时间序列预测分解技术固有的"信息泄漏"问题.
- 提高水质预测模型的准确性和可靠性.
主要方法:
- 开发了三个改进的类EMD分解方法:SW-EMD,STMP-EMD和MTMP-EMD.
- 将这些方法与CEEMDAN-MSBTCN-BiLSTM-DMAttention模型结构集成,其中包含基于等号相似性的依赖矩阵.
- 应用了新型模型来预测水质指标:pH,溶解氧 (DO) 和甲 (KMnO4).
主要成果:
- 提出的模型在预测pH,DO和KMnO4水平方面表现良好.
- 与主流LSTM模型相比,精度提高了1.958% (RMSE) 和0.853% (MAPE).
- 改进的分解方法有效地缓解了"信息泄漏",并减少了模型过度装配.
结论:
- 提出的三个分解改进策略成功地解决了时间序列预测中的"信息泄露"问题.
- 创新的CEEMDAN-MSBTCN-BiLSTM-DMAttention模型结构,结合改进的分解,为水质预测提供了一个强大的解决方案.
- 这项研究为水质预测提供了有效的实验框架,提高了模型的可靠性,并解决了过度拟合的挑战.
相关概念视频
What is a Mode?
18.1K
The mode is one of the commonly used measures of a central tendency. It is defined as the most frequent value in a data set.
There can be more than one mode in a data set if multiple values have the same highest frequency. For instance, suppose that the Statistics exam scores of 20 students are: 50; 53; 59; 59; 63; 63; 72; 72; 72; 72; 72; 76; 78; 81; 83; 84; 84; 84; 90; 93. Here, the mode is 72, as it occurs most frequently, five times.
A data set with two modes is called bimodal. For example,...
There can be more than one mode in a data set if multiple values have the same highest frequency. For instance, suppose that the Statistics exam scores of 20 students are: 50; 53; 59; 59; 63; 63; 72; 72; 72; 72; 72; 76; 78; 81; 83; 84; 84; 84; 90; 93. Here, the mode is 72, as it occurs most frequently, five times.
A data set with two modes is called bimodal. For example,...
18.1K
Skewness
10.9K
The measures of central tendency calculated from a data set may not reveal much about its intrinsic distribution. If a plot is made of the data set’s values, the mean and the median may not only differ, but also the plot may have more values on one side of the central tendencies. Such a data set is said to be skewed towards that side.
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
10.9K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Noncompartmental Analysis: Statistical Moment Theory
89
Noncompartmental analyses leverage statistical moment theory to examine time-related changes in macroscopic events, encapsulating the collective outcomes stemming from the constituent elements in play. Statistical moment theory is a mathematical approach used to describe the time course of drug concentration in the body without assuming a specific compartmental model. SMT provides insights into drug absorption, distribution, metabolism, and elimination by treating drug concentration versus time...
89
Random Error
832
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
832
Propagation of Uncertainty from Systematic Error
484
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
484


