数据归算对空气质量预测问题的影响
Van Hua1,2,3, Thu Nguyen4, Minh-Son Dao5
1Faculty of Mathematics and Computer Science, University of Science, Ho Chi Minh City, Vietnam.
PloS one
|September 12, 2024
概括
准确的空气质量预测对于环境健康至关重要. 本研究回顾了空气质量数据时间序列的缺失数据归算技术,评估了八种方法来指导从业者.
科学领域:
- 环境科学 环境科学
- 数据科学数据科学数据科学
- 时间序列分析时间序列分析
背景情况:
- 准确的空气质量预测对于公共卫生和环境政策至关重要.
- 空气质量数据通常基于时间序列,经常包含缺失值.
- 处理缺失数据对于可靠的空气质量分析和预测至关重要.
研究的目的:
- 提供对空气质量预测和缺失数据归算技术的全面审查,用于时间序列数据.
- 在不同的空气质量数据集上实证评估八种不同的归算方法的性能.
- 为环境环境中处理缺少时间序列数据提供实用建议.
主要方法:
- 对空气质量预测和时间序列归算方法的文献综述.
- 对八种归算技术的实证评估:平均值,中位数,kNNI,MICE,SAITS,BRITS,MRNN和变压器.
- 使用各种全球空气质量数据集进行评估.
主要成果:
- 该研究系统地比较了不同归算方法对空气质量时间序列数据的有效性.
- 在不同的数据集中分析了归算技术之间的性能变化.
- 确定适用于解决环境数据中缺失值的方法.
结论:
- 归算方法的选择显著影响空气质量数据分析和预测准确度.
- 经验评估提供了对各种技术的优缺点的见解.
- 提供了建议,以帮助从业人员选择环境时间序列数据的适当归算策略.
相关概念视频
Steps in Outbreak Investigation
114
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
114
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Random Error
848
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
848
Measurement of Air Content in Concrete
123
Air content measurement in concrete is critical for ensuring structural integrity and durability of concrete structures, especially in environments prone to severe weather conditions. Accurate air content analysis optimizes concrete's resistance to freeze-thaw cycles and enhances its workability and strength. Several methods are standardized under ASTM guidelines to measure the air content in fresh concrete, each suitable for different concrete types and conditions.
The pressure method,...
The pressure method,...
123
Detection of Gross Error: The Q Test
5.7K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.7K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
117
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
117


