一个空间时空XGBoost模型用于PM2.5度预测及其在上海的应用
Zidong Wang1, Xianhua Wu1, You Wu1
1School of Economics and Management, Shanghai Maritime University, Shanghai 201306, China.
Heliyon
|December 7, 2023
概括
这项研究开发了一个先进的框架来预测上海的PM2.5水平,通过增强的XGBoost建模,提高了17%的预测准确度,并减少了28%的错误.
科学领域:
- 环境科学 环境科学
- 数据科学数据科学数据科学
- 大气化学 大气化学
背景情况:
- 颗粒物 (PM2.5) 污染对公众健康和环境构成重大风险.
- 准确预测PM2.5度对于有效的环境管理和政策实施至关重要.
- 现有的预测模型往往难以准确,特别是在极端事件和不同空间分布期间.
研究的目的:
- 开发和验证一个创新的分析和预测框架,用于预测上海16个地区的PM2.5度水平.
- 通过优化 XGBoost 模型并采用先进的分解和融合技术,提高 PM2.5 预测的准确性和稳定性.
- 调查上海PM2.5度的季节性和空间模式.
主要方法:
- 使用了增强的XGBoost模型,具有参数调整,经验模式分解 (特别是由SSA验证的EEMD) 和模型融合.
- 应用了混合型号组件的可变重量分配和通过IDW方法验证的Kriging插值.
- 将地理坐标转换为球形表面弧度,并使用层次分析来监测站的重要性.
主要成果:
- 与原始模型相比,修改后的XGBoost模型显示,适合度增加了17%,根平均平方误差减少了28%.
- 确定了对PM2.5变化的显著季节性影响,周期为3个月,冬季极端值频率更高.
- 与农村地区相比,上海市中心的PM2.5度较高,从中心向外的趋势下降.
结论:
- 开发的框架显著提高了PM2.5预测的准确性和稳定性,有效地避免了极端点的预测偏差.
- 上海的PM2.5度呈现出明显的季节和空间模式,冬季和城市中心地区的度更高.
- 该研究的创新方法论贡献提高了PM2.5预测模型在城市环境中的可靠性和适用性.
相关概念视频
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Sampling Plans
187
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
187
End Point Prediction: Gran Plot
337
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
337
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
72
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
72


