在不平衡的空气污染数据中,使用具有相关性加权 (MBB-RW) 的移动块引导方法对极端值进行修改的重新采样策略
Mahiran Muhammad1, Ahmad Zia Ul-Saufie2, Noor Fadhilah Ahmad Radi1
1Faculty of Computer and Mathematical Sciences, Universiti Teknologi MARA, Shah Alam, 40450, Selangor, Malaysia.
Scientific reports
|December 12, 2025
概括
一个新的重新采样策略,即移动块引导与相关权重 (MBB-RW),有效地解决了空气污染数据的不平衡. 这种方法可以提高极端事件的预测准确度,这对于公共卫生和环境监测至关重要.
科学领域:
- 环境科学 环境科学
- 数据科学数据科学数据科学
- 统计建模 统计建模
背景情况:
- 空气污染数据集经常显示由于极端事件导致右倾分布,导致数据不平衡.
- 不平衡的回归模型在极端事件中扎,显示对正常条件的偏差,并在关键污染事件中表现不佳.
研究的目的:
- 为解决空气污染数据不平衡回归提出一个新的重新采样策略,即移动块引导与相关权重 (MBB-RW),以解决空气污染数据的不平衡回归问题.
- 提高模型的预测准确度,特别是对于极端空气污染事件.
主要方法:
- 开发了一个修改后的重新采样策略,MBB-RW,集成移动块引导与相关性加权.
- 保持时间序列依赖性,同时强调空气污染数据集中的极端数据点.
- 将MBB-RW应用于极端梯度提升 (XGBoost) 模型进行PM$_{10}$预测.
主要成果:
- MBB-RW有效地缓解了数据不平衡,并改善了极端事件的模型性能.
- 在PM$_{10}$预测中,根平均平方误差 (RMSE) 从108.3010降至39.1846 (63.8188%) 和平均绝对误差 (MAE) 从85.1041降至27.1082 (68.14700%) 显著降低.
- 在极端空气污染事件期间预测PM$_{10}$度的准确性提高.
结论:
- 该MBB-RW重新采样策略是改善不平衡回归数据集中的极端值预测的关键贡献,特别是对于空气质量数据.
- 这种方法提高了PM$_{10}$度预测的准确性,为管理极端空气污染事件提供了宝贵的见解.
相关概念视频
Bootstrapping
787
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
787
Sampling Plans
861
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
861
Quantifying and Rejecting Outliers: The Grubbs Test
3.4K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.4K
Strategies for Assessing and Addressing Confounding
336
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
336
Cluster Sampling Method
13.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.9K
Weighted Mean
6.2K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
6.2K


