Related Experiment Video
Updated: May 28, 2025

10:29
Calibrated Passive Sampling - Multi-plot Field Measurements of NH3 Emissions with a Combination of Dynamic Tube Method and Passive Samplers
Published on: March 21, 2016
12.3K
Enhancing PM2.5 prediction by mitigating annual data drift using wrapped loss and neural networks.
Md Khalid Hossen1,2,3, Yan-Tsung Peng2, Meng Chang Chen3
1Social Networks and Human-Centered Computing, Taiwan International Graduate Program, Academia Sinca, Taipei, Taiwan.
Plos One
|February 11, 2025
Summary
This study addresses data drifting in deep learning by analyzing annual temperature data and proposing new models for PM2.5 prediction. The novel Front-loaded and Back-loaded connection models significantly improve prediction accuracy, outperforming traditional neural networks.
Area of Science:
- Environmental Science
- Data Science
- Meteorology
Background:
- Deep learning models often assume training data is from a single distribution, which is not always true for real-world data.
- Data collected over time or from different contexts can exhibit distribution shifts, impacting model performance.
- Annual temperature variations and air quality data (PM2.5) exemplify such context-dependent data drifting.
Purpose of the Study:
- To analyze data drifting phenomena in meteorological and air quality datasets.
- To develop and evaluate novel deep learning models for PM2.5 prediction that account for data drifting.
- To compare the performance of proposed models against existing deep learning architectures.
Main Methods:
- Utilized three statistical techniques to calculate P-values for identifying data drifting in annual station data (2014-2018).
- Proposed two models, Front-loaded connection (FLC) and Back-loaded connection (BLC), designed to handle data drifting characteristics.
- Incorporated a wrapped loss function to further enhance model training and accuracy.
- Evaluated model performance using Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE).
Main Results:
- Statistical analysis identified stations with significant data drifting.
- The proposed FLC and BLC models demonstrated superior performance in PM2.5 prediction compared to baseline BiLSTM and CNN models.
- Performance enhancements ranged from 24.1% to 16% and 12% to 8.3% for FLC/BLC against BiLSTM.
- Improvements against CNN were 24.6% to 11.8% and 10% to 10.2% for FLC/BLC.
Conclusions:
- Data drifting is a critical issue in deep learning for time-series data like PM2.5.
- The proposed FLC and BLC models effectively address data drifting, leading to more accurate PM2.5 predictions.
- The wrapped loss function further improves model robustness and accuracy in the presence of data drift.

