Related Experiment Video
Updated: Aug 14, 2026

Real-World M3-BREATHE: Toward Multimodal Mobile Monitoring of Behaviour, Respiration, and Exposures for Treatment and Health Evaluation
Published on: June 5, 2026
Interpretable short-term PM2.5 forecasting using meteorology-pollution coupling across multiple Beijing monitoring
1Guangzhou College of Technology and Business, Guangzhou, China. wuyf.admin@gmail.com.
None:
Accurate short-term PM2.5 forecasting is important for urban environmental monitoring because it supports early warning and short-term emission-management decisions. Recent research has advanced air-quality prediction through statistical, machine-learning, deep-learning, and spatiotemporal architectures. Complementing these developments, this study evaluates leakage-safe temporal validation, cross-station transferability, and operational interpretability using the public Beijing Multi-Site Air Quality benchmark, comprising 420,768 hourly observations from 12 stations during 2013-2017. The predictor space integrates pollutant lags, rolling statistics, meteorological covariates, cyclic calendar encodings, trigonometric wind-direction components, and station indicators. Persistence, ridge regression, random forest, XGBoost, LightGBM, and CatBoost were evaluated at 1-, 6-, and 24-h horizons under a strict chronological train-validation-test design, together with Diebold-Mariano testing, feature-group ablation, and station-holdout transfer evaluation. The 6-h horizon was treated as the principal operational setting because it lies between the persistence-dominated 1-h task and the more uncertain 24-h task. CatBoost achieved the best 6-h performance (RMSE = 51.24 µg/m3, MAE = 30.87 µg/m3, R2 = 0.617), reducing RMSE by 6.73% relative to persistence; the improvement was statistically significant. In the station-holdout experiment, LightGBM achieved an RMSE of 53.11 µg/m3, 5.65% lower than persistence. Feature analyses identified recent PM2.5 history, wind speed and direction, diurnal cyclicity, and the moisture-related temperature-dew-point gap as the dominant predictors. Performance nevertheless deteriorated during high-pollution episodes, indicating that unresolved episodic emissions and atmospheric processes remain difficult to capture. Because the data represent a historical pollution regime, contemporary deployment would require retraining, recalibration, or domain adaptation. The framework provides a reproducible and interpretable benchmark for future model comparisons under equally rigorous chronological evaluation.
