Related Experiment Video
Updated: Jul 2, 2026

Measuring Sub-23 Nanometer Real Driving Particle Number Emissions Using the Portable DownToTen Sampling System
Published on: May 22, 2020
Comparison of High Spatial Resolution PM2.5, PM10, and NO2 Estimates Using a Deep Ensemble Machine Learning Framework
Christine T Cowie1,2,3, Ivan C Hanigan4,3,5, Wenhua Yu6
1Woolcock Institute of Medical Research, Macquarie University, Macquarie Park, New South Wales 2113, Australia.
Abstract:
Recent studies report improved performance of ensemble models over machine learning (ML) models for air pollution estimation, although there is little evidence of their added value in settings with sparsely monitored data. We developed and compared three ML models, a supervised linear regression (SLR) model, and an ensemble model, for estimating annual average PM2.5, PM10, and NO2 in NSW, Australia, a relatively sparsely monitored, low pollution, and large geographic region. We assembled pollutant data from government and project monitors and data on 236 predictors including land use, population, traffic, and satellite observations. We used a three-stage DEML framework: (1) base-ML models (Random Forest (RF), XGBoost, GBM) and a SLR model; (2) meta-learner (RF, XGBoost, GBM, GLMNet); and (3) ensemble. We reserved 10% of data for hold-out validation and conducted 10-fold Cross validation (CV) using the training data sets. We evaluated the models using CV-R2 and RMSE. DEML models resulted in the best fit for all pollutants; however, improvements over base ML models were modest, indicating the latter, with lower cost of implementation, are valuable for low pollution settings with heterogeneous monitoring density. Choice of CV methods substantially impacted model performance and should be considered, along with setting constraints, when choosing modeling methods.
Related Concept Videos
Linear Approximations
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...