A data-driven supervised machine learning approach to estimating global ambient air pollution concentrations with
Liam Jordan Berrisford1,2, Hugo Barbosa3, Ronaldo Menezes4,5
1Department of Mathematics, University of Exeter, Exeter, UK.
Royal Society Open Science
|July 25, 2025
Summary
A new machine learning framework fills gaps in global air pollution data, providing comprehensive hourly estimates for nitrogen dioxide (NO2), ozone (O3), and particulate matter (PM2.5, PM10). This empowers detailed environmental and health studies worldwide.
Area of Science:
- Environmental Science
- Data Science
- Atmospheric Chemistry
Background:
- Global air pollution monitoring relies on sparse, unevenly distributed stations.
- Temporal data gaps are common due to technical and power issues.
- Accurate, comprehensive air quality data is crucial for public health and environmental assessments.
Purpose of the Study:
- To develop a scalable, data-driven machine learning framework to address data gaps in air pollution monitoring.
- To generate a comprehensive global dataset of key air pollutants (NO2, O3, PM10, PM2.5, SO2) with high spatial and temporal resolution.
- To provide prediction intervals for each estimate to quantify uncertainty.
Main Methods:
- Developed a supervised machine learning framework for imputing missing air quality data.
- Trained models to estimate concentrations of NO2, O3, PM10, PM2.5, and SO2.
- Generated global concentration estimations at 261,377 locations with 0.25° spatial resolution and hourly intervals.
Main Results:
- Created a comprehensive global dataset of air pollutant concentrations.
- Provided hourly estimates with prediction intervals for over 260,000 locations worldwide.
- Examined model performance across diverse geographical regions.
Conclusions:
- The machine learning framework effectively imputes missing air quality data, creating a valuable resource.
- The high-resolution dataset supports detailed downstream assessments for various stakeholders.
- Analysis of performance provides insights for optimizing future air quality monitoring station placement.
More Related Videos
Related Concept Videos
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Steps in Outbreak Investigation
207
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
207
Mechanistic Models: Compartment Models in Individual and Population Analysis
87
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
87
Sampling Plans
274
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
274
Regression Analysis
6.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.0K


