Machine learning driven by environmental covariates to estimate high-resolution PM2.5 in data-poor regions
XiaoYe Jin1,2, Jianli Ding1,2,3, Xiangyu Ge1,2
1Department of MOE Key Laboratory of Oasis Ecology, Xinjiang University, Urumqi, China.
Peerj
|April 5, 2022
Summary
This study maps fine particulate matter (PM2.5) air pollution in Xinjiang using advanced modeling. Results reveal high PM2.5 in southern Xinjiang, particularly the Tarim Basin, with winter peaks, offering a solution for data-scarce regions.
Area of Science:
- Environmental Science
- Atmospheric Science
- Public Health
Background:
- Particulate Matter (PM2.5) affects air quality and public health.
- Spatial distribution of PM2.5 is poorly understood in data-scarce regions.
- Xinjiang faces challenges in PM2.5 monitoring due to limited stations.
Purpose of the Study:
- To estimate the spatial distribution of PM2.5 concentrations in Xinjiang from 2015-2020 at 1 km resolution.
- To compare the performance of Random Forest (RF) and bagging algorithm models for PM2.5 estimation.
- To identify spatial and seasonal patterns of PM2.5 in Xinjiang.
Main Methods:
- Developed and validated Random Forest (RF) and bagging algorithm models.
- Utilized ground-monitored PM2.5 data, aerosol optical depth (AOD), meteorological data, and geographical variables.
- Employed 10-fold cross-validation (CV) for model verification.
Main Results:
- The RF model demonstrated superior performance for high-resolution PM2.5 estimation.
- PM2.5 concentrations were higher in southern Xinjiang (Tarim Basin) and lower in northern Xinjiang.
- Significant seasonality was observed, with winter concentrations highest (71.95 µg m⁻³) and summer lowest (43.40 µg m⁻³).
Conclusions:
- The RF model effectively estimates PM2.5 in data-scarce regions.
- The study provides crucial insights into Xinjiang's air quality patterns.
- This approach supports air quality monitoring and sustainable development efforts.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
89
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
89
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
109
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
109
Sampling Plans
305
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
305


