Evaluating drivers of PM2.5 air pollution at urban scales using interpretable machine learning
Yali Hou1, Qunwei Wang2, Tao Tan3
1College of Information Engineering, Nanjing Xiaozhuang University, Nanjing 211171, China.
Waste Management (New York, N.Y.)
|December 2, 2024
Summary
China
Area of Science:
- Environmental Science
- Data Science
- Urban Planning
Background:
- Urban fine particulate matter (PM2.5) pollution hinders China's Sustainable Development Goals.
- Identifying PM2.5 drivers is crucial for effective pollution reduction strategies.
Purpose of the Study:
- Develop and validate a machine learning model to analyze urban PM2.5 concentrations.
- Identify key socioeconomic and industrial drivers of PM2.5 pollution.
- Propose city-specific strategies for PM2.5 reduction.
Main Methods:
- Utilized a combined CatBoost and Tree-Structured Parzen Estimator (TPE) machine learning model.
- Employed SHapley Additive exPlanations (SHAP) to determine factor importance.
- Analyzed PM2.5 data from 297 Chinese cities between 2000 and 2021.
Main Results:
- The model achieved high accuracy (R² = 96.44%) in predicting urban PM2.5.
- Socioeconomic factors and industrial activities were identified as primary drivers.
- Nitrogen oxide emissions, technology investment, population density, and energy consumption were significant contributors in 2000.
Conclusions:
- Targeted strategies, focusing on population density and mining development, are recommended for future pollution control.
- The framework provides a robust tool for evaluating pollution factors and tailoring reduction strategies.
- Significant improvements in air quality were observed due to reduced nitrogen oxide emissions.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
27
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
27
Sampling Plans
167
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
167


