Related Experiment Video
Updated: Sep 26, 2025

Analysis of the Ambient Particulate Matter-induced Chromosomal Aberrations Using an In Vitro System
Published on: December 21, 2016
A hybrid satellite and land use regression model of source-specific PM2.5 and PM2.5 constituents
Md Mostafijur Rahman1, George Thurston2
1Department of Environmental Medicine, New York University Grossman School of Medicine, New York, NY 10010, United States.
Abstract:
Although PM2.5 mass varies in source and composition over time and space, most health effects assessment have made the inherent assumption that all PM2.5 mass has the same health implications, irrespective of composition. Nationwide estimates of source-specific PM2.5 mass and constituents at local-scale would allow for epidemiological studies and health effects assessments that consider the variability in PM2.5 characteristics in their health impact assessments. In response, we developed US models of annual exposures at the census tract level for five major PM2.5 sources (traffic, soil, coal, oil, and biomass combustion) and six trace elements (elemental carbon, sulfur, silicon, selenium, nickel, and non-soil potassium) for 2001 through 2014. We employed Absolute Factor Analysis (APCA) to derive the source-specific PM2.5 impacts at monitoring stations. Random forest algorithms that incorporated predictors derived from satellite, chemical transport model, and census tract resolution land-use data on traffic, meteorology, and emissions, which were rigorously tested by 10-fold cross-validation (CV), were then employed to estimate elemental and source-specific PM2.5 levels at non-monitoring site census-tracts over the study years. Model performances were moderate to good, with CV R2 ranging from 0.41 to 0.95. For PM2.5 sources, the highest CV R2 was attained for traffic PM2.5 (CV R2 = 0.73), followed by coal (CV R2 = 0.65), oil (CV R2 = 0.62), soil (CV R2 = 0.60), and biomass (CV R2 = 0.41). Among constituents, the CV was highest for sulfur (CV R2 = 0.95). Our analyses provided highly resolved spatial estimates of annual elemental and source-specific PM2.5 concentrations at the census-tract level, for 2001 through 2014. This dataset offers exposure estimates in support of future nationwide long-term health effects studies of source-specific PM2.5 mass and constituents, enabling epidemiological research that addresses the fact that not all particles are the same.
More Related Videos
12:26Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM
Published on: October 11, 2016
09:44Use of Principal Components for Scaling Up Topographic Models to Map Soil Redistribution and Soil Organic Carbon
Published on: October 16, 2018
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as: