Related Experiment Video
Updated: Oct 5, 2025

07:14
Automated, High-resolution Mobile Collection System for the Nitrogen Isotopic Analysis of NOx
Published on: December 20, 2016
11.8K
Scalable penalized spatiotemporal land-use regression for ground-level nitrogen dioxide
Kyle P Messier1, Matthias Katzfuss2
1National Toxicology Program, National Institute of Environmental Health Sciences.
The Annals of Applied Statistics
|January 24, 2022
Summary
This study introduces a new method for accurately mapping daily nitrogen dioxide (NO2) air pollution across the US. The approach improves predictions for health and environmental risk assessments.
Area of Science:
- Environmental Science
- Public Health
- Statistical Modeling
Background:
- Nitrogen dioxide (NO2) is a major air pollutant from traffic with known health risks.
- Accurate spatiotemporal NO2 data is crucial for assessing exposure and health impacts.
- Existing land-use regression (LUR) models have limitations in handling complex spatial correlations.
Purpose of the Study:
- To develop a scalable and accurate method for estimating daily, nationwide NO2 concentrations.
- To improve variable selection and estimation in LUR models for air pollution.
- To provide reliable data for epidemiological and risk assessment studies.
Main Methods:
- Developed a scalable approach combining Gaussian-process approximation with penalized LUR coefficients.
- Implemented simultaneous variable selection and estimation for spatiotemporally correlated errors.
- Validated the method using simulated data and applied it to daily US-wide NO2 data.
Main Results:
- The new approach demonstrated superior model selection (specificity and sensitivity) and prediction accuracy (calibration and sharpness) compared to existing methods.
- Applied to US NO2 data, the method yielded more accurate, sparser, and interpretable models.
- Generated daily NO2 predictions revealing significant spatial variations within and between US cities.
Conclusions:
- The developed method offers a significant advancement in modeling air pollution spatiotemporal distribution.
- The accurate and interpretable NO2 predictions are valuable for public health research and risk assessment.
- This approach facilitates national-scale, daily exposure assessments for epidemiological studies.
Keywords:
Gaussian processKrigingair pollutiongeneral Vecchia approximationspatial statisticsvariable selectionMore Related Videos
Related Concept Videos
Levels of Use of a GIS
113
Geographic Information Systems (GIS) operate across three levels of application, each representing an increasing degree of complexity: data management, analysis, and prediction. These levels reflect the expanding functionality and versatility of GIS technology in handling spatial data for diverse purposes.Data ManagementAt its foundational level, GIS serves as a tool for data management, enabling the input, storage, retrieval, and organization of spatial data. This level is often employed in...
113
Calibration Curves: Linear Least Squares
2.9K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
2.9K
Regression Analysis
6.3K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
6.3K
Residuals and Least-Squares Property
8.0K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.0K
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Sampling Plans
316
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
316

