Related Experiment Video
Updated: Jan 16, 2026

Use of Principal Components for Scaling Up Topographic Models to Map Soil Redistribution and Soil Organic Carbon
Published on: October 16, 2018
Machine learning-based land-use regression models for predicting carbon dioxide concentrations in San Francisco Bay
Anna C Smith1, Linfeng Li1, Jiansheng Xiang1
1Department of Earth Science and Engineering, Imperial College London, SW7 2AZ, London, United Kingdom.
None:
Carbon dioxide (CO2) is a key driver of anthropogenic climate change and cities have been identified as major sources of emissions. Urbanization and land use change are associated with rising urban CO2 emissions, highlighting the need to study spatiotemporal trends in intraurban CO2 to inform sustainable city planning. This study investigates the use of land use regression (LUR) to predict intraurban CO2 concentrations, using data from the BEACO2N monitoring network in the San Francisco Bay Area. Additionally, LUR is compared to machine learning (ML) algorithms capable of capturing non-linear relationships, representing a two-fold novel contribution. Model performance is evaluated using reserved data from training sensors as well as unseen sensor locations. For training sensors, extreme gradient boosting (XGBoost) and a convolutional neural network (CNN) achieved the highest predictive accuracy (R²=0.58), outperforming traditional LUR (R²=0.34). XGBoost and CNN also outperformed traditional LUR for unseen sensor locations, accounting for up to 42% of the variability in observed CO2 concentrations. These models offer insight into urban land use and carbon dynamics, supporting more informed approaches to urban planning and decarbonization.
Supplementary Information:
The online version contains supplementary material available at 10.1007/s12665-025-12582-w.
Related Concept Videos
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Design Example: Analyzing Capacity Contours for Flood Risk Assessment
Mechanistic Models: Compartment Models in Individual and Population Analysis
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...

