Related Experiment Video
Updated: Jan 16, 2026

Use of Principal Components for Scaling Up Topographic Models to Map Soil Redistribution and Soil Organic Carbon
Published on: October 16, 2018
Machine learning-based land-use regression models for predicting carbon dioxide concentrations in San Francisco Bay
Anna C Smith1, Linfeng Li1, Jiansheng Xiang1
1Department of Earth Science and Engineering, Imperial College London, SW7 2AZ, London, United Kingdom.
Machine learning models like XGBoost and CNN accurately predict urban carbon dioxide (CO2) emissions, outperforming traditional methods. This research supports sustainable city planning and decarbonization efforts by understanding CO2 dynamics.
Area of Science:
- Environmental Science
- Urban Planning
- Climate Change Research
Background:
- Cities are major sources of anthropogenic carbon dioxide (CO2) emissions, driving climate change.
- Urbanization and land use changes are linked to increasing CO2 emissions within cities.
- Understanding intraurban CO2 spatiotemporal trends is crucial for sustainable city planning.
Purpose of the Study:
- To investigate the efficacy of land use regression (LUR) for predicting intraurban CO2 concentrations.
- To compare LUR with machine learning (ML) algorithms for CO2 prediction.
- To evaluate model performance using both training and unseen sensor locations.
Main Methods:
- Utilized data from the BEACO2N monitoring network in the San Francisco Bay Area.
- Applied traditional land use regression (LUR) models.
- Employed machine learning algorithms, including extreme gradient boosting (XGBoost) and convolutional neural networks (CNN).
- Evaluated predictive accuracy using reserved training data and data from unseen sensor locations.
Main Results:
- For training sensors, XGBoost and CNN achieved the highest accuracy (R²=0.58), surpassing traditional LUR (R²=0.34).
- XGBoost and CNN also outperformed LUR for unseen sensor locations, explaining up to 42% of CO2 concentration variability.
- Machine learning models demonstrated superior ability in capturing non-linear relationships in urban CO2 data.
Conclusions:
- Advanced ML models like XGBoost and CNN are effective tools for predicting intraurban CO2 concentrations.
- These models provide valuable insights into urban land use and carbon dynamics.
- The findings support the development of informed urban planning and decarbonization strategies.
Related Concept Videos
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Design Example: Analyzing Capacity Contours for Flood Risk Assessment
Mechanistic Models: Compartment Models in Individual and Population Analysis
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Calibration Curves: Linear Least Squares
For data that follow a straight line, the standard method for fitting is the linear...

