对中国二氧化碳排放预测的统计和机器学习模型进行比较研究
Xiangqian Li1, Xiaoxiao Zhang2
1School of Statistics, Capital University of Economics and Business, Beijing, 100070, People's Republic of China.
Environmental science and pollution research international
|October 22, 2023
概括
长短期记忆 (LSTM) 模型在预测中国每日二氧化碳 (CO2) 排放方面表现出色,优于统计和其他机器学习方法. 这种先进的预测有助于制定有效的排放控制策略和可持续的未来.
科学领域:
- 环境科学 环境科学
- 气候变化研究 气候变化研究
- 数据科学数据科学数据科学
背景情况:
- 增加的二氧化碳 (CO2) 排放是全球变暖的主要驱动因素.
- 准确的二氧化碳排放预测对于有效的气候变化缓解战略至关重要.
- 接近实时的每日二氧化碳排放数据对于及时的政策干预至关重要.
研究的目的:
- 确定中国近乎实时的每日二氧化碳排放的最佳预测模型.
- 为了比较统计和机器学习模型对二氧化碳排放预测的性能.
- 使用平均平方误差,根平均平方误差,平均绝对误差,平均绝对百分比误差和确定系数来评估模型.
主要方法:
- 利用从2020年1月1日到2022年9月30日的单变量每日时间序列数据.
- 提出并评估了六种预测模型:灰色预测 (GM(1,1)),ARIMA,SARIMAX,人工神经网络 (ANN),随机森林 (RF) 和长短期记忆 (LSTM).
- 使用五个关键统计指标 (MSE,RMSE,MAE,MAPE,R2) 评估模型性能.
主要成果:
- 机器学习模型 (ANN,RF,LSTM) 的表现始终优于统计模型 (GM(1,1),ARIMA,SARIMAX).
- 在所有评估指标上,LSTM 模型表现出了卓越的性能.
- 在LSTM的测试中,MSE达到3.5179e-04,RMSE达到0.0187,MAE达到0.0140,MAPE达到14.8291%,R2达到0.9844.
结论:
- 在中国,LSTM模型是最适合近乎实时的每日二氧化碳排放预测的.
- 在捕捉复杂的排放模式方面,LSTM的稳定性支持其在排放预测中的应用.
- 来自LSTM的准确CO2预测可以为有效的缓解策略和环境政策提供数据驱动的决策信息.
相关概念视频
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Statistical Methods for Analyzing Epidemiological Data
385
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
385
Calculating and Interpreting the Linear Correlation Coefficient
6.0K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
6.0K
Mechanistic Models: Compartment Models in Individual and Population Analysis
45
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
45
Correlation and Regression
1.3K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.3K


