在空气质量预测模型中应用拉索正规化技术以缓解过度装配
Abbas Pak1, Abdullah Kaviani Rad2, Mohammad Javad Nematollahi3
1Department of Computer Sciences, Shahrekord University, Shahrekord, Iran.
Scientific reports
|January 2, 2025
概括
这项研究应用了拉索规范化来改善德黑兰的空气质量预测模型,成功地减少了颗粒物 (PM) 和气态污染物的过. 虽然对PM预测有效,但气态污染物预测由于其动态性而存在局限性.
科学领域:
- 环境科学与工程环境科学与工程
- 数据科学和机器学习
- 大气化学和物理大气化学和物理
背景情况:
- 空气污染对公共卫生和生态可持续性构成重大全球挑战.
- 机器学习 (ML) 模型越来越多地用于空气质量预测,但受到过度装配的影响.
- 过度装配降低了ML模型的有效性和通用性,需要规范化技术.
研究的目的:
- 通过应用最小绝对收缩和选择运营商 (Lasso) 规范化技术来提高空气质量预测模型的精度.
- 为了减轻过,并提高预测PM2.5,PM10,CO,NO2,SO2和O3度的模型的概括性.
- 用德黑兰的空气质量数据评估拉索在特征选择和模型性能改进方面的有效性.
主要方法:
- 利用来自伊朗德黑兰16个传感器的综合数据集,涵盖2013年至2023年的时间.
- 将拉索规范化技术应用于用于空气污染物度预测的机器学习模型.
- 使用R平方 (R2),平均绝对误差 (MAE),平均平方误差 (MSE),根平均平方误差 (RMSE) 和规范平均平方误差 (NMSE) 指数评估模型性能.
主要成果:
- 拉索规范化通过减少过拟合和确定颗粒物的主要预测特征 (PM2.5:R2=0.80,PM10:R2=0.75) 显著提高了模型可靠性.
- 气态污染物 (CO,NO2,SO2,O3) 的预测表现仍然不令人满意 (R2在0.35到0.55之间),这是由于它们的高活力和复杂的化学相互作用.
- 对于PM预测的强表现很可能是由于缺少的数据很少,而气态污染物模型面临着其性质和所选择的模型架构固有的挑战.
结论:
- 拉索规范化是一种高度有效的技术,用于减轻过度装配和选择空气质量预测模型中的重要特征.
- 该研究强调了与颗粒物相比,准确预测气体污染物的挑战,这是由于固有的大气复杂性.
- 拉索的成功应用表明它在开发强大可靠的空气质量预测系统方面具有广泛采用的强大潜力.
相关概念视频
Regression Toward the Mean
6.2K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.2K
Regression Analysis
5.4K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.4K
Residuals and Least-Squares Property
7.2K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.2K
Calibration Curves: Linear Least Squares
1.1K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
1.1K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Multiple Regression
2.8K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.8K


