相关实验视频
Updated: Sep 17, 2025

12:44
Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
8.1K
提高废水处理厂的能源消耗预测和可解释性:一个新的时间差权重重重新采样框架,用于不平衡回归的交叉验证
Kangrong Tang1, Anlei Wei1, Hanxiao Shi1
1Xi'an Key Laboratory of Environmental Simulation and Ecological Health in the Yellow River Basin, College of Urban and Environmental Sciences, Northwest University, Xi'an, 710127, China; Institute of Environmental Sciences, Northwest University, Xi'an, Shaanxi, 710127, China.
Journal of environmental management
|June 28, 2025
概括
本研究引入了一种新的机器学习 (ML) 框架,使用时间差权重重重新采样 (TDWR) 来改善废水处理厂 (WWTP) 的能源消耗预测. 随机低采样 (SUS) 方法显著提高了预测准确度,并减少了错误.
科学领域:
- 环境工程 环境工程
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 精确的能源消耗预测对于优化废水处理厂 (WWTP) 运营至关重要.
- 由于可变影响条件,在WWTP中常见的不平衡数据集通常会阻碍机器学习 (ML) 模型的性能.
- 现有的ML模型在预测能源使用时与固有的数据不平衡作斗争.
研究的目的:
- 开发和评估一种新的ML框架,以解决WWTP能源消耗预测中的不平衡回归问题.
- 介绍和比较三个时间差权重重复采样 (TDWR) 方法:门样本不足 (TUS),随机样本不足 (SUS) 和反向组图样本不足 (IHS).
- 提高ML模型的准确性和可解释性,以提高WWTP中智能低碳能源管理的准确性和可解释性.
主要方法:
- 提出了一个新的ML框架,使用三个TDWR方法 (TUS,SUS,IHS) 来处理不平衡回归.
- 员工内部验证 (80/20分) 和外部交叉测试,以进行可靠的模型评估.
- 使用XGBoost,支持向量回归,人工神经网络和随机森林模型,使用R2,RMSE,MAPE和残余分析进行评估. 为了解释性,使用了SHAP.
主要成果:
- 使用抽样因子为6的随机不足抽样 (SUS-6) 与XGBoost相结合,产生了优异的性能 (R2=0.9998,RMSE=0.0833,MAPE=0.14%).
- 与原始数据相比,TDWR显著提高了预测准确性:R2增加了高达27.6%,RMSE减少了约87%,MAPE减少了96.07%.
- 剩余分析显示95%的置信区间缩小了70%左右. 在其他ML模型中也出现了类似的改善.
结论:
- 提议的TDWR框架有效地解决了WWTP能源消耗预测中的数据不平衡,大大提高了ML模型的准确性.
- SUS-6表现出最好的表现,在预测指标和残余分布方面取得了实质性的改进.
- SHAP分析确定了影响能源使用的关键通风相关特征 (BOD,COD,NH3-N),为WWTP优化和低碳管理提供了可操作的见解.
相关概念视频
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K

