相关实验视频
Updated: Feb 22, 2026

12:44
Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
8.7K
制造浪潮:通过自我关联和基线模型的被忽视的作用,重新思考废水废水质量预测中的机器学习
Yijie Wang1, Damien Batstone2, Zhenju Sun3
1School of Civil and Environmental Engineering, Nanyang Technological University, 50 Nanyang Avenue 639798, Singapore; Nanyang Environment & Water Research Institute, Nanyang Technological University, 1 Cleantech Loop 637141, Singapore.
Water research
|February 20, 2026
概括
对于污水处理厂废水质量预测的机器学习 (ML) 模型可能会因为目标自相关性而高估准确性. 持久性模型通常会超过ML,特别是在低波动性条件下,突出了更好的基线的需要.
科学领域:
- 环境工程环境工程
- 机器学习应用 机器学习应用
- 废水处理 废水处理
背景情况:
- 时间序列机器学习 (ML) 被广泛用于预测废水处理厂 (WWTP) 废水质量,并强调预测准确度.
- 废水质量数据表现出固有的自相关性,这可能会膨胀ML模型的感知性能.
- 像R平方这样的标准指标可能会误导,因为简单的持久性模型可能会达到高分,需要重新评估适当的基线.
研究的目的:
- 通过考虑目标自相关性,重新评估ML模型对WWTP废水质量预测的性能.
- 引入和验证持久性模型作为评估ML模型解释性和准确性的关键基准.
- 提出一种新的指数来量化数据波动及其对模型性能的影响.
主要方法:
- 用三项已发表的研究和两个额外的WWTP数据集的数据评估自回归方法.
- 综合SHAP分析的应用,以量化历史目标值对预测的影响.
- 将ML模型的性能与不同预测时间和波动性场景的持久性模型基准进行比较.
主要成果:
- SHAP分析显示,历史目标值显著主导其他参数 (64%-396%更重要),表明报告的高准确性可能来自自相关性.
- 持久性模型在预测化学氧气需求 (COD),总 (TN) 和总 (TP) 方面经常超过ML模型,特别是在低挥发性场景中.
- 一个新的指数PN-MAROC有效量化数据波动,并与持久性 (R2 = 0.93) 和ML模型 (R2 = 0.75) 性能强烈相关.
结论:
- 该研究强调了在解释ML模型在WWTP废水质量预测中的性能时考虑目标自相关性和采用强大的基线的关键需要.
- 持久性模型作为一个重要的基准,往往优于复杂的ML模型,特别是当废水质量表现出低波动时.
- 拟议的PN-MAROC指数为评估数据特征提供了一个实用的工具,并指导废水管理中的预测模型的选择和解释.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
290
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
290
Multiple Regression
4.1K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.1K
Steps in Outbreak Investigation
622
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
622
Residuals and Least-Squares Property
9.6K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.6K