Related Experiment Video
Updated: Feb 22, 2026

Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
Making waves: Rethinking machine learning in wastewater effluent quality prediction through the overlooked roles of
Yijie Wang1, Damien Batstone2, Zhenju Sun3
1School of Civil and Environmental Engineering, Nanyang Technological University, 50 Nanyang Avenue 639798, Singapore; Nanyang Environment & Water Research Institute, Nanyang Technological University, 1 Cleantech Loop 637141, Singapore.
None:
Time-series machine learning (ML) approaches have been increasingly used to predict effluent quality in wastewater treatment plants (WWTPs), with a principal focus being on accuracy. However, as wastewater effluent quality is by nature autocorrelated, the interpretation of reported good ML performance could be overestimated by prediction target autocorrelation. Since standard performance metrics (e.g., R2) can yield high values even for simple persistence models, a re-evaluation using proper baselines and metrics is essential for model interpretations. This study addresses these gaps by highlighting the role of target autocorrelation and introducing the persistence model as a benchmark method. Autoregressive methods are evaluated using data from three published studies and two additional WWTP datasets. The aggregated SHAP analyses demonstrate that the importance of the historical target exceeds the second most influential parameter by 64 %-396 %, suggesting that a substantial portion of the reported "high accuracy" in short-term effluent quality prediction can be attributed to strong target autocorrelation rather than the effective learning of complex input-output relationships. Moreover, the persistence model frequently outperforms ML models in COD, TN, and TP predictions regardless of prediction horizon, especially in low-volatility scenarios. A novel index, PN-MAROC, is proposed to quantify data volatility and shows a strong correlation with performance of both persistence models (R2 = 0.93) and ML models (R2 = 0.75). This research highlights the need to consider autocorrelation and appropriate baselines in WWTP effluent quality predictions, and seeks to provide practical guidance for model interpretation and assessment.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Steps in Outbreak Investigation
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...