Related Experiment Video
Updated: Jun 14, 2025

Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
Machine learning framework for energy consumption prediction in wastewater treatment plants combining resampling and
Kangrong Tang1, Anlei Wei2, Zixuan Wang1
1Xi'an Key Laboratory of Environmental Simulation and Ecological Health in the Yellow River Basin, College of Urban and Environmental Sciences, Northwest University, Xi'an 710127, China.
Abstract:
This study aims to enhance energy consumption prediction in wastewater treatment plants (WWTPs) by developing a robust machine learning framework. To address data imbalance and poor interpretability, the framework integrates three resampling methods with cohort Shapley additive explanations (SHAP). Specifically, stochastic under-sampling with sampling weights to the power of 3 (SUS-3) is combined with a random forest model, which achieves the best performance, yielding a determination coefficient of 0.928 and a root-mean-square error of 4.255 in predicting total energy consumption. In unit-energy-consumption classification, SUS-3 improves accuracy and precision by over 30% and reduces model uncertainty by 35%. Cohort-SHAP analyses reveal stable ammonia-nitrogen and biochemical oxygen demand aeration energy impacts, with season/load-specific energy patterns of chemical oxygen demand, suspended solids, total phosphorus, total nitrogen, and inflow rate, peaking in spring/summer and emphasizing phosphorus in winter. This framework offers a transferable, interpretable solution necessitating adaptive strategies to optimize energy use in WWTPs.
Related Concept Videos
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

