Related Experiment Video
Updated: May 5, 2026

Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
Tackling data scarcity in wastewater treatment resulting from time-delay: A semi-supervised collaborative MTN-MRR
Jing Wu1, Jinwei Zhou1, Lili Tang2
1School of Data Science and Information Engineering, Guizhou Minzu University, Guiyang 550025, China.
None:
Accurate prediction of quality parameters is essential for the efficient operation of wastewater treatment plants (WWTPs). However, in real-world WWTPs, measurements of key quality parameters often experience significant delays due to the lags associated with laboratory analysis. Compared to variables that can be measured in real-time, this time delay results in a scarcity of data for crucial effluent parameters such as BOD5. To address these issues, this paper proposes a novel semi-supervised co-training framework with batch consistency confidence designed for soft sensor modeling in data-scarce environments. The framework begins with an Exponential Moving Average (EMA)-based preprocessing module that can denoise and purify raw input data to improve training reliability. Subsequently, a Multi-Scale Temporal Network (MTN), which uses adaptive spatio-temporal convolutions to extract hierarchical dynamic features across multiple time scales. To further enhance robustness in noisy and sparse settings, we incorporate a MiniRocket Ridge Regression (MRR) module which can combine fast MiniRocket transformations with Ridge Regression. Additionally, a Batch-Consistency Confidence-based Semi-Supervised Learning strategy is proposed to maximize the utility of unlabeled data through a confidence-aware batch consistency mechanism. Finally, extensive experiments on both the Benchmark Simulation Model No 2 (BSM2) simulation and real-world datasets demonstrate that our proposed soft sensor significantly improves prediction accuracy and generalization. Compared to the best baseline, the proposed method achieves a 23.38 % reduction in RMSE on the BSM2 dataset and a 24.27 % reduction on real-world data. Moreover, it consistently outperforms other methods in terms of MAE, R, R2 and RMSSD.

