関連する実験動画
Updated: Feb 22, 2026

12:44
Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
8.7K
波を起こす: 自動相関とベースラインモデルの見過ごされた役割を通して,廃水流水質の予測における機械学習を再考する
Yijie Wang1, Damien Batstone2, Zhenju Sun3
1School of Civil and Environmental Engineering, Nanyang Technological University, 50 Nanyang Avenue 639798, Singapore; Nanyang Environment & Water Research Institute, Nanyang Technological University, 1 Cleantech Loop 637141, Singapore.
Water research
|February 20, 2026
まとめ
排水処理所の排水質の予測のための機械学習 (ML) モデルは,標的の自己相関性により,精度を過大評価することがあります. 持続性モデルは,特に低揮発性条件では,しばしばMLを上回り,より良いベースラインの必要性を強調します.
科学分野:
- 環境工学環境工学とは
- 機械学習 アプリケーション
- 排水処理 排水処理 排水処理
背景:
- タイムシリーズの機械学習 (ML) は,排水処理施設 (WWTP) の排水質を予測するために広く使用されており,予測精度に重点を置いています.
- 排水質データには固有の自己相関性があり,MLモデルの認識されたパフォーマンスを膨らませることができます.
- R-squaredのような標準的なメトリックは,単純な持続性モデルが高いスコアを達成し,適切なベースラインで再評価を必要とするため,誤解を招く可能性があります.
研究 の 目的:
- ターゲットの自己相関を考慮して,WWTPの排水質の予測のためのMLモデルのパフォーマンスを再評価する.
- MLモデルの解釈性と正確性を評価するための重要な基準として,持続性モデルを導入し,検証する.
- データボラティリティとモデルのパフォーマンスへの影響を定量化するための新しいインデックスを提案する.
主な方法:
- 3つの公表された研究と2つの追加のWWTPデータセットのデータを用いて,自動回帰的な方法の評価.
- 予測に対する過去の目標値の影響を定量化するために,集約されたSHAP分析の適用.
- MLモデルのパフォーマンスを,異なる予測期間とボラティリティシナリオにおける持続性モデルのベンチマークと比較する.
主要な成果:
- SHAPの分析では,過去の目標値が他のパラメータを大幅に支配していることが明らかになった (64%-396%の重要度が高い),これは,報告された高い精度が自動相関から生じる可能性があることを示している.
- 持続性モデルは,化学酸素需要 (COD),総窒素 (TN),総リン (TP) の予測において,特に低揮発性シナリオにおいて,MLモデルを頻繁に上回りました.
- 新しい指数PN-MAROCは,データの変動を効果的に定量化し,持続性 (R2 = 0.93) とMLモデル (R2 = 0.75) のパフォーマンスの両方と強く相関しています.
結論:
- この研究は,WWTPの排水質の予測におけるMLモデルのパフォーマンスを解釈する際に,ターゲットの自己相関を考慮し,堅固なベースラインを採用する必要性を強調しています.
- 持続性モデルは,重要な基準として機能し,しばしば複雑なMLモデルを上回ります,特に排水質が低揮発性を示す場合です.
- 提案されたPN-MAROCインデックスは,データ特性を評価し,排水管理における予測モデルの選択と解釈を導くための実用的なツールを提供します.
キーワード:
自動相関は,自己相関である.ベンチマーキング (Benchmarking) とは機械学習 (Machine Learning) とは,機械学習 (Machine Learning) というものです.持続性モデル (persistence model) とは,持続性モデル (persistence model) とは,持続性モデル (persistence model) とは,持続性モデル (persistence model) とは,持続性モデル (persistence model) とは,持続性モデル (persistence model) とは,持続性のモデル (persistence model) とは,持続性のモデル (persistence model) とは,持続性のモデル (persistence model) とは,持続性のモデル (persistence model) とは,持続性のモデル (persistence model) とは,持続性のモデル (persistence model) とは,タイムシリーズのモデルモデル排水処理施設は,排水処理施設として利用できます.関連する概念動画
Mechanistic Models: Compartment Models in Individual and Population Analysis
290
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
290
Multiple Regression
4.1K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.1K
Steps in Outbreak Investigation
622
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
622
Residuals and Least-Squares Property
9.6K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.6K