Related Experiment Video
Updated: Aug 26, 2025

Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
Predicting in-stream water quality constituents at the watershed scale using machine learning.
Itunu C Adedeji1, Ebrahim Ahmadisharaf1, Yanshuo Sun2
1Department of Civil and Environmental Engineering, Resilient Infrastructure and Disaster Response Center, Florida A&M University-Florida State University College of Engineering, 2525 Pottsdamer St., Tallahassee, FL 32310, USA.
Machine learning models can predict key water quality indicators like nutrients and suspended solids using publicly available data. However, fecal coliform bacteria prediction remains challenging, requiring further data integration for accurate forecasting.
Area of Science:
- Environmental Science
- Water Quality Monitoring
- Machine Learning Applications
Background:
- Accurate in-stream water quality prediction is crucial for effective waterbody management and restoration.
- Data-driven modeling, particularly machine learning, offers a powerful approach for water quality assessment.
- Publicly available datasets present an opportunity for scalable water quality prediction frameworks.
Purpose of the Study:
- To evaluate the efficacy of machine learning algorithms in predicting key in-stream water quality constituents using various data inputs.
- To assess the performance of models using only publicly available data versus those incorporating antecedent conditions and additional water quality observations.
- To determine the feasibility of predicting total phosphorus (TP), total nitrogen (TN), total suspended solids (TSS), dissolved oxygen (DO), and fecal coliform bacteria (FC).
Main Methods:
- Five machine learning algorithms were implemented: Support Vector Machines, Random Forest (RF), eXtreme Gradient Boost (XGB), ensemble RF-XGB, and Artificial Neural Network (ANN).
- Four modeling scenarios were tested, varying input data from publicly available datasets to include antecedent conditions and in-stream water quality observations (e.g., pH, turbidity).
- Model performance was evaluated using Nash-Sutcliffe efficiency (NSE) and percent bias (PBIAS) metrics in Bullfrog Creek, Florida.
Main Results:
- Models using only publicly available data satisfactorily predicted TP, TN, DO, and TSS (NSE > 0.75, PBIAS < 10%).
- Fecal coliform bacteria (FC) predictions were unsatisfactory using only public data (NSE < 0.49, PBIAS > 25%), indicating a need for more specific data.
- Antecedent conditions slightly improved predictions and reduced uncertainty, especially when combined with other water quality observations.
- Model performance showed minimal seasonal dependence for all constituents.
Conclusions:
- Publicly available datasets are sufficient for satisfactory prediction of key water quality parameters like nutrients and TSS.
- Predicting fecal coliform bacteria requires more comprehensive data beyond standard publicly available sources.
- The developed machine learning framework is adaptable and can serve as a valuable tool for routine water quality monitoring and prediction.
More Related Videos
Related Concept Videos
Steps in Outbreak Investigation
Testing Water Quality
Typical Model Studies
Quality of Water
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Design Example: Analyzing Capacity Contours for Flood Risk Assessment

