Related Experiment Video
Updated: Aug 26, 2025

Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
Predicting in-stream water quality constituents at the watershed scale using machine learning
Itunu C Adedeji1, Ebrahim Ahmadisharaf1, Yanshuo Sun2
1Department of Civil and Environmental Engineering, Resilient Infrastructure and Disaster Response Center, Florida A&M University-Florida State University College of Engineering, 2525 Pottsdamer St., Tallahassee, FL 32310, USA.
Abstract:
Predicting in-stream water quality is necessary to support the decision-making process of protecting healthy waterbodies and restoring impaired ones. Data-driven modeling is an efficient technique that can be used to support such efforts. Our objective was to determine if in-stream concentrations of contaminants, nutrients-total phosphorus (TP) and total nitrogen (TN) -total suspended solids (TSS), dissolved oxygen (DO), and fecal coliform bacteria (FC) can be predicted satisfactorily using machine learning (ML) algorithms based on publicly available datasets. To achieve this objective, we evaluated four modeling scenarios, differing in terms of the required inputs (i.e., publicly available datasets (e.g., land-use/land cover)), antecedent conditions, and additional in-stream water quality observations (e.g., pH and turbidity). We implemented five ML algorithms-Support Vector Machines, Random Forest (RF), eXtreme Gradient Boost (XGB), ensemble RF-XGB, and Artificial Neural Network (ANN) -and demonstrated our modeling framework in an inland stream-Bullfrog Creek, located near Tampa, Florida. The results showed that, while including additional water quality drivers improved overall model performance for all target constituents, TP, TN, DO, and TSS could still be predicted satisfactorily using only publicly available datasets (Nash-Sutcliffe efficiency [NSE] > 0.75 and percent bias [PBIAS] < 10%), whereas FC could not (NSE < 0.49 and PBIAS >25%). Additionally, antecedent conditions slightly improved predictions and reduced the predictive uncertainty, particularly when paired with other water quality observations (6.9% increase in NSE for FC, and 2.7% for TP, TN, DO, and TSS). Also, comparable model performances of all water quality constituents in wet and dry seasons suggest minimal season-dependence of the predictions (<4% difference in NSE and < 10% difference in PBIAS). Our developed modeling framework is generic and can serve as a complementary tool for monitoring and predicting in-stream water quality constituents.
More Related Videos
Related Concept Videos
Steps in Outbreak Investigation
Testing Water Quality
Typical Model Studies
Quality of Water
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Design Example: Analyzing Capacity Contours for Flood Risk Assessment

