Related Experiment Video
Updated: Jul 3, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Development of a basin-scale total nitrogen prediction model by integrating clustering and regression methods
Su Han Nam1, Siyoon Kwon2, Young Do Kim1
1Department of Civil and Environmental Engineering, Myongji University, Yongin, South Korea.
Abstract:
Nutrient runoff into rivers caused by human activity has led to global eutrophication issues. The Nakdong River in South Korea is currently facing significant challenges related to eutrophication and harmful algal blooms, underscoring the critical importance of managing total nitrogen (T-N) levels. However, traditional methods of indoor analysis, which depend on sampling, are labor-intensive and face limitations in collecting high-frequency data. Despite advancements in sensor allowing for the measurement of various parameters, sensors still cannot directly measure T-N, necessitating surrogate regression methods. Therefore, we conducted T-N predictions using a water quality dataset collected from 2018 to 2022 at 157 observatories within the Nakdong River basin. To account for the water quality characteristics of each location, we employed a clustering technique to divide the basin and compared a Gaussian mixture model with K-means clustering. Moreover, optimal regressor for each cluster was selected by comparing multiple linear regression (MLR), random forest, and XGBoost. The results showed that forming four clusters via K-means clustering was the most suitable approach and MLR was reasonably accurate for all clusters. Subsequently, recursive feature elimination cross-validation was used to identify suitable parameters for T-N prediction, thus leading to the construction of high-accuracy T-N prediction models. Clustering was useful not only for improving the regressors but also for spatially analyzing the water quality characteristics of the Nakdong River. The MLR model can reveal causal relationships and thus is useful for decision-making. The results of this study revealed that the combination of a simple linear regression model and clustering method can be applied to a wide watershed. The clustering-based regression model showed potential for accurately predicting T-N at the basin level and is expected to contribute to nationwide water quality management through future applications in various fields.
Related Concept Videos
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...

