通过整合集群和回归方法开发一个盆地规模的总预测模型
Su Han Nam1, Siyoon Kwon2, Young Do Kim1
1Department of Civil and Environmental Engineering, Myongji University, Yongin, South Korea.
The Science of the total environment
|February 10, 2024
概括
在河流中管理总 (T-N) 对于防止缩至关重要. 这项研究使用集群和回归开发了准确的T-N预测模型,改善了Nakdong河流域的水质管理.
科学领域:
- 环境科学 环境科学
- 水质管理水质管理
- 数据科学数据科学数据科学
背景情况:
- 全球肥沃化是由营养物质的排放驱动的,由于高总 (T-N) 的原因,Nakdong河面临着严重的藻类繁殖挑战.
- 传统的TN监测依赖于劳动密集型采样,限制高频数据收集.
- 直接T-N传感器不可用,需要替代回归方法来准确估计.
研究的目的:
- 使用历史水质数据集,为Nakdong河流域开发准确的总 (T-N) 预测模型.
- 为了评估聚类技术的有效性与T-N估计的各种回归模型相结合.
- 为改善流域水质管理提供一个框架.
主要方法:
- 从纳克东河流域 (2018-2022) 的157个观测站收集了水质数据.
- 应用K-means集群,根据水质特征将盆地划分为四个不同的区域.
- 在每个集群中比较多重线性回归 (MLR),随机森林和XGBoost模型用于T-N预测,选择MLR作为最佳. 使用递归特征消除来进行参数选择.
主要成果:
- 将K-means分成四组为空间分析和提高回归精度提供了最合适的方法.
- 在所有集群中,MLR模型表现出合理的准确性,递归特征消除识别了高准确性TN预测的关键参数.
- 开发的基于集群的回归模型有效地预测了盆地规模的T-N水平.
结论:
- 结合K-means集群和MLR,提供了一个强大的,准确的方法来预测T-N在大型河流盆地.
- 这种方法增强了对水质的空间理解,并支持有效管理水资源的知情决策.
- 该研究强调了在全国水质监测和管理战略中广泛应用的潜力.
相关概念视频
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Sampling Plans
181
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
181
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
54
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
54


