对比量子回归spline分析和监督机器学习,用于沿海海洋水产养殖设施的环境质量评估
Kleopatra Leontidou1, Verena Rubel1, Thorsten Stoeck1
1Ecology Group, Rheinland-Pfälzische Technische Universität Kaiserslautern-Landau, Kaiserslautern, Germany.
PeerJ
|June 19, 2023
概括
监督机器学习 (SML) 和量子回归线 (QRS) 通过使用细菌eDNA元编码有效地推断海洋环境质量. 在监测水产养殖影响方面,SML表现出更高的准确性和稳定性.
科学领域:
- 海洋生态海洋生态学
- 环境DNA (eDNA) 转基因编码
- 水产养殖影响评估
背景情况:
- 海洋鱼水产养殖导致有机丰富,这是沿海生态系统的局部压力因素.
- 传统的地大无脊椎动物生物监测是耗时且昂贵的.
- 细菌eDNA元编码为环境质量评估提供了一种快速,具有成本效益的替代方案.
研究的目的:
- 为了比较量子回归线 (QRS) 和监督机器学习 (SML) 的性能,从细菌元编码数据推断环境质量.
- 评估这些方法是否适用于监测水产养殖对有机丰富的影响.
- 将QRS和SML的准确性和稳定性与参考动物质量指数 (IQI) 进行评估.
主要方法:
- 从挪威和苏格兰的鱼养殖场收集了230个水产养殖样本,沿着有机丰富梯度.
- 使用eDNA元编码分析了细菌群落.
- 从盆地宏观动物数据计算了参考动物质量指数 (IQI).
- 应用QRS来识别细菌指标并计算分子IQI.
- 开发了一个SML (随机森林) 模型,直接预测基于宏观动物的IQI.
主要成果:
- 无论是QRS还是SML都准确地推断出环境质量,分别准确率为89%和90%.
- 对于两种方法,参考IQI和分子IQI之间发现了很高的对应性 (p < 0.001).
- 与QRS相比,SML在处理自然变异方面表现出更高的确定系数和更大的稳定性.
- 在SML确定的20个最重要的ASV中,15个与两个地区的QRS指标一致.
结论:
- 监督机器学习 (SML) 是一种强大的方法,可以使用eDNA元编码数据推断海洋环境质量.
- SML显示了监测水产养殖对海洋生态系统的影响的巨大潜力.
- 建议进行进一步的研究和纳入更多样本,以改进SML模型并减少来自时空变化的噪声.
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Quantitative Analysis
336
Quantitative analysis is a technique for measuring the amount of specific constituents in a sample. When the sample's composition is unknown, qualitative analysis is performed first to identify its components, which ensures that the correct substances are measured during the quantitative phase.
In quantitative analysis, two key measurements are made: the sample quantity and a property proportional to the amount of the analyte (the substance being analyzed). This forms the basis of the...
In quantitative analysis, two key measurements are made: the sample quantity and a property proportional to the amount of the analyte (the substance being analyzed). This forms the basis of the...
336
Response Surface Methodology
190
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
The process of RSM involves several key steps:
190
Testing Water Quality
144
When the quality of water for concrete preparation is uncertain, its impact on the setting time of cement and compressive strength of mortar is assessed by comparison with de-ionized or distilled water benchmarks. American Society for Testing and Materials (ASTM) C1602 requires the setting times to be within 90 minutes of the control, British Standard (BS) 3146:1980 allows a 30-minute variance in the initial setting, while British Standards European Norm (BS EN) 1008 specifies initial setting...
144
Mechanistic Models: Compartment Models in Individual and Population Analysis
66
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
66


