机器学习与回归模型对比,用于预测医院水网中军团菌污染的风险
Osvalda De Giglio1, Fabrizio Fasano1, Giusy Diella1
1Interdisciplinary Department of Medicine, Hygiene Section, University of Bari Aldo Moro, Bari, Italy.
Annali di igiene : medicina preventiva e di comunita
|July 11, 2024
概括
机器学习模型可以高准确地预测医院水系统中的军团菌污染. 这种方法有助于预防军团病爆发,保护患者和医疗保健工作者.
科学领域:
- 环境微生物学环境微生物学
- 医院的水系统是医院的.
- 传染病流行病学 传染病流行病学
背景情况:
- 军团菌在医院的水道网络中构成重大风险,可能会在脆弱的患者和工作人员中引起军团菌病.
- 定期监测对于实施针对军团菌的预防措施至关重要.
- 预测污染风险的标准化方法是提高医院安全的必要条件.
研究的目的:
- 为了标准化医院供水中的军团菌污染的预测方法.
- 为了比较机器学习,传统和组合模型在风险预测中的性能.
- 为了确定影响莱吉欧内拉感染风险的关键因素.
主要方法:
- 在15个月内 (2021年7月至2022年10月) 分析了来自意大利一家医院的1,053个水样.
- 收集了58个与水网结构和环境相关的参数.
- 开发和测试机器学习,回归和组合模型,使用70%的数据并对30%进行验证.
主要成果:
- 5.4%的样本对军团菌呈阳性.
- 最有效的机器学习模型实现了93.4%的准确性,43.8%的灵敏度和96%的特异性.
- 关键影响参数包括水网类型,过器更换和大气温度.
- 机器学习模型在准确性和灵敏性方面表现优于传统和组合模型.
结论:
- 机器学习模型在预测军团菌污染风险方面表现出卓越的表现.
- 需要进一步的研究,通过结合更多的数据和参数来完善预测模型.
- 改进的预测模型可以加强对医院获得的军团病的预防策略.
相关概念视频
Steps in Outbreak Investigation
122
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
122
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Typical Model Studies
354
Fluid mechanics model studies often utilize scaled-down systems to predict fluid behavior in full-scale environments, such as river flows, dam spillways, and structures interacting with open surfaces. Maintaining Froude number similarity in river models is crucial, as it replicates surface flow features like wave patterns and velocities.
354
Mechanistic Models: Compartment Models in Individual and Population Analysis
36
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
36


