一种可解释的时间序列机器学习方法,用于基于废水的流行病学中的不同预测和现在预测长度
Mallory Lai1, Shaun S Wulff1, Yongtao Cao2
1Department of Mathematics and Statistics, University of Wyoming, 1000 E University Ave, Laramie, WY, USA.
MethodsX
|October 12, 2023
概括
本研究引入了一种时间序列机器学习方法,使用废水数据预测COVID-19病例. 这种方法通过分析环境变量来进行准确的预测来加强疾病监测.
科学领域:
- 环境科学环境科学
- 流行病学 流行病学
- 机器学习 机器学习
背景情况:
- 基于废水的流行病学是公共卫生监测的日益增长的方法.
- 监测像COVID-19这样的传染病对于人口健康管理至关重要.
研究的目的:
- 介绍一种新的时间序列机器学习 (TSML) 方法来预测COVID-19病例.
- 为基于废水的疾病监测开发一个可解释和以假设为导向的框架.
主要方法:
- 利用功能工程来创建可解释的预测器,例如特定站点的交付时间.
- 采用特征选择来识别最佳预测变量用于现在预测和预测.
- 实施了先决评估,以确保可靠的绩效评估,并防止数据泄露.
主要成果:
- 通过废水数据,TSML方法在预测COVID-19病例方面表现出有效性.
- 该框架成功地将环境变量集成到预测模型中.
- 该方法允许灵活的现在预测和不同长度的预测.
结论:
- 开发的TSML方法为基于废水的流行病学提供了可行的工具.
- 这种方法提高了疾病监测系统的准确性和可解释性.
- 这些发现支持在公共卫生中使用机器学习来预测传染病趋势.
相关概念视频
Steps in Outbreak Investigation
135
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
135
Noncompartmental Analysis: Mean Residence Time
161
According to statistical moment theory, mean residence time (MRT) is an important measure in pharmacokinetics. MRT can be defined as the expected mean of a probability density function distribution. It provides valuable insights into drug disposition in the body.
After the administration of a drug through intravenous bolus injection, the drug molecules are distributed throughout the body and remain there for varying periods. The MRT represents the average time these drug molecules stay in the...
After the administration of a drug through intravenous bolus injection, the drug molecules are distributed throughout the body and remain there for varying periods. The MRT represents the average time these drug molecules stay in the...
161
Statistical Methods for Analyzing Epidemiological Data
385
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
385
Mechanistic Models: Compartment Models in Individual and Population Analysis
59
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
59
Introduction To Survival Analysis
252
Survival analysis is a statistical method used to study time-to-event data, where the "event" might represent outcomes like death, disease relapse, system failure, or recovery. A unique feature of survival data is censoring, which occurs when the event of interest has not been observed for some individuals during the study period. This requires specialized techniques to handle incomplete data effectively.
The primary goal of survival analysis is to estimate survival time—the time...
The primary goal of survival analysis is to estimate survival time—the time...
252
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K


