模拟稀疏的裂谷热发病率数据:对零膨胀的自我激发和自动回归模型的贝叶斯观点
Alexandros Angelakis1,2, Bryan O Nyawanda1,2, Penelope Vounatsou3,4
1Swiss Tropical and Public Health Institute, Kreuzstrasse 2, Allschwil, CH-4123, Basel-Land, Switzerland.
BMC infectious diseases
|September 30, 2025
概括
零膨胀模型显著提高了裂谷热 (RVF) 预测准确度,特别是在稀疏的监测数据下. 使用降雨数据的零膨胀负二项式模型为RVF流行地区提供了最可靠的预测.
科学领域:
- 流行病学 流行病学
- 数学建模的数学建模
- 兽医公共卫生 兽医公共卫生
背景情况:
- 裂谷热 (RVF) 是一种动物传播疾病,监测数据稀少,使预测建模复杂化.
- 人类和牲畜监测中零计数的高频率是常见的挑战.
- 零膨胀模型和自我激发的时间模型是稀疏计数数据的潜在框架.
研究的目的:
- 为了比较三种贝叶斯式零膨胀时间计数模型对RVF的预测性能.
- 为了评估模拟数据集的模型性能,具有不同的稀疏度级别.
- 通过使用现实数据,确定最准确的RVF发病率预测模型.
主要方法:
- 三个零膨胀贝叶斯模型的比较:零膨胀的负二项式 (ZINB) 与自行回归的时间随机效应,自我激发的负二项式 (SE-NB) 和一般化的自行回归移动平均负二项式 (GARMA-NB).
- 使用模拟数据集与受控稀疏度进行评估.
- 适用于肯尼亚北部 (2018-2024) 每月的RVF发病率数据,包括三个月的降雨滞后.
主要成果:
- 零通货膨胀在特定稀疏度值内的所有测试模型的预测性能得到了显著的改善.
- 在ZINB模型:29-94.5%的改进;在SE-NB模型:25-93%的改进;在GARMA-NB模型:30-95%的改进.
- 带有三个月降雨滞后的ZINB模型在预测肯尼亚北部RVF发病率方面表现出最高的准确性.
结论:
- 零膨胀负二项式模型对于准确的RVF预测至关重要,特别是在数据稀疏的环境中.
- 基于气候的共同变量,如降雨量,显著提高了这些模型的早期预警能力.
- 这些发现支持将先进的统计建模集成为改善在特有地区的RVF监测和控制.
相关概念视频
Steps in Outbreak Investigation
492
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
492
Statistical Methods for Analyzing Epidemiological Data
898
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
898
Parametric Survival Analysis: Weibull and Exponential Methods
1.0K
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
1.0K
Clearance Models: Noncompartmental Models
244
Clearance is a pharmacokinetic parameter traditionally defined by compartment models, signifying the rate at which a drug is expelled from the body. However, a noncompartmental model offers an alternative method for assessing clearance, primarily employing empirical data obtained after administering a single drug dose.
The noncompartmental approach capitalizes on extensive sampling data, correlating the volume of distribution to systemic exposure and the administered dosage. This method enables...
The noncompartmental approach capitalizes on extensive sampling data, correlating the volume of distribution to systemic exposure and the administered dosage. This method enables...
244
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
242
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
242
Distributions to Estimate Population Parameter
5.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
5.0K


