基于流动性的COVID-19病例预测模型的公平性评估
Abdolmajid Erfani1, Vanessa Frias-Martinez2,3
1Department of Civil, Environmental, and Geospatial Engineering, Michigan Technological University, Houghton, MI, United States of America.
PloS one
|October 18, 2023
概括
使用移动数据的COVID-19预测模型显示偏差,对老年,贫穷和农村人口的表现不太准确. 这凸显了需要更具代表性的数据收集,以确保所有人口群体之间公平的模型性能.
科学领域:
- 流行病学 流行病学
- 数据科学数据科学数据科学
- 公共卫生 公共卫生
背景情况:
- 人类流动性分析对于了解COVID-19传播和干预效果至关重要.
- 公共可用的移动数据有助于研究时空趋势和疾病传播.
- 基于流动性的预测模型在人口统计学中的公平表现仍然是一个开放的问题.
研究的目的:
- 为了调查基于移动性的COVID-19预测模型是否在不同的人口群体中表现差异.
- 测试这种假设,即移动数据中的偏差有助于不公平的模型准确性.
- 确定与模型性能变化相关的特定社会人口统计特征.
主要方法:
- 在美国县级应用了两种基于移动性的COVID-19感染预测模型.
- 利用SafeGraph移动数据进行模型培训和评估.
- 与县级社会人口统计数据相关联的模型性能指标.
主要成果:
- 在人口统计学特征中观察到模型性能存在显著的系统偏差.
- 模型对大,受过教育,富裕,年轻和城市县的准确性更高.
- 对于老年,贫穷,受教育程度较低和农村人口的预测,准确性较低.
结论:
- 当前预测模型中使用的移动数据可能不足以代表某些人群,导致偏见的COVID-19预测.
- 需要改进数据收集和抽样策略,以确保全面的流动模式代表性.
- 解决数据偏差对于开发公平,准确的公共卫生监测工具至关重要.
相关概念视频
Steps in Outbreak Investigation
135
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
135
Bias in Epidemiological Studies
314
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
314
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Mechanistic Models: Compartment Models in Individual and Population Analysis
45
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
45
Relative Risk
191
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
191


