在国家COVID队列协作中预测长期COVID使用超级学习者:队列研究
Zachary Butzin-Dozier1, Yunwen Ji1, Haodong Li1
1Division of Biostatistics, University of California Berkeley School of Public Health, Berkeley, CA, United States.
JMIR public health and surveillance
|August 15, 2024
概括
使用机器学习,可以预测COVID-19 (PASC) 后急性后果风险. 医疗保健的使用,人口统计和呼吸系统因素是长期COVID的关键预测因素.
科学领域:
- 计算生物学是一种计算生物学.
- 流行病学 流行病学
- 医疗信息学 医疗信息学
背景情况:
- 后急性COVID-19 (PASC) 或长期COVID的后续症状在生物系统中呈现出各种长期症状.
- 鉴定PASC风险因素和病因是具有挑战性的,因为症状异质.
- 对PASC的预测特征对早期识别和预防有价值.
研究的目的:
- 使用临床数据预测PASC诊断的个体风险.
- 确定PASC发展的关键预测因素.
- 为了利用电子健康记录进行PASC风险评估.
主要方法:
- 利用了Super Learner,一个整体机器学习算法,结合了梯度增强和随机森林.
- 分析了来自国家COVID队列协作的55257名患者 (1:4 PASC对照比率) 的数据.
- 使用Shapley值在个体特征,时间窗口和临床领域中评估变量重要性.
主要成果:
- 实现了准确的PASC诊断预测,曲线下的面积为0.874.4.
- 最重要的预测因素包括观察期长度,急性COVID-19期间的医疗互动,以及下呼吸道感染.
- 基线特征最具预测性,其次是医疗保健使用,人口统计和呼吸系统因素.
结论:
- 开发了一种使用电子健康记录数据进行PASC风险预测的开源方法.
- 医疗保健使用成为一个强有力的预测因素,需要在观察性研究中仔细考虑.
- 在急性COVID-19之前进行早期风险评估,专注于基线和呼吸系因素,可以加强干预措施.
相关概念视频
Longitudinal Research
11.9K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
11.9K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Longitudinal Studies
148
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
148
Steps in Outbreak Investigation
115
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
115
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Cancer Survival Analysis
334
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
334


