随机生存森林用于预测多个生理风险因素对所有原因死亡率的综合影响
Bu Zhao1, Vy Kim Nguyen2,3, Ming Xu4
1School for Environment and Sustainability, University of Michigan, Ann Arbor, MI, USA. zhaobu@umich.edu.
Scientific reports
|July 6, 2024
概括
本研究使用机器学习确定了所有原因死亡的关键生理风险因素. 它强调吸烟,功能,葡萄糖水平,性别和白细胞计数是死亡风险的关键预测因素.
科学领域:
- 公共卫生 公共卫生
- 生物统计学 生物统计学
- 流行病学 流行病学
背景情况:
- 了解所有原因死亡的综合风险因素对于有针对性的干预至关重要.
- 现有的研究往往低估了多种生理风险因素对死亡率的协同效应.
研究的目的:
- 采用生存树机器学习模型来分析综合生理风险因素和全因死亡率.
- 确定影响死亡率的五大因素,并可视化它们的综合影响.
- 将已识别的死亡率切线与当前的临床值进行比较.
主要方法:
- 利用了1999-2014年NHANES调查的数据,与国家死亡指数数据 (17,790名成年人) 相关联.
- 基于生存树的应用机器学习,用于风险因素相互作用的灵活,非参数分析.
- 确定并排列影响因素,可视化它们对死亡率的综合影响.
主要成果:
- 确定了五个最重要的死亡因素,即丁素 (吸烟生物标志物),淋巴膜过率 (GFR),血葡萄糖,性别和白细胞计数.
- 死亡风险增加与男性性别,活跃吸烟,低GFR,高血葡萄糖和高白细胞计数有关.
- 确定的死亡率切线在很大程度上与现有的临床值和研究保持一致.
结论:
- 机器学习模型有效地确定了关键的生理风险因素及其对所有原因死亡率的综合影响.
- 这些发现为增强风险预测提供了基础,为临床实践和精准医学策略提供了信息.
- 对于关键风险因素的确定的切线为风险分层和干预设计提供了宝贵的见解.
相关概念视频
Survival Tree
79
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
79
Comparing the Survival Analysis of Two or More Groups
175
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
175
Mechanistic Models: Compartment Models in Individual and Population Analysis
36
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
36
Assumptions of Survival Analysis
121
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
121
Statistical Methods for Analyzing Epidemiological Data
347
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
347
Odds Ratio
123
The odds ratio (OR) is a statistical measure used extensively in epidemiology and research to quantify the strength of association between exposure and outcome across different groups. Unlike relative risk, which compares the probabilities of an event occurring, the odds ratio compares the odds of an event occurring in the exposed group to the odds of it occurring in the unexposed group. The odds, in this context, are calculated as the probability of the event happening divided by the...
123


