一个可解释的机器学习模型与人口变量和饮食模式用于ASCVD识别:来自美国NHANES 1999-2018
Qun Tang1, Yong Wang1, Yan Luo2
1Department of Cardiovascular Medicine, Wuhu City Second People's Hospital, Wuhu, 241000, China.
BMC medical informatics and decision making
|March 3, 2025
概括
机器学习模型确定了动脉样硬化心血管疾病 (ASCVD) 风险的关键因素. 男性的性别,年龄和吸烟增加了风险,而乳制品摄入量减少了风险,为ASCVD预防提供了洞察力.
科学领域:
- 心血管健康 心血管健康
- 机器学习应用 机器学习应用
- 营养流行病学 营养流行病学
背景情况:
- 关于与动脉样硬化心血管疾病 (ASCVD) 的人口和饮食联系的研究有限.
- 了解这些关联对于制定有针对性的预防策略至关重要.
- 现有的模型在预测ASCVD风险时往往缺乏透明度和准确性.
研究的目的:
- 开发和验证一个透明的机器学习 (ML) 算法来预测ASCVD.
- 确定与ASCVD相关的显著的人口和饮食因素.
- 分析使用大型,代表性美国人口数据集的相关性.
主要方法:
- 利用了美国国家健康和营养检查调查 (1999-2018) 的数据,其中有40,298名参与者.
- 开发并比较了五种ML模型来预测ASCVD风险.
- 选择了Extreme梯度提升 (XGBoost) 模型,因为它的性能优越.
主要成果:
- XGBoost模型在预测ASCVD时实现了0.8143的AUC和88.4%的准确性.
- 确定了男性性别,年龄,吸烟和ASCVD风险之间的正相关性.
- 发现乳制品摄入量的负相关性;低精制谷物摄入量没有降低风险;贫困收入比率和卡路里摄入量的非线性关联.
结论:
- XGBoost模型有效地确定了ASCVD的人口和饮食预测因素.
- 男性性别,年龄较大和吸烟是显著的危险因素.
- 饮食因素,如乳制品摄入量发挥作用,突出了个性化营养指导在ASCVD预防中的潜力.
相关概念视频
Genome-wide Association Studies-GWAS
12.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.3K
Mechanistic Models: Compartment Models in Individual and Population Analysis
23
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
23


