机器学习算法用调查数据预测美国的重度插曲性饮酒情况.
Laura Llamosas-Falcón1,2, Charlotte Probst1,2,3,4, Kevin Shield1,3,5
1Institute for Mental Health Policy Research, Centre for Addiction and Mental Health, Toronto, Canada.
Drug and alcohol review
|November 5, 2025
概括
机器学习算法使用调查数据准确地预测重型情节性饮酒 (HED). XGBoost表现出卓越的表现,确定每日平均饮酒量和年龄作为公共卫生政策的关键预测因素.
科学领域:
- 公共卫生 公共卫生
- 数据科学数据科学数据科学
- 医疗信息学 医疗信息学
背景情况:
- 严重的偶发性饮酒 (HED) 是一个重大的公共卫生挑战,由于调查的局限性,经常被低估.
- 在直接测量不可靠的情况下,预测建模为估计个人HED风险提供了一个解决方案.
- 机器学习算法 (MLA) 在预测HED时可能会超过传统的后勤回归.
研究的目的:
- 为了比较各种MLA的预测性表现,以识别重度插曲性饮酒.
- 为了确定HED预测中最准确和最强大的MLA.
- 通过使用SHAP值来评估HED预测的特征重要性.
主要方法:
- 利用了来自国家健康访谈调查 (1997-2018) 的数据.
- 培训和交叉验证了六个MLA:后勤回归,天真的海湾,k-最近的邻居,支持矢量机,随机森林和XGBoost.
- 采用了夏普利添加式解释 (SHAP) 方法来解释可解释性和特征排序.
主要成果:
- XGBoost表现出最高的性能 (精度为0.92,灵敏度为0.80,精度为0.83).
- 通过正确排序HED实例的概率来衡量模型性能,范围从0.85到0.97.
- 通过SHAP分析,平均每日饮酒量和年龄被确定为HED最重要的预测因素.
结论:
- 选择的特征可以产生强大的HED预测,证明MLA对健康行为建模的潜力.
- 将这些模型集成到模拟框架中,可以为HED制定有效的公共卫生政策.
- 未来的研究应该专注于外部验证和偏见调查,以提高预测准确度.
更多相关视频
08:05A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
10.1K
05:12Chronic Intermittent Ethanol Vapor Exposure Paired with Two-Bottle Choice to Model Alcohol Use Disorder
Published on: June 23, 2023
1.4K
相关概念视频
Hypothesis Test for Test of Independence
7.4K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
7.4K
Steps in Outbreak Investigation
481
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
481
Determination of Expected Frequency
2.5K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.5K
