探索可解释的机器学习算法,以模拟2018年至2023年间撒哈拉以南非洲男性烟草使用的预测因素
Mequannent Sharew Melaku1, Nebebe Demis Baykemagn2, Lamrot Yohannes3
1Department of Health Informatics, Institute of Public Health, University of Gondar, Gondar, Ethiopia. Mequannent.sharew@uog.edu.et.
机器学习确定了撒哈拉以南非洲男性烟草使用的关键预测因素. 年龄,教育和财富是重要因素,为戒烟提供了有针对性的公共卫生干预信息.
科学领域:
- 公共卫生 公共卫生
- 流行病学 流行病学
- 数据科学数据科学数据科学
背景情况:
- 在撒哈拉以南非洲,烟草吸烟对公共卫生构成了重大挑战.
- 烟草使用的普遍性受到各种人口和社会经济因素的影响.
研究的目的:
- 在撒哈拉以南非洲 (2018-2023) 的男性中模拟烟草使用的预测因素.
- 利用机器学习算法来识别烟草消费的关键决定因素.
主要方法:
- 人口和健康调查数据的分析 (n=147,466名男性).
- 机器学习模型的应用:决策树,物流回归,随机森林,KNN,XGBoost,AdaBoost.
- 通过随机搜索进行超参数优化,并进行十倍交叉验证;使用SHAP分析评估预测器意义.
主要成果:
- 综合使用烟草的流行率为14.73%,莫桑比克,赞比亚,贝宁,马里,毛里塔尼亚,塞内加尔,几内亚,塞拉利昂和利比里亚的流行率显著.
- XGBoost模型实现了98%的准确性和97%的AUC.
- 确定了关键预测因素:年龄,教育,财富指数,宗教,居住地,互联网使用,职业,第一次性行为的年龄,性伴侣数量和婚姻状况.
结论:
- 机器学习有效地识别了撒哈拉以南非洲男性烟草使用的重要预测因素.
- 人口,社会经济和行为因素对于了解烟草使用模式至关重要.
- 这些发现支持制定有针对性的公共卫生战略,以减轻吸烟的流行.
更多相关视频
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
相关概念视频
Statistical Methods for Analyzing Epidemiological Data
Steps in Outbreak Investigation
Survival Tree
Building a Survival Tree
Constructing a...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
