机器学习算法的应用用于预测艾滋病毒检测,使用来自2002-2017年南非成人人口调查的证据:一种艾滋病毒检测预测模型
Musa Jaiteh1, Edith Phalane1, Yegnanew A Shiferaw2
1South African Medical Research Council/University of Johannesburg Pan African Centre for Epidemics Research Extramural Unit, Faculty of Health Sciences, University of Johannesburg, Johannesburg 2006, South Africa.
Tropical medicine and infectious disease
|June 25, 2025
概括
这项研究使用机器学习来确定南非艾滋病毒检测的关键预测因素. 调查结果强调,对测试地点,人口统计和社会经济地位的了解对于有针对性的干预至关重要.
科学领域:
- 公共卫生 公共卫生
- 流行病学 流行病学
- 机器学习 机器学习
背景情况:
- 南非人口中很大一部分人未知艾滋病毒状况,这阻碍了疫情控制工作.
- 现有的预测模型不足以指导南非的针对性艾滋病毒检测干预措施.
- 机器学习 (ML) 提供了识别艾滋病毒感染风险较高的个体的潜力,为测试建议提供了信息.
研究的目的:
- 通过使用监督机器学习 (SML) 算法,确定南非成年人中艾滋病毒检测的一致预测因素.
- 评估四个SML算法 (决策树,随机森林,SVM,物流回归) 在基于人口的调查数据上的性能.
- 为提高艾滋病毒检测效率和实现联合国艾滋病规划署2030年目标提供数据驱动的政策倡议信息.
主要方法:
- 利用南非国家艾滋病毒流行,发病率,行为和沟通调查 (SABSSM) 数据集,从五个横截面周期.
- 应用了四种SML算法:决策树,随机森林,支持向量机 (SVM) 和后勤回归.
- 雇佣了80%的培训和20%的测试,每个数据集的5倍交叉验证.
主要成果:
- 随机森林在所有数据集中表现出卓越的性能,实现了最高的准确性,精度,F1得分和AUC.
- 艾滋病毒检测的关键预测因素包括对检测地点的了解,女性,年龄较小,高社会经济地位和数字信息获取.
- SVM显示了高回忆率,但精度较低,而物流回归和决策树的性能中等,决策树容易过度拟合.
结论:
- 监督机器学习,特别是随机森林,有效地识别了南非艾滋病毒检测的预测因素.
- 有针对性的干预措施应侧重于提高对测试场所的认识,利用数字平台,并解决社会经济因素.
- 通过数据驱动的策略来提高艾滋病毒检测的有效性,对于南非朝着联合国艾滋病规划署2030年目标的进步至关重要.
相关概念视频
Steps in Outbreak Investigation
215
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
215
Statistical Methods for Analyzing Epidemiological Data
550
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
550


