Related Experiment Video
Updated: Sep 18, 2025

Comprehensive & Cost Effective Laboratory Monitoring of HIV/AIDS: an African Role Model
Published on: October 31, 2010
The Application of Machine Learning Algorithms to Predict HIV Testing Using Evidence from the 2002-2017 South African
Musa Jaiteh1, Edith Phalane1, Yegnanew A Shiferaw2
1South African Medical Research Council/University of Johannesburg Pan African Centre for Epidemics Research Extramural Unit, Faculty of Health Sciences, University of Johannesburg, Johannesburg 2006, South Africa.
This study used machine learning to identify key predictors of HIV testing in South Africa. Findings highlight knowledge of testing sites, demographics, and socioeconomic status as crucial for targeted interventions.
Area of Science:
- Public Health
- Epidemiology
- Machine Learning
Background:
- A significant portion of South Africa's population has an unknown HIV status, hindering epidemic control efforts.
- Existing predictive models are insufficient for guiding targeted HIV testing interventions in South Africa.
- Machine learning (ML) offers potential for identifying individuals at higher risk for HIV infection, informing testing recommendations.
Purpose of the Study:
- To identify consistent predictors of HIV testing among adults in South Africa using supervised machine learning (SML) algorithms.
- To evaluate the performance of four SML algorithms (decision trees, random forest, SVM, logistic regression) on population-based survey data.
- To inform data-driven policy initiatives for enhancing HIV testing efficacy and achieving the UNAIDS 2030 goal.
Main Methods:
- Utilized the South African National HIV Prevalence, Incidence, and Behavior and Communication Survey (SABSSM) datasets from five cross-sectional cycles.
- Applied four SML algorithms: decision trees, random forest, support vector machines (SVM), and logistic regression.
- Employed an 80% training and 20% testing split with 5-fold cross-validation for each dataset.
Main Results:
- Random forest demonstrated superior performance across all datasets, achieving the highest accuracy, precision, F1-score, and AUC.
- Key predictors of HIV testing included knowledge of testing sites, being female, younger age, high socioeconomic status, and digital information access.
- SVM showed high recall but lower precision, while logistic regression and decision trees had moderate performance, with decision trees prone to overfitting.
Conclusions:
- Supervised machine learning, particularly random forest, effectively identifies predictors of HIV testing in South Africa.
- Targeted interventions should focus on increasing awareness of testing sites, leveraging digital platforms, and addressing socioeconomic factors.
- Improving HIV testing efficacy through data-driven strategies is crucial for South Africa's progress towards the UNAIDS 2030 goal.
Related Concept Videos
Steps in Outbreak Investigation
Statistical Methods for Analyzing Epidemiological Data

