Artificial Intelligence and Machine Learning-based prediction of tuberculosis treatment failure: A systematic review
Rogers Kamulegeya1, Rose Nabatanzi1, Derrick Semugenze1,2
1Department of Immunology and Molecular Biology, Makerere University College of Health Sciences (MakCHS), Kampala, Uganda.
Background:
Tuberculosis (TB) remains a leading cause of infectious disease mortality worldwide, and treatment failure contributes to ongoing transmission, drug resistance, and poor clinical outcomes. Artificial intelligence (AI) and machine learning (ML) approaches have attracted growing interest in predicting TB treatment outcomes, but the literature is heterogeneous and lacks a comprehensive synthesis.
Methods:
We systematically searched PubMed/MEDLINE and Embase (January 2000-October 2025) for studies developing or validating AI/ML models to predict TB treatment failure. Two reviewers independently screened records and extracted data on study characteristics, predictor modalities, algorithms, validation strategies, and performance metrics, i.e., area under the curve (AUC), sensitivity, specificity, and confidence intervals. Studies reporting AUC with confidence intervals or sufficient data for calculation were included in random-effects meta-analysis. Missing standard errors were estimated from sample sizes and event rates using established methods. Risk of bias was assessed using PROBAST. Subgroup analyses and meta-regression explored heterogeneity, and publication bias was assessed using funnel plots, Egger's test, and trim-and-fill analysis. The study is registered with PROSPERO (CRD420251101443).
Results:
Thirty-four studies met the inclusion criteria. Publications increased markedly from 2019 onwards (91% of studies). Tree-based methods predominated (52.9%), and multimodal models (≥3 data types) were used in 41.2% of the studies. Nineteen studies (100,790 participants) contributed to the meta-analysis. The pooled AUC was 0.836 (95% CI 0.799-0.868), with substantial heterogeneity (I² = 97.9%). In subgroup analyses, studies including HIV-positive participants showed lower discrimination (AUC 0.748) than those excluding them (0.924). Only eight studies (23.5%) performed external validation, and only one study (2.9%) was rated low risk of bias overall (PROBAST), primarily due to analytical domain deficiencies. Egger's test suggested publication bias (p = 0.024). Major evidence gaps included underrepresentation of high-burden countries, HIV-affected populations, social determinants, pediatric TB, and extrapulmonary disease.
Conclusions:
AI/ML models for predicting TB treatment failure show promising discrimination but are not yet ready for routine clinical implementation. Performance varies substantially across populations and settings, and methodological limitations, including inadequate validation, poor calibration assessment, and high risk of bias, limit confidence in current estimates. Future research should prioritize rigorous external validation, calibration assessment, and development in underrepresented populations, particularly HIV-affected and high TB burden settings.
