Explainable machine learning for predicting childhood anemia in Sub-Saharan Africa using population-based DHS Data

Andualem Enyew Gedefaw1, Amanuel Worku2, Abraham Keffale Mengistu3

  • 1Department of Health Informatics, Institute of Public Health, College of Medicine and Health Sciences, University of Gondar, Gondar, Ethiopia.

Insights

Childhood anemia prediction in Sub-Saharan Africa shows machine learning models offer moderate performance. While statistically significant improvements over traditional methods were observed, low sensitivity necessitates complementary screening strategies for effective intervention.

Area of Science:

  • Public Health
  • Computational Epidemiology
  • Pediatrics

Background:

  • Childhood anemia is a critical public health issue in Sub-Saharan Africa, impacting child development and survival.
  • Prevalence exceeds 60% across 26 countries, demanding scalable prediction tools for targeted interventions.

Purpose of the Study:

  • To evaluate the performance of various machine learning models in predicting childhood anemia using Demographic and Health Survey (DHS) data.
  • To identify key predictors of childhood anemia for improved public health strategies.

Main Methods:

  • Pooled DHS data from 26 Sub-Saharan African countries (110,251 children aged 6-59 months) were analyzed.
  • Eight machine learning models were trained and evaluated using hyperparameter tuning and 5-fold cross-validation.
  • Model performance was assessed via accuracy, precision, recall, F1-score, ROC-AUC, with SHAP for interpretability.

Main Results:

  • The CatBoost model achieved the highest performance (ROC-AUC = 0.84), outperforming logistic regression significantly (p < 0.001).
  • Models showed moderate discriminatory ability with limited sensitivity (recall ≈ 0.33-0.41).
  • Key predictors included residence type, height-for-age z-score, country, and child age.

Conclusions:

  • Machine learning models offer moderate to strong predictive performance for childhood anemia using DHS data.
  • Despite statistical significance, low sensitivity restricts their use as standalone screening tools.
  • Findings emphasize early nutrition, undernutrition reduction, and integrated predictive analytics for resource allocation in high-burden regions.