Related Experiment Video
Updated: Apr 1, 2026

Author Spotlight: Developing a Point-of-Care Hemoglobin Estimation Method for Anemia Management
Published on: January 19, 2024
Predicting childhood anaemia in Ghana with explainable machine learning: A national survey analysis
Yahye Sheikh Abdulle Hassan1, Mohamed Abdirahim Omar2, Julius Kwabena Karikari3,4
1Faculty of Medicine and Health Sciences, Jamhuriya University of Science and Technology, Mogadishu, Somalia.
Introduction:
Childhood anaemia remains a major public health problem in Ghana, with marked regional and socioeconomic disparities. Conventional regression may not fully capture complex, non-linear relationships among biological, maternal, and household factors. We used supervised machine learning to predict anaemia among children aged 6-59 months using nationally representative survey data.
Methods:
We analysed the 2022 Ghana Demographic and Health Survey, including de facto children aged 6-59 months with valid haemoglobin and complete covariates (weighted N = 3,382). Anaemia was defined as altitude-adjusted haemoglobin <11.0 g/dL. Twenty-one predictors were included. Data were split into training (80%) and testing (20%) sets using stratified sampling. Six models (logistic regression, decision tree, random forest, gradient boosting, support vector machine, and artificial neural network) were tuned via grid search with 10-fold cross-validation.
Results:
The weighted prevalence of childhood anaemia was 48.95% (n = 1,655). Gradient boosting showed the best overall discrimination (AUC = 0.72; F1 = 68.99%; accuracy = 66.27%). Support vector machine and logistic regression achieved the highest sensitivity (recall = 72.73% and 71.74%). Random forest showed overfitting (100% training accuracy; test accuracy = 65.23%). Decision tree and neural network performed poorly (AUC = 0.57 and 0.63). Key predictors across models and SHAP were child age, malaria status, maternal anaemia, region, and household wealth (with feature rankings varying by algorithm).
Conclusion:
Machine learning models achieved moderate predictive performance for childhood anaemia in Ghana. Gradient boosting provided the strongest discrimination, while support vector machine and logistic regression offered higher sensitivity for screening. However, these sensitivities imply that approximately 28-30% of anaemic children may be missed, which should be considered when applying these models in public health screening. Identified determinants support targeted, malaria-integrated nutrition and maternal-child interventions in high-risk groups.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020