Related Experiment Video
Updated: Jun 2, 2025

Comparative Analysis of Automatic Fecal Analyzer versus Direct Wet Smear Microscopy for Detecting Parasitic Infections in Stool Samples
Published on: April 25, 2025
Derivation and validation of a clinical predictive model for longer duration diarrhea among pediatric patients in
Billy Ogwel1,2, Vincent H Mzazi3, Alex O Awuor4
1Kenya Medical Research Institute- Center for Global Health Research (KEMRI-CGHR), P.O Box 1578-40100, Kisumu, Kenya. ogwelbill@gmail.com.
Insights
Machine learning models can predict longer duration diarrhea (LDD) in children. This tool helps identify at-risk children for better management and improved health outcomes.
Area of Science:
- Pediatric infectious diseases
- Computational epidemiology
- Clinical decision support systems
Background:
- Longer duration diarrhea (LDD) in children is associated with adverse health outcomes.
- Current clinical tools for identifying children at risk of LDD are lacking.
- Machine learning (ML) offers a novel approach for developing predictive models for LDD.
Purpose of the Study:
- To derive and validate a machine learning (ML) predictive model for identifying children at increased risk of longer duration diarrhea (LDD).
- To assess the performance and calibration of various ML algorithms in predicting LDD.
- To identify key predictors of LDD in young children.
Main Methods:
- Utilized de-identified data from two African studies (N=1,482 and N=682) for model development and temporal validation.
- Applied seven ML algorithms, including random forest, to predict LDD (≥7 days).
- Employed split-sampling, K-fold cross-validation, and over-sampling; used explainable AI for predictor importance.
Main Results:
- The random forest model demonstrated the best performance with an AUC of 83.0% in development and 71.0% in validation.
- Key predictors of LDD included pre-enrolment diarrhea duration, modified Vesikari score, age, and vomiting.
- The model's calibration was good and not statistically significant (Brier score=0.17, p=0.219).
Conclusions:
- ML-derived algorithms can effectively identify children at higher risk of LDD.
- Integrating these ML models into clinical practice can facilitate targeted management and closer observation for at-risk children.
- This approach has the potential to improve clinical decision-making and patient outcomes for pediatric diarrhea.
Background:
Despite the adverse health outcomes associated with longer duration diarrhea (LDD), there are currently no clinical decision tools for timely identification and better management of children with increased risk. This study utilizes machine learning (ML) to derive and validate a predictive model for LDD among children presenting with diarrhea to health facilities.
Methods:
LDD was defined as a diarrhea episode lasting ≥ 7 days. We used 7 ML algorithms to build prognostic models for the prediction of LDD among children < 5 years using de-identified data from Vaccine Impact on Diarrhea in Africa study (N = 1,482) in model development and data from Enterics for Global Health Shigella study (N = 682) in temporal validation of the champion model. Features included demographic, medical history and clinical examination data collected at enrolment in both studies. We conducted split-sampling and employed K-fold cross-validation with over-sampling technique in the model development. Moreover, critical predictors of LDD and their impact on prediction were obtained using an explainable model agnostic approach. The champion model was determined based on the area under the curve (AUC) metric. Model calibrations were assessed using Brier, Spiegelhalter's z-test and its accompanying p-value.
Results:
There was a significant difference in prevalence of LDD between the development and temporal validation cohorts (478 [32.3%] vs 69 [10.1%]; p < 0.001). The following variables were associated with LDD in decreasing order: pre-enrolment diarrhea days (55.1%), modified Vesikari score(18.2%), age group (10.7%), vomit days (8.8%), respiratory rate (6.5%), vomiting (6.4%), vomit frequency (6.2%), rotavirus vaccination (6.1%), skin pinch (2.4%) and stool frequency (2.4%). While all models showed good prediction capability, the random forest model achieved the best performance (AUC [95% Confidence Interval]: 83.0 [78.6-87.5] and 71.0 [62.5-79.4]) on the development and temporal validation datasets, respectively. While the random forest model showed slight deviations from perfect calibration, these deviations were not statistically significant (Brier score = 0.17, Spiegelhalter p-value = 0.219).
Conclusions:
Our study suggests ML derived algorithms could be used to rapidly identify children at increased risk of LDD. Integrating ML derived models into clinical decision-making may allow clinicians to target these children with closer observation and enhanced management.
Related Concept Videos
Steps in Outbreak Investigation
Kaplan-Meier Approach

