Related Experiment Videos
Leakage-Controlled and Survey-Weighted Machine Learning for Neonatal Mortality Risk Prediction Using NFHS-5 Data
Moumita Mukherjee1, Talha Ali Khan2, Raja Hashim Ali2
1Institute of International Health, Charité-Universitätsmedizin, 13353 Berlin, Germany.
Abstract:
Background: Neonatal mortality remains uneven across Indian states, while prediction studies using survey data are often limited by class imbalance, data leakage, inadequate calibration, and insufficient consideration of complex survey design. This study developed and rigorously evaluated survey-aware machine-learning models for neonatal mortality risk prediction using NFHS-5 data. Methods: Data from 33,338 children in Bihar, Chhattisgarh, and Uttarakhand were analysed. Household-grouped development/test splitting, repeated grouped nested cross-validation, DHS sampling weights, fold-contained preprocessing, socioeconomic clustering, particle swarm optimisation, resampling, and out-of-fold feature augmentation were applied. Logistic regression, random forest, histogram gradient boosting (HGB), and artificial neural networks were compared using PR-AUC as the primary metric. Calibration, household-bootstrap confidence intervals, decision-curve analysis, prediction timepoint analysis, and leave-one-state-out validation were performed. Results: HGB achieved the highest repeated cross-validation PR-AUC (0.189) and ROC-AUC (0.778). On the untouched test set, ROC-AUC was 0.755 (95% CI 0.718-0.796), and PR-AUC was 0.202 (0.138-0.267), with sensitivity 0.813, specificity 0.517, PPV 0.063, and NPV 0.989. Clustering, PSO, SMOTE, and augmentation added little value. Antenatal performance was weaker, and state-wise transportability varied. Conclusions: Survey-weighted HGB provided the strongest predictive performance, but low PPV and heterogeneous state-level results restrict its use to low-cost screening. Prospective validation is required before deployment.