Related Experiment Videos
Development and internal validation of an interpretable machine learning model using first-24-h postpartum nursing
1Department of Obstetrics, The Affiliated Hospital, Southwest Medical University, Luzhou, China.
Background:
Delayed lactogenesis II is a common early postpartum breastfeeding problem after cesarean delivery and may interfere with the establishment of exclusive breastfeeding. Although demographic and obstetric risk factors have been widely studied, the predictive value of routinely recorded early postpartum nursing process variables remains insufficiently explored. This study aimed to develop and internally validate a dynamic 24-h postpartum landmark interpretable machine learning model using clinical, perioperative, neonatal, and first-24-h nursing process variables to predict delayed lactogenesis II by 72 h after cesarean delivery.
Methods:
This single-center retrospective cohort study included women who underwent cesarean delivery at a tertiary hospital between January 2021 and December 2024 and had documented intention to breastfeed. Delayed lactogenesis II was defined as the absence of standardized evidence of lactogenesis onset by 72 h postpartum, based on maternal responses and supporting indicators documented by trained nurses at predefined postpartum assessments. The prediction landmark was set at 24 h postpartum; therefore, only variables available before delivery, during cesarean delivery, or within the first 24 h postpartum were used for model development. Because women who had already experienced lactogenesis II by 24 h were no longer at risk at the prediction landmark, an additional sensitivity analysis was restricted to women without lactogenesis II onset by the 24-h assessment. Candidate predictors were extracted from electronic medical records, perioperative records, neonatal records, nursing assessment documentation, and breastfeeding assessment forms. A complete-case approach was used; 32 women (2.2% of those initially screened) were excluded because of incomplete key predictor data, and no statistical imputation was performed. The final cohort was randomly divided into a training cohort and an internal validation cohort at a ratio of 7:3. Five models were developed and compared: logistic regression, random forest, support vector machine (SVM), XGBoost, and LightGBM. Hyperparameters were selected within the training cohort using repeated stratified five-fold cross-validation, and the internal validation cohort remained untouched until final model evaluation. Model performance was assessed using the area under the receiver operating characteristic curve, accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1 score. Calibration curve analysis, decision curve analysis, and SHAP-based interpretation were performed for the final model.
Results:
A total of 1,286 women were included, of whom 326 developed delayed lactogenesis II, corresponding to an event rate of 25.3%. Among the included women, 71 (5.5%) had already experienced lactogenesis II by the 24-h assessment and had been retained in the primary cohort. The training and validation cohorts included 900 and 386 women, respectively. In the validation cohort, all five models showed acceptable discrimination, with AUCs ranging from 0.816 to 0.836. The support vector machine achieved the highest validation AUC of 0.836, followed by LightGBM, XGBoost, random forest, and logistic regression. XGBoost showed a training AUC of 0.872 and a validation AUC of 0.823, with validation accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1 score of 0.756, 0.694, 0.778, 0.515, 0.882, and 0.591, respectively. Although SVM achieved the numerically highest validation AUC, its discrimination did not differ significantly from that of XGBoost (AUC difference, 0.013; 95% CI, -0.019 to 0.045; DeLong P = 0.421). XGBoost was retained as the final interpretable model because it showed a smaller training-to-validation AUC difference and enabled direct SHAP-based characterization of predictor contributions. In the internal validation cohort, the XGBoost model had a calibration intercept of 0.058, a calibration slope of 1.684 (95% CI, 1.304-2.064), and a Brier score of 0.143. Although the intercept indicated limited systematic miscalibration, the calibration slope and its confidence interval were entirely above the ideal value of 1, indicating non-ideal calibration and under-dispersion of predicted probabilities, with predictions compressed toward the mean. Decision curve analysis indicated greater net benefit than the treat-all and treat-none strategies across clinically relevant threshold probabilities. SHAP analysis identified first breastfeeding time, breastfeeding frequency within 24 h, first skin-to-skin contact time, intraoperative blood loss, and postoperative pain score at 24 h as the most influential predictors. After excluding the 71 women with lactogenesis II onset by 24 h, the landmark-restricted sensitivity cohort comprised 1,215 women. In this restricted population, the XGBoost model achieved a validation AUC of 0.817, and the five leading predictors remained unchanged.
Conclusions:
A dynamic 24-h postpartum landmark interpretable machine learning model incorporating clinical, perioperative, neonatal, and first-24-h nursing process variables showed acceptable discrimination for predicting delayed lactogenesis II by 72 h after cesarean delivery, but its calibration was non-ideal. Although SVM achieved the numerically highest validation AUC, its discrimination did not differ significantly from that of XGBoost. XGBoost was retained as the final interpretable model on the basis of its overall balance of validation discrimination, apparent stability, and direct SHAP-based interpretability. First breastfeeding time, breastfeeding frequency, skin-to-skin contact, intraoperative blood loss, and postoperative pain score may serve as early postpartum risk markers to help identify women who could benefit from intensified breastfeeding support and individualized nursing care. Because the model has undergone only single-center internal validation, external validation and recalibration are required before it can be used to estimate individual risk or support individual-level clinical decision-making.