Related Experiment Videos
Development and validation of a machine learning-based risk prediction model for PICC-associated bloodstream
Ruiqing Song1, Zhirui Li1, Denghui Ma1
1Department of Neonatal Intensive Care, The Third Affiliated Hospital of Zhengzhou University, Zhengzhou, China.
Objective:
To develop and validate a machine learning-based model for predicting the risk of peripherally inserted central catheter (PICC)-associated bloodstream infection (CRBSI) in preterm infants.
Methods:
This retrospective multicenter study included 151 preterm infants with CRBSI and 302 matched controls from a tertiary hospital (2017-2024), randomly divided into a training set (n = 317) and an internal validation set (n = 136). An additional 96 cases from four tertiary hospitals were used for external validation. Eight significant predictors identified by univariate analysis were used to construct five models: Logistic Regression (LR), Extreme Gradient Boosting (XGBoost), Decision Tree (DT), Support Vector Machine (SVM), and Random Forest (RF). Model performance was evaluated using the area under the curve (AUC), accuracy, precision, recall, and F1 score. The Shapley Additive Explanations (SHAP) was applied for model interpretation.
Results:
Eight variables were identified as key predictors, including puncture duration, catheterized vein, duration and frequency of mechanical ventilation, catheter dwell time, fetal distress, hypoalbuminemia, and antibiotic use within 24 h after birth. Infection rates increased markedly with prolonged puncture duration (6% for < 30 min vs 78% for 30-60 min vs 93% for >60 min) and catheter dwell time (21% for <14 days vs 32% for 14-21 days vs 47% for >21 days). Femoral vein catheterization showed the highest infection rate (82%). In the internal validation set, the AUCs of LR, XGBoost, DT, SVM, and RF were 0.89, 0.86, 0.87, 0.88, and 0.95, respectively; in the external validation set, they were 0.89, 0.88, 0.87, 0.93, and 0.96. The RF model achieved the highest accuracy in both internal (0.91) and external (0.90) validation sets, while SVM showed the highest external accuracy (0.92). Precision ranged from 0.79 to 0.89, recall from 0.27 to 0.31, and F1 scores from 0.42 to 0.45. SHAP analysis showed that puncture duration was the most important predictor, followed by catheterized vein and mechanical ventilation-related variables.
Conclusions:
The RF model demonstrated superior performance in predicting CRBSI risk in preterm infants with PICC placement. This model may facilitate early identification of high-risk patients and support clinical decision-making.