Related Experiment Video
Updated: Sep 24, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Performance comparison of different machine learning models to predict the risk of postcraniotomy nausea and
Lingyi Wu1, Hongying Pan1, Yihong Xu1
1Department of Nursing, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, Hangzhou, Zhejiang, China.
Background:
Machine learning (ML) models were constructed and validated to predict the risk of postoperative nausea and vomiting (PONV) after craniotomy. We compared different algorithms to provide a reliable tool for clinical decision-making.
Methods:
Data from patients who underwent craniotomy at a tertiary hospital in Zhejiang Province between January 1, 2021 and December 31, 2022 were retrospectively collected as the training set. LASSO (Least Absolute Shrinkage and Selection Operator) regression with 10-fold cross-validation was used to screen for key predictors, and four prediction models were established: Logistic Regression (LR), Random Forest (RF), eXtreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM). Patients from September 1, 2023 to December 22, 2023 served as the temporal validation set. Evaluation of model performance was based on accuracy, precision, specificity, recall, F1-score, Brier score, and the area under the receiver operating characteristic curve (AUC). Temporal validation was performed to assess the generalization of the models and decision curve analysis (DCA) was applied to evaluate clinical utility.
Results:
A total of 768 eligible samples were included, and 14 core predictors were screened by LASSO regression. Of the four constructed risk prediction models, temporal validation based on ML algorithms confirmed that LightGBM achieved optimal comprehensive performance with the highest accuracy, recall, F1-score, and AUC, along with lowest Brier score, indicating that this model had significant advantages in discriminative, calibration, and generalization capabilities. The LR model demonstrated stable performance with high specificity but low recall, indicating a certain risk of missed diagnosis, while the performance of the RF model decreased significantly from the training to validation set, suggesting obvious overfitting and limited clinical application value. The XGBoost model achieved moderate performance, but had a relatively higher Brier score, suggesting moderate prediction error. The DCA results showed that LightGBM offered greater and more stable net benefits across a wider range of clinically relevant threshold probabilities, further supporting its clinical utility.
Conclusion:
LightGBM shows the best overall performance for risk stratification of PONV after craniotomy, and SHAP visualization endows the model with favorable clinical interpretability to facilitate individualized risk prediction.