Related Experiment Video
Updated: Sep 15, 2026

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Comparing machine learning algorithms for predicting postoperative medical management in prolactinoma surgery: a
Karthik Papisetty1, Chris Donghyun Kim1, Thomas McCaffery1
1Department of Neurosurgery, Emory University, Atlanta, GA, USA.
Purpose:
Surgical resection is an important treatment modality for prolactinomas, yet approximately 20% of patients require postoperative dopamine agonist therapy (DAT) due to prolactin rebound. Accurate preoperative identification of this patient population would better inform surgical decision-making and longitudinal care. We aimed to develop a parsimonious, interpretable machine learning model to predict postoperative DAT requirement following prolactinoma resection.
Methods:
A retrospective cohort of 138 patients who underwent surgical prolactinoma resection at a single institution between 2000 and 2024 was analyzed. Eight classification algorithms were evaluated on 75 preoperative features, including four clinically motivated engineered features. A five-feature model was identified through Gini-impurity-based feature selection. Nested cross-validation with a five-fold outer loop and five-fold inner loop was used for unbiased performance estimation. Shapley Additive Explanations (SHAP) were computed to characterize feature contributions to individual predictions.
Results:
The cohort was nearly evenly split between patients requiring (49.3%) and not requiring (50.7%) postoperative DAT. The five selected features were preoperative serum prolactin, tumor volume, FSH, maximum tumor diameter, and prolactin density. Under nested cross-validation, Random Forest achieved the highest AUROC of 0.792 (95% CI: 0.656-0.928), with XGBoost (0.784), Gradient Boosting (0.780), and Logistic Regression (0.775) performing comparably. A pre-specified clinical logistic-regression baseline achieved an AUROC of 0.735, indicating a modest incremental gain from the machine learning model. Calibration (Brier score 0.199; calibration slope 0.906, intercept -0.022) and decision-curve analysis supported clinical utility across the relevant threshold range. SHAP analysis identified elevated preoperative prolactin, larger tumor volume, and lower FSH as the primary drivers of predicted medication requirements.
Conclusion:
A parsimonious five-feature model achieved modest, consistent discrimination for postoperative DAT requirement following prolactinoma resection, with several algorithms performing comparably and only a modest gain over a conventional clinical baseline. Nested cross-validation and SHAP analysis address key methodological limitations in existing pituitary machine learning literature. External multicenter validation is necessary prior to clinical implementation.