Machine Learning Models for Predicting Response to Percutaneous Transhepatic Biliary Drainage in Malignant
Pankaj Gupta1, Sonali Thakur2, Abhinandan Kumar2
1Post Graduate Institute of Medical Education and Research, Chandigarh, India. pankajgupta959@gmail.com.
Background And Aims:
Percutaneous transhepatic biliary drainage (PTBD) is a standard palliative intervention for malignant obstructive jaundice, yet 60-75% of patients fail to achieve a clinically meaningful biochemical response. Pre-procedural identification of responders has direct implications for patient selection and shared decision-making. We developed and validated machine learning (ML) models to predict biochemical response to PTBD using pre-procedure clinical and laboratory variables.
Methods:
We conducted a retrospective analysis of 243 patients with malignant obstructive jaundice who underwent PTBD at a tertiary care referral center. Response was defined as ≥ 50% reduction in serum total bilirubin at two weeks post-procedure. Twenty-six pre-procedural features - including demographics, liver biochemistry, renal function, haematological indices, and clinical severity scores - were evaluated. Five ML algorithms (Logistic Regression, Gradient Boosting, XGBoost, LightGBM, and stacking ensembles) were trained with synthetic minority oversampling (SMOTE variants) to address class imbalance. Performance was evaluated by stratified 5-fold cross-validation and a held-out test set (20%).
Results:
Response occurred in 63 of 243 patients (25.9%). The best-performing model (Logistic Regression with L1 regularisation, SMOTE oversampling) achieved an area under the receiver operating characteristic curve (AUC) of 0.714 (95% CI 0.529-0.870) on the held-out test set, with sensitivity 76.9% and specificity 61.1%. Tree-based models demonstrated severe cross-validation overfitting to synthetic samples (CV AUC 0.83-0.92) with substantial test-set degradation (AUC 0.58-0.70), a finding attributable to SMOTE-induced distribution shift in small, imbalanced datasets. There was no statistically significant AUC difference between logistic regression and tree-based models. Calibration assessment revealed overconfident predictions (Brier score 0.24, calibration slope 0.46). SHAP analysis identified prothrombin time index, ascites, creatinine, cancer type, and cholangitis severity as the most influential predictors.
Conclusions:
Pre-procedural ML-based prediction of PTBD response is feasible with a modest AUC. Logistic regression generalises most reliably in this small, class-imbalanced dataset. External validation in larger prospective cohorts is required before clinical deployment.
