Related Experiment Video
Updated: Apr 7, 2026

A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
Multicentre validation and clinical interpretation of an explainable gradient-boosting model for dental-implant
Waad Kheder1, Moussa Leblouba2, Renita Rego1
1College of Dental Medicine, University of Sharjah, P.O.Box. 27272, Sharjah, United Arab Emirates.
Objective:
To evaluate, on unseen multicentre data, the discrimination, calibration, clinical utility, and transparency of an interpretable gradient-boosting model for implant survival/failure prediction, and to report centre-wise transportability with an internal-external (leave-one-centre-out) validation.
Materials And Methods:
We analysed 910 unique implants from three centres (A, B, C) and created a stratified 80/20 hold-out test set (N=182). Within the training split, we performed randomized hyper-parameter search (five-fold cross-validation; macro-F1 scoring) and tuned a single decision threshold for failure-positive classification by maximizing validation macro-F1. Discrimination was assessed with accuracy, class-wise precision/recall/F1, ROC-AUC, and PR-AUC; uncertainty was quantified by 1,000-sample bootstrap confidence intervals. Calibration was evaluated with Brier score and logistic calibration (generalized linear models for intercept-in-the-large and slope) plus reliability diagrams. Clinical utility was examined using decision-curve analysis with bootstrap confidence bands. Generalizability was tested with leave-one-centre-out (LOCO). Post-hoc interpretability used SHAP for global and case-level explanations. All analyses were implemented in a single public script and exported to a machine-readable workbook for replication.
Results:
On the pooled hold-out test (N=182), the model achieved accuracy 0.8736; macro-F1 0.8736; failure-class precision 0.8901 and recall 0.8617; success-class precision 0.8571 and recall 0.8864. ROC-AUC was 0.9253 and PR-AUC 0.9090. Bootstrap 95 % CIs were accuracy 0.824-0.923 and ROC-AUC 0.881-0.963. Internal-external validation showed robust centre-wise performance: accuracy 0.8599-0.9211 and ROC-AUC 0.9034-0.9352 across centres A-C. Calibration analyses indicated acceptable agreement between predicted and observed risks on pooled and centre-held-out sets, and decision-curve analysis demonstrated positive net benefit versus "treat-all/none" across clinically relevant thresholds. SHAP identified diabetes status, bone density, and tobacco smoking among the most influential variables, and a clinician-oriented case study illustrated how feature values drove an individual failure prediction.
Conclusion:
An interpretable gradient-boosting model delivered strong and balanced discrimination on unseen multicentre data, acceptable calibration, and net clinical benefit over a wide threshold range. Consistent LOCO results support transportability across centres, and the alignment between SHAP explanations and known risk factors favours clinical trust. All code, tuned parameters, and exports are provided to facilitate external verification.
Clinical Relevance:
Clinics can select operating thresholds that match their tolerance for missed failures versus false alarms; at commonly used ranges, decision-curve analysis indicates measurable net benefit. SHAP explanations provide patient-specific rationale to support counselling, follow-up planning, and shared decision-making.

