Related Experiment Videos
Methodological quality and performance of artificial intelligence and machine learning models for preoperative risk
Anisha R Kumar1, Suryakant Singh2, Masaru Ishii3
1Division of Otolaryngology, Department of Surgery, Stony Brook University School of Medicine, Stony Brook, New York, United States; AI Innovation Institute, Stony Brook University, Stony Brook, New York, United States.
Abstract:
Artificial intelligence (AI) and machine learning (ML) are increasingly being applied to preoperative risk prediction in plastic surgery; however, the methodological quality and clinical readiness of these models are yet to be systematically evaluated. This systematic review assessed the quality, risk of bias, and predictive performance of AI/ML preoperative risk prediction models in plastic surgery using the PROBAST+AI framework. Five databases were searched from inception through October 2025. Ten studies met the inclusion criteria, encompassing autologous breast reconstruction (n = 2), alloplastic breast reconstruction (n = 5), head and neck reconstruction (n = 1), burn surgery (n = 1), and aesthetic surgery (n = 1). Random forest was the most frequently used algorithm (n = 4), followed by neural networks (n = 2), deep forest with RUSBoost (n = 1), support vector machine (n = 1), and logistic regression (n = 1). AUC ranged from 0.66 to 0.82 among the 8 studies reporting discrimination. Critical methodological limitations were identified: only 2 studies (20%) performed external validation, 5 of 7 development studies (71.4%) had events per variable <10 indicating inadequate sample size, and 7 studies (70%) did not report model calibration. Pre-reconciliation inter-rater reliability across 102 paired domain-level ratings yielded a Cohen's kappa of 0.240 and Prevalence-Adjusted Bias-Adjusted Kappa of 0.039, consistent with published benchmarks for PROBAST-based systematic reviews. All discrepancies were resolved via structured consensus. Current AI/ML models for preoperative risk prediction in plastic surgery demonstrate variable performance and substantial methodological limitations that preclude clinical implementation. Multi-institutional prospective validation studies with rigorous methodology are needed before clinical adoption.