Related Experiment Videos
A Multidisciplinary Team-Based Large Language Model Framework for Predicting Postoperative Neurological Complications
Jili Li1, Julin Zhang2, Xingrui Tao3
1Department of Thoracic Surgery and Institute of Thoracic Oncology, West China Hospital, Sichuan University, Chengdu, Sichuan, China.
Background:
Postoperative neurological complications (PNCs) after acute type A aortic dissection (ATAAD) surgery are clinically emergent and require multidimensional perioperative risk assessment. Large language models (LLMs) have shown potential in clinical prediction, but the incremental value of structured multiagent collaboration remains unclear.
Objective:
This study aimed to develop and validate a multidisciplinary team (MDT)-based LLM framework for predicting PNC after ATAAD surgery and to compare its performance with that of traditional machine learning (ML) models and single-agent LLM settings.
Methods:
A retrospective cohort from January 2020 to June 2024 (N=763) was randomly divided into a training set (n=533) and an internal validation set (n=230). A prospective cohort from July 2024 to June 2025 (n=120) was used for prospective validation. The outcome was PNC, defined as stroke, cerebral hemorrhage, paraplegia, or coma. Population-level in-context learning used outcome-stratified summary statistics from the training cohort, including predictor distributions in patients with and without PNC and between-group P values. Two LLMs (DeepSeek-V3 and ChatGPT [GPT-5]) were evaluated under 4 settings: no MDT without in-context learning, MDT without in-context learning, no MDT with in-context learning, and MDT with in-context learning. Model performance was assessed using the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, accuracy, F1-score, and Brier score.
Results:
PNC occurred in 13.0% (99/763) of patients in the retrospective cohort and 15.8% (19/120) in the prospective cohort. Among the ML models, the random forest achieved the highest AUC in the internal validation set (AUC 0.7857, 95% CI 0.6945-0.8768). Among the LLM configurations, GPT-5 with MDT and in-context learning achieved the highest AUC (AUC 0.8419, 95% CI 0.7398-0.9440), with a sensitivity of 80.00%, specificity of 87.00%, and Brier score of 0.0876. Although this configuration showed a significantly higher AUC than the fully unaided GPT-5 baseline (P=.006), this improvement reflected the combined effect of MDT and in-context learning, as adding the MDT framework within matched settings did not significantly improve AUC over corresponding single-agent settings in either cohort. The AUC of GPT-5 with MDT and in-context learning was also not significantly higher than that of the random forest model (P=.18). Word frequency analysis showed that the generated rationales following in-context learning more frequently mentioned variables that were statistically significant in the provided context.
Conclusions:
The GPT-5 configuration combining MDT-style collaboration with population-level in-context learning showed promising discrimination for PNC prediction after ATAAD surgery. However, the MDT layer did not produce a statistically significant incremental improvement in AUC over the corresponding single-agent settings, and its predictive contribution remains to be confirmed. The framework generated structured, role-specific rationales, supporting its further evaluation as a proof-of-concept approach for postoperative risk stratification. Larger multicenter studies are required before routine clinical implementation.