Related Experiment Video
Updated: Jun 12, 2026

The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve
Published on: September 7, 2022
Clinical Implications of ChatGPT-assisted Multimodal Pre-operative Assessment in Elderly Pertrochanteric Fracture
Mitsuaki Noda1, Shunsuke Takahara2, Shinya Hayashi3
1Department of Orthopedics, Himeji Central Hospital, Himeji, Japan.
Introduction:
Given the high mortality associated with femoral trochanteric fractures, reliable pre-operative estimation of post-operative death risk is essential. Commonly used scoring systems demonstrate only moderate predictive ability, largely because they rely on limited and sometimes outdated clinical parameters. ChatGPT (Generative Pre-Training Transformer) may be capable of synthesizing diverse pre-operative clinical information within a single multimodal analytical framework. Therefore, this study aimed to (i) Explore whether ChatGPT-based multimodal integration of pre-operative data can generate clinically interpretable patterns and (ii) Assess reproducibility of generated outputs across two repeated evaluations.
Materials And Methods:
Patients with pertrochanteric fractures were retrospectively reviewed. Demographic variables and image sets, including laboratory data, cardiac reports, and medication records, were uploaded to ChatGPT (GPT-4o and GPT-5). The same standardized prompt was used to generate a 2-month post-operative mortality risk (%) after multimodal integration of demographic and clinical image-based information, and each model was tested twice. Patients were classified as Alive or Dead based on 2-month post-operative status. Generated output values (%) were then examined for between-group separation. Agreement between repeated estimates was assessed using Bland-Altman analysis.
Results:
A total of 134 patients were included. The Alive group comprised 129 patients (106 females; mean age, 87 years), and the Death group comprised 5 patients (3 females; mean age 90 years). ChatGPT-4o showed no significant difference between groups. GPT-5 generated higher output values in the Death group across both repeated evaluations, indicating separation between groups after multimodal data integration. Agreement analysis showed a mean difference (bias) of 1.28% (95% confidence interval [CI], 0.09-2.47) for GPT-4o and 0.18% (95% CI, -0.41-0.76) for GPT-5.
Conclusions:
ChatGPT, particularly GPT-5, demonstrated separation of generated output values between survivors and non-survivors, after integrating diverse pre-operative data. However, this should not be interpreted as evidence of predictive accuracy, primarily due to the severe imbalance in group size. Rather, these findings suggest potential clinical relevance of AI-assisted multimodal assessment as a pathway for addressing complex medical questions in daily practice.
