Related Experiment Videos
AI Clinical Decision Making Compared with Fellowship-Trained Hand Surgeons: A Case-Based Study
Claire A Donnelley1, Chris J Lee1, Andrea Halim1
1Department of Orthopaedics and Rehabilitation, Yale School of Medicine, New Haven, CT.
Purpose:
Large language model-based artificial intelligence (AI) platforms have attracted substantial interest as potential clinical adjuncts within surgical specialties, particularly as patients have begun using AI for clinical advice. This study compared the management concordance of four commercially available AI platforms to practicing fellowship-trained hand surgeons using a series of clinical case vignettes within the field of hand surgery.
Methods:
Nine hand surgery clinical vignettes with multiple-choice answer options were developed, including representative radiographs and clinical images. Four cases were classified as straightforward ("easy") and five as clinically complex ("hard"). The cases were then distributed anonymously to fellowship-trained hand surgeons from three institutions using an online survey. Four AI platforms were queried using a standardized prompt: ChatGPT-4 (OpenAI), Claude Opus 4.6 (Anthropic), Gemini (Google DeepMind), and Open Evidence. All AI platforms were accessed in March 2026. Each platform was queried once per case without regeneration to reflect real-world single-query usage. The primary outcome was concordance between each AI platform and the surgeon plurality response for each case.
Results:
Fifteen fellowship-trained hand surgeons completed the survey. Surgeon consensus was high on easy cases (73% to 100% plurality agreement). All four AI platforms achieved concordance with the surgeon plurality on all four easy cases (100%). Performance diverged on complex cases: Claude Opus 4.6 and Gemini each achieved concordance on 3 of 5 complex cases (60%), whereas ChatGPT-4 and Open Evidence each achieved concordance on 1 of 5 (20%). Overall concordance was 7 of 9 (78%) for Claude Opus 4.6 and Gemini and 5 of 9 (56%) for ChatGPT-4 and Open Evidence.
Conclusions:
AI platforms were in agreement with surgeon responses in the majority of low-complexity cases, but had reduced agreement in clinically complex cases.