Related Experiment Video
Updated: Aug 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative performance of large language models for foot and ankle clinical decision support: A multi-rater
Kwang Hwan ParK1, Dong Woo Shim1, Jin Woo Lee1
1Department of Orthopaedic Surgery, Yonsei University College of Medicine, Seoul, Republic of Korea.
Background:
This study compared seven large language models (LLMs) to identify the optimal model for foot and ankle clinical decision support.
Methods:
The LLMs answered 20 multiple-choice questions (MCQs) and 20 open-ended clinical questions. MCQ accuracy was recorded. Three blinded foot and ankle surgeons evaluated the open-ended responses based on accuracy, completeness, and clinical relevance (total score range: 3-21).
Results:
GPT-o3 and GPT-5 Thinking achieved the highest MCQ accuracy (95%). For open-ended evaluations, mean total scores differed significantly across the models (p < 0.001). GPT-5 Thinking (19.98 ± 0.66) and GPT-o3 (19.67 ± 0.51) attained the highest scores, significantly outperforming the other five models, with no statistical difference between these top two performers.
Conclusion:
LLM performance for foot and ankle disorders differed substantially across models. GPT-5 Thinking and GPT-o3 exhibited superior performance under the tested conditions, emphasizing the necessity of targeted model selection for clinical decision-making. However, because LLMs evolve rapidly, these findings should be interpreted as a time-specific snapshot rather than as fixed or generalizable rankings of model performance.