Related Experiment Video
Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing the Third Wave of Generative AI: Performance of Advanced Models on Text-based Questions From the 2024
Daniel S Hayes1, Alyssa Barré, Ryan D Muchow
1From the University of Kentucky College of Medicine (Hayes), and the University of Kentucky Department of Orthopaedic Surgery and Sports Medicine, Lexington, KY (Barré and Muchow).
Introduction:
Learners are rapidly using generative artificial intelligence (AI) models in their education. We assessed the performance of recently released or updated models on the 2024 American Academy of Orthopaedic Surgeons Orthopaedic In-training Examination for their potential applications in orthopaedic education.
Methods:
Eleven models with recent enhancements in reasoning and research capabilities were evaluated. A total of 119 text-based questions were entered verbatim into each model. Model outputs were recorded as correct or incorrect. Additional analyses included reasoning time, citation accuracy, confidence in answer selection, and comparison with orthopaedic resident performance. References generated by the top performing model were compared with American Academy of Orthopaedic Surgeon Recommended Readings for incorrectly answered questions.
Results:
Ten of 11 AI models exceeded the American Board of Orthopaedic Surgery minimal passing standard (67.7%). Eight models surpassed PGY5 resident performance. OpenAI's o1Pro with Deep Research achieved the highest accuracy (90.8%), outperforming the mean performance of PGY5 residents by 17.8%. More than half of the ResStudy Recommended Readings were cited as supporting references by the top performing model on questions it answered incorrectly. Subspecialty performance varied, with highest accuracy in Shoulder and Elbow and Sports Medicine questions. Longer reasoning times generally correlated with improved accuracy.
Discussion:
Advanced AI models demonstrated substantial improvements over previous generations, with more than half of the tested models exceeding senior orthopaedic resident performance. Improved reasoning and research capabilities highlight the evolving capabilities of AI in medical education, although increased understanding of their utilization by learners is needed. Variability among subspecialties may suggest differences in training data or reasoning capabilities.
Conclusions:
Modern AI models exhibit high proficiency on the Orthopaedic In-training Examination and may serve as valuable supplemental educational tools. Ongoing evaluation is warranted to understand their optimal integration into orthopaedic training while recognizing limitations in clinical reasoning, lived experience, and imaging interpretation.
