Related Experiment Video
Updated: May 24, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
475
Comparitive performance of artificial intelligence-based large language models on the orthopedic in-training
Andrew Y Xu1, Manjot Singh1, Mariah Balmaceno-Criss1
1Warren Alpert Medical School, Brown University, Providence, RI, USA.
Journal of Orthopaedic Surgery (Hong Kong)
|March 3, 2025
Summary
GPT-4 demonstrated superior performance on orthopedic board-style questions, surpassing the passing threshold for the 2022 Orthopedic In-Training Examination (OITE). This advanced large language model (LLM) outperformed GPT-3.5 and Google Bard, highlighting its potential in clinical applications.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Orthopedic Surgery Training
Background:
- Large language models (LLMs) offer potential clinical applications, but their comparative efficacy on specialized medical examinations is not well-established.
- The performance of different LLMs on orthopedic board-style questions requires thorough investigation to understand their capabilities and limitations.
Purpose of the Study:
- To comparatively evaluate the performance of three leading large language models (LLMs) on orthopedic board-style questions.
- To assess the proficiency of GPT-4, GPT-3.5, and Google Bard against orthopedic resident performance benchmarks and specific question types.
Main Methods:
- Three LLMs (GPT-4, GPT-3.5, and Google Bard) were evaluated using 189 official 2022 Orthopedic In-Training Examination (OITE) questions.
- Performance was analyzed against orthopedic resident scores, with specific attention to higher-order, image-associated, and subject-specific questions.
Main Results:
- GPT-4 exceeded the 2022 OITE passing threshold, performing at the PGY-3 to PGY-5 resident level and significantly outperforming GPT-3.5 and Bard (p < .001).
- GPT-3.5 and Bard did not meet the passing threshold, performing at PGY-1/PGY-2 and PGY-1/PGY-3 levels, respectively.
- GPT-4 demonstrated superior performance on image-associated and higher-order questions compared to GPT-3.5 and Bard (p < .001).
Conclusions:
- The AI-based LLM GPT-4 shows strong capability in answering diverse OITE questions, exceeding the minimum score for the 2022 examination.
- GPT-4 significantly outperforms its predecessors GPT-3.5 and Google Bard on orthopedic board-style questions.
- These findings underscore the potential of advanced LLMs like GPT-4 as valuable tools in medical education and assessment.
Keywords:
ChatGPTGoogle bardartificial intelligencelarge language modelsmedical educationresidency educationsurgical education
