Related Experiment Video
Updated: Jul 1, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
555
Performance of Two Artificial Intelligence Generative Language Models on the Orthopaedic In-Training Examination
Orthopedics
|March 11, 2024
Summary
Artificial intelligence (AI) language models show varied performance on the Orthopaedic In-Training Examination (OITE). ChatGPT demonstrated higher accuracy, exceeding the resident average, suggesting potential in orthopedic education.
Area of Science:
- Artificial Intelligence
- Medical Education
- Orthopaedic Surgery
Background:
- Generative AI large language models offer potential in healthcare education.
- The Orthopaedic In-Training Examination (OITE) assesses resident progress.
Purpose of the Study:
- To evaluate the performance of AI language models on the 2022 OITE.
- To compare the accuracy of ChatGPT and Bard on orthopaedic in-training assessments.
Main Methods:
- Administered the 2022 OITE to OpenAI's ChatGPT and Google's Bard.
- Inputted questions with and without text descriptions of accompanying images.
Main Results:
- ChatGPT achieved 69.1% accuracy, improving to 77.8% with image descriptions.
- Bard achieved 49.8% accuracy, improving to 58% with image descriptions (P<.0001).
- ChatGPT outperformed the resident average (66%) and excelled in shoulder-related questions.
Conclusions:
- Publicly available AI models exhibit variable accuracy on the OITE.
- AI tools have potential roles in orthopedic education, such as simulating cases and personalizing learning.
- Further research is needed for safe adoption and risk mitigation of AI in orthopedics.

