Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Assessment of Airway, Skin Color, and Use of Accessory Muscles01:30

Assessment of Airway, Skin Color, and Use of Accessory Muscles

1.9K
A thorough assessment of respiratory health is paramount in clinical settings to identify and manage respiratory distress and ensure adequate oxygenation. This article elaborates on the critical aspects of respiratory evaluation, including airway assessment, skin color examination, and the observation of accessory muscle use, which are integral to effectively diagnosing and managing patients with respiratory conditions.
Introduction
The initial evaluation of a patient's respiratory system...
1.9K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The Long-term Radiographic Fate of the Chronically ACL-Deficient Knee: Response.

The American journal of sports medicine·2026
Same author

Is Higher Surgeon Volume Associated With Lower Complication and Revision Risk After Distal Radius Fracture Surgery? A Population-based Cohort Study of 13,389 Patients.

Clinical orthopaedics and related research·2026
Same author

Comparing the Efficacy and Safety of Intra-articular Injection Treatments for Hip Osteoarthritis: A Systematic Review and Network Meta-analysis.

Orthopaedic journal of sports medicine·2026
Same author

Assessment of Laxity in Multiligamentous Knee Injuries Using Stress Radiographs: A Simple Technique for a Busy Clinic.

Video journal of sports medicine·2026
Same author

Recombinant human growth hormone (rHGH) for muscle enhancement in knee osteoarthritis: protocol for a pilot, randomised placebo-controlled trial.

BMJ open·2026
Same author

The Long-term Radiographic Fate of the Chronically ACL-Deficient Knee: A Systematic Review and Meta-analysis of Matched Cohort Studies.

The American journal of sports medicine·2026

Related Experiment Video

Updated: Apr 9, 2026

Utilizing a 3D Printed Laparoscopic Nissen Fundoplication Model to Shorten a Resident's Learning Curve
08:21

Utilizing a 3D Printed Laparoscopic Nissen Fundoplication Model to Shorten a Resident's Learning Curve

Published on: August 15, 2025

825

Large Language Models Outperform PGY-5 Residents on the Orthopaedic In-Training Examination: A Comparative Analysis

Rushil Dave1, Nimit Vediya, Naman Sharma

  • 1From the Temerty Faculty of Medicine (Dave, Vediya, Sharma), University of Toronto, Toronto, ON, Canada, the Division of Orthopaedic Surgery (Shah, Whelan, Wolfstadt), Department of Surgery, University of Toronto, Toronto, ON, Canada, the Division of Orthopaedic Surgery (Whelan), Department of Surgery, St. Michael's Hospital, Unity Health Network, Toronto, ON, Canada, the University of Toronto Orthopaedic Sports Medicine (Whelan), Toronto, ON, Canada, and the Division of Orthopaedic Surgery (Wolfstadt), Department of Surgery, Mount Sinai Hospital, Toronto, ON, Canada.

The Journal of the American Academy of Orthopaedic Surgeons
|April 8, 2026
PubMed
Summary

ChatGPT demonstrated superior performance on the 2024 Orthopaedic In-Training Examination (OITE), achieving postgraduate year five resident level. This highlights the potential of advanced large language models (LLMs) in medical education and orthopaedic practice.

More Related Videos

The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve
09:51

The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve

Published on: September 7, 2022

3.8K
Mechanical Ventilation Boot Camp Curriculum
07:36

Mechanical Ventilation Boot Camp Curriculum

Published on: March 12, 2018

10.8K

Related Experiment Videos

Last Updated: Apr 9, 2026

Utilizing a 3D Printed Laparoscopic Nissen Fundoplication Model to Shorten a Resident's Learning Curve
08:21

Utilizing a 3D Printed Laparoscopic Nissen Fundoplication Model to Shorten a Resident's Learning Curve

Published on: August 15, 2025

825
The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve
09:51

The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve

Published on: September 7, 2022

3.8K
Mechanical Ventilation Boot Camp Curriculum
07:36

Mechanical Ventilation Boot Camp Curriculum

Published on: March 12, 2018

10.8K

Area of Science:

  • Artificial Intelligence in Medicine
  • Medical Education Technology
  • Orthopaedic Surgery Assessment

Background:

  • Large language models (LLMs) show increasing use in medical education and assessments.
  • Previous LLMs performed at a first-year resident level on the 2022 Orthopaedic In-Training Examination (OITE).
  • Recent LLM advancements and image processing capabilities warrant further evaluation.

Purpose of the Study:

  • To evaluate the performance of six contemporary large language models (LLMs) on the 2024 Orthopaedic In-Training Examination (OITE).

Main Methods:

  • Six LLMs (ChatGPT, Gemini, Grok, Mistral, DeepSeek, Llama) were assessed on 203 multiple-choice questions.
  • Models capable of image and text processing were evaluated alongside text-only models.
  • Performance metrics included accuracy, image interpretation, and logical consistency, stratified by question difficulty.

Main Results:

  • ChatGPT achieved the highest accuracy (74.9%), logical consistency (72.4%), and image interpretation (73.8%).
  • Performance declined across all models as question difficulty increased.
  • Logical consistency strongly correlated with overall accuracy and correctness (P < 0.00001).

Conclusions:

  • ChatGPT demonstrated the highest performance, comparable to a postgraduate year five resident.
  • These findings suggest significant potential for LLM integration in orthopaedic clinical settings, electronic medical records, and surgical planning.
  • Future research should focus on improving performance on difficult questions and developing specialized orthopaedic LLMs.