Related Experiment Video
Updated: Apr 8, 2026

Orthopedic Robot-Assisted Femoral Neck System in the Treatment of Femoral Neck Fracture
Published on: March 3, 2023
Improving Emergency Department Efficiency With Large Language Model-Guided Orthopaedic Triage for Proximal Humerus
Lucy Zhao1,2,3, Ethan Bott1,2, Arya S Rao1,2
1Harvard Medical School, Boston, MA.
Objectives:
To evaluate whether large language models can reduce consults for proximal humerus fractures that do not meet institutional consult criteria.
Design:
Retrospective review.
Setting:
Single-center Level 1 trauma center.
Patient Selection Criteria:
Adults presenting to the emergency department (ED) with isolated proximal humerus fractures over a 2-year period were included. Exclusion criteria were polytrauma, concomitant orthopaedic injuries, pathologic fractures, lack of in-house ED imaging, and fractures missed in the ED.
Outcome Measures And Comparisons:
Generative Pretrained Transformer-4o (GPT-4o) and o4-mini were provided history of present illnesses, physical examinations, and X-ray reports and asked whether orthopaedics consultation was indicated based on institutional criteria (open fracture, tenting skin, neurovascular compromise, or humeral head dislocation). A gold standard was determined by 2 independent authors who retrospectively reviewed each case and reached consensus on consult necessity based on these criteria. Large language models alignment with this standard was compared with performance of real-world providers using generalized linear models. Consult wait time and work relative value unit (wRVU) savings were estimated using the cohort's average wait time and Current Procedural Terminology-based wRVUs for a 30-minute low-to-moderate complexity outpatient consult.
Results:
Three-hundred fifteen patients (99 males and 216 females) were included (average age: 65.1 years, range: 20-100 years). Alignment with consult criteria was 92.4% [95% confidence interval (CI) (88.9%, 94.8%)] for GPT-4o, 94.9% [95% CI (91.9%, 96.9%)] for o4-mini and 32.7% [95% CI (27.7%, 38.1%)] for ED providers. From a baseline of 240 consults, 327.3 wait hours, and 432 wRVUs, GPT-4o could have saved 179 consults, 295.3 wait hours, and 322.2 wRVUs over 2 years. o4-mini could have saved 183 consults, 302.0 wait hours, and 329.4 wRVUs.
Conclusions:
Large language models accurately identified uncomplicated proximal humerus fractures, potentially conserving unnecessary ED orthopaedic consults.
Level Of Evidence:
Prognostic, Level III. See Instructions for Authors for a complete description of levels of evidence.
More Related Videos
05:57A Teleoperated Robotic System-Assisted Percutaneous Transiliac-Transsacral Screw Fixation Technique
Published on: January 6, 2023
04:41Treatment with Locking Intramedullary Nailing for Intertrochanteric Fracture of the Femur Utilizing a New Awl with a Distal Positioner
Published on: June 6, 2025