Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Artificial Intelligence in Surgical Education: A Pilot Study Using ASCRS Guideline-Derived Questions.

The American surgeon·2026
Same author

Cross-Modality Differences In Pancreatic Cyst Size Measurements May Affect Surgical Decision-Making.

The American surgeon·2026
Same author

AI Assessment and Management of Visceral Aneurysms Using ChatGPT-4o-mini: A Pilot Study Examining the Feasibility of Automating the AI Validation Process.

The American surgeon·2026
Same author

Disseminated peritoneal coccidioidomycosis mimicking malignancy in an immunosuppressed patient: a case report.

Journal of surgical case reports·2025
Same author

Telehealth tDCS to reduce cannabis use: A pilot RCT in multiple sclerosis as a framework for generalized use.

Drug and alcohol dependence·2025
Same author

The unmet need for cannabis use disorder treatment in multiple sclerosis: Insights from a nationwide pilot study.

Multiple sclerosis and related disorders·2025

Related Experiment Video

Updated: May 7, 2026

Simulator Training for Endovascular Neurosurgery
08:08

Simulator Training for Endovascular Neurosurgery

Published on: May 6, 2020

3.6K

Artificial Intelligence in the Trauma Bay: A Pilot Comparison With Surgical Trainees.

Shivam Pandya1, Marco Romo1, Tamir Bresler1

  • 1Department of Surgery, Los Robles Regional Medical Center, Thousand Oaks, CA, USA.

The American Surgeon
|May 5, 2026
PubMed
Summary

Large language models (LLMs) show comparable accuracy to junior surgical residents on trauma knowledge. This study suggests artificial intelligence may be a valuable adjunct in surgical education, warranting further validation.

Keywords:
LLMsartificial intelligenceclinical practice guidelinessurgical educationtrauma surgery

More Related Videos

Emergency Undocking in Robotic Surgery: A Simulation Curriculum
06:48

Emergency Undocking in Robotic Surgery: A Simulation Curriculum

Published on: May 20, 2018

9.6K
Creation of a High-Fidelity, Low-Cost, Intraosseous Line Placement Task Trainer via 3D Printing
11:45

Creation of a High-Fidelity, Low-Cost, Intraosseous Line Placement Task Trainer via 3D Printing

Published on: August 17, 2022

2.0K

Related Experiment Videos

Last Updated: May 7, 2026

Simulator Training for Endovascular Neurosurgery
08:08

Simulator Training for Endovascular Neurosurgery

Published on: May 6, 2020

3.6K
Emergency Undocking in Robotic Surgery: A Simulation Curriculum
06:48

Emergency Undocking in Robotic Surgery: A Simulation Curriculum

Published on: May 20, 2018

9.6K
Creation of a High-Fidelity, Low-Cost, Intraosseous Line Placement Task Trainer via 3D Printing
11:45

Creation of a High-Fidelity, Low-Cost, Intraosseous Line Placement Task Trainer via 3D Printing

Published on: August 17, 2022

2.0K

Area of Science:

  • Medical Education
  • Artificial Intelligence in Surgery
  • Trauma Care Guidelines

Background:

  • Large language models (LLMs) excel in general medical knowledge but their accuracy in high-acuity surgical settings like trauma bays is unclear.
  • Assessing LLM performance against human expertise is crucial for understanding their potential applications.
  • Guideline-driven environments require precise and up-to-date knowledge application.

Purpose of the Study:

  • To compare the accuracy of Google Gemini, a contemporary LLM, against junior general surgery residents.
  • To evaluate performance on trauma knowledge questions derived from national practice management guidelines.
  • To determine if LLMs can match resident accuracy in a critical surgical domain.

Main Methods:

  • Thirty multiple-choice questions were developed from current trauma guidelines and validated by trauma surgeons.
  • Six junior general surgery residents (PGY-1-2) completed the assessment.
  • Google Gemini was tested on the same questions under standardized conditions, with accuracy compared using a two-proportion z-test.

Main Results:

  • Residents achieved 87.2% accuracy, while the LLM achieved 90.0% accuracy.
  • No statistically significant difference in accuracy was found between the LLM and junior residents (P = .67).
  • This pilot study indicates comparable performance in this specific knowledge domain.

Conclusions:

  • LLMs demonstrate comparable accuracy to junior surgical residents on trauma guideline-based questions.
  • Guideline-grounded AI shows potential as an adjunct in surgical education.
  • Further validation and power studies are necessary to confirm these preliminary findings and explore broader applications.