Related Experiment Video
Updated: May 7, 2026

08:08
Simulator Training for Endovascular Neurosurgery
Published on: May 6, 2020
3.6K
Artificial Intelligence in the Trauma Bay: A Pilot Comparison With Surgical Trainees
Shivam Pandya1, Marco Romo1, Tamir Bresler1
1Department of Surgery, Los Robles Regional Medical Center, Thousand Oaks, CA, USA.
The American Surgeon
|May 5, 2026
Summary
Large language models (LLMs) show comparable accuracy to junior surgical residents on trauma knowledge. This study suggests artificial intelligence may be a valuable adjunct in surgical education, warranting further validation.
Area of Science:
- Medical Education
- Artificial Intelligence in Surgery
- Trauma Care Guidelines
Background:
- Large language models (LLMs) excel in general medical knowledge but their accuracy in high-acuity surgical settings like trauma bays is unclear.
- Assessing LLM performance against human expertise is crucial for understanding their potential applications.
- Guideline-driven environments require precise and up-to-date knowledge application.
Purpose of the Study:
- To compare the accuracy of Google Gemini, a contemporary LLM, against junior general surgery residents.
- To evaluate performance on trauma knowledge questions derived from national practice management guidelines.
- To determine if LLMs can match resident accuracy in a critical surgical domain.
Main Methods:
- Thirty multiple-choice questions were developed from current trauma guidelines and validated by trauma surgeons.
- Six junior general surgery residents (PGY-1-2) completed the assessment.
- Google Gemini was tested on the same questions under standardized conditions, with accuracy compared using a two-proportion z-test.
Main Results:
- Residents achieved 87.2% accuracy, while the LLM achieved 90.0% accuracy.
- No statistically significant difference in accuracy was found between the LLM and junior residents (P = .67).
- This pilot study indicates comparable performance in this specific knowledge domain.
Conclusions:
- LLMs demonstrate comparable accuracy to junior surgical residents on trauma guideline-based questions.
- Guideline-grounded AI shows potential as an adjunct in surgical education.
- Further validation and power studies are necessary to confirm these preliminary findings and explore broader applications.

