Related Experiment Video
Updated: Apr 21, 2026

06:48
Emergency Undocking in Robotic Surgery: A Simulation Curriculum
Published on: May 20, 2018
10.6K
Artificial Intelligence in Surgical Education: A Pilot Study Using ASCRS Guideline-Derived Questions
Shivam Pandya1, Tyler Wilson1, Ryan Meyer1
1Department of Surgery, Los Robles Regional Medical Center, Thousand Oaks, CA, USA.
The American Surgeon
|April 20, 2026
Summary
Large language models (LLMs) accurately interpreted specialized surgical guidelines, demonstrating high performance on questions from the American Society of Colon and Rectal Surgeons (ASCRS) guidelines. Both Google Gemini and OpenEvidence showed near-perfect accuracy in this focused study.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Medical Informatics
Background:
- Large language models (LLMs) show promise in general medical assessments.
- Their ability to interpret and apply subspecialty clinical practice guidelines is not well-understood.
- Evaluating LLM performance on specific surgical guidelines is crucial for clinical integration.
Purpose of the Study:
- To assess the accuracy and consistency of Google Gemini and OpenEvidence.
- To evaluate LLM performance on the 2022 American Society of Colon and Rectal Surgeons (ASCRS) Clinical Practice Guidelines.
- To compare LLM accuracy against chance performance and inter-model agreement.
Main Methods:
- Developed 30 multiple-choice questions (MCQs) from ASCRS guidelines for anorectal conditions.
- Validated MCQs independently by surgeon reviewers.
- Presented MCQs to both LLMs under identical conditions for analysis.
Main Results:
- Both Gemini and OpenEvidence achieved 96.7% accuracy (29/30 questions correct).
- Performance significantly exceeded chance (p < .0001).
- Models demonstrated perfect inter-model agreement (Cohen's kappa = 1.0), missing the same question.
Conclusions:
- Contemporary LLMs exhibit near-perfect accuracy in applying specific subspecialty surgical guidelines.
- Findings suggest LLMs can accurately interpret guidelines within a limited domain.
- Further research across multiple guidelines is needed to confirm generalizability.