Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Apr 21, 2026

Emergency Undocking in Robotic Surgery: A Simulation Curriculum
06:48

Emergency Undocking in Robotic Surgery: A Simulation Curriculum

Published on: May 20, 2018

10.6K

Artificial Intelligence in Surgical Education: A Pilot Study Using ASCRS Guideline-Derived Questions.

Shivam Pandya1, Tyler Wilson1, Ryan Meyer1

  • 1Department of Surgery, Los Robles Regional Medical Center, Thousand Oaks, CA, USA.

The American Surgeon
|April 20, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Evaluating the Accuracy of ChatGPT-4o in Addressing Complex Clinical Questions Based on NCCN Guidelines for Rectal Adenocarcinoma.

Journal of surgical oncology·2026
Same author

The complete Genome Sequence of <i>Glaucopsyche lygdamus palosverdesensis</i>, Palos Verdes Blue Butterfly.

Biodiversity genomes·2026
Same author

Demonstrating conservation impacts in California Marine Protected Areas using large-scale participatory science data.

PloS one·2026
Same author

Complete Genome Sequence of Laguna Mountains Skipper (<i>Pyrgus ruralis lagunae</i>).

Biodiversity genomes·2026
Same author

The complete genome sequence of Apodemia mormo langei, Lange's Metalmark Butterfly.

Biodiversity genomes·2026
Same author

Artificial Intelligence in the Trauma Bay: A Pilot Comparison With Surgical Trainees.

The American surgeon·2026

Large language models (LLMs) accurately interpreted specialized surgical guidelines, demonstrating high performance on questions from the American Society of Colon and Rectal Surgeons (ASCRS) guidelines. Both Google Gemini and OpenEvidence showed near-perfect accuracy in this focused study.

Area of Science:

  • Artificial Intelligence in Medicine
  • Clinical Decision Support Systems
  • Medical Informatics

Background:

  • Large language models (LLMs) show promise in general medical assessments.
  • Their ability to interpret and apply subspecialty clinical practice guidelines is not well-understood.
  • Evaluating LLM performance on specific surgical guidelines is crucial for clinical integration.

Purpose of the Study:

  • To assess the accuracy and consistency of Google Gemini and OpenEvidence.
  • To evaluate LLM performance on the 2022 American Society of Colon and Rectal Surgeons (ASCRS) Clinical Practice Guidelines.
  • To compare LLM accuracy against chance performance and inter-model agreement.

Main Methods:

  • Developed 30 multiple-choice questions (MCQs) from ASCRS guidelines for anorectal conditions.
Keywords:
colorectalresident educationsurgical educationsurgical oncology

Related Experiment Videos

Last Updated: Apr 21, 2026

Emergency Undocking in Robotic Surgery: A Simulation Curriculum
06:48

Emergency Undocking in Robotic Surgery: A Simulation Curriculum

Published on: May 20, 2018

10.6K
  • Validated MCQs independently by surgeon reviewers.
  • Presented MCQs to both LLMs under identical conditions for analysis.
  • Main Results:

    • Both Gemini and OpenEvidence achieved 96.7% accuracy (29/30 questions correct).
    • Performance significantly exceeded chance (p < .0001).
    • Models demonstrated perfect inter-model agreement (Cohen's kappa = 1.0), missing the same question.

    Conclusions:

    • Contemporary LLMs exhibit near-perfect accuracy in applying specific subspecialty surgical guidelines.
    • Findings suggest LLMs can accurately interpret guidelines within a limited domain.
    • Further research across multiple guidelines is needed to confirm generalizability.