Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Endoscopic Studies II: Thoracocentesis01:26

Endoscopic Studies II: Thoracocentesis

522
Thoracentesis(Thoracocentesis), commonly known as pleural tap, is a medical procedure where a 22 gauge needle is inserted into the pleural space, the area between the lung and chest wall. This procedure is commonly performed to diagnose or treat various respiratory disorders.
Description
Excess pleural fluid or air may accumulate in some respiratory disorders in the thoracic cavity. To treat pleural effusion, a physician conducts thoracentesis by carefully piercing the chest wall and entering...
522

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Retraction Note: GPT-4 as a Board-Certified Surgeon: A Pilot Study.

Medical science educator·2026
Same author

The collaboration of surgical education fellows' (CoSEF) guide to designing, executing, and publishing surgical education research.

Global surgical education : journal of the Association for Surgical Education·2026
Same author

Losing oneself: Lack of self in depression and its recurrence.

Journal of affective disorders·2025
Same author

The Era "or Error" of Second Localization Procedures.

The breast journal·2025
Same author

Attitudes toward Sustainability among Surgeons at Ten Academic Hospitals in the United States: Are we Ready, Willing, and Incentivized to Change?

Annals of surgery·2025
Same author

GPT-4: A Creative Copilot for Navigating Academic Surgery.

Surgical innovation·2023

Related Experiment Video

Updated: Sep 16, 2025

Eye Tracking During A Complex Aviation Task For Insights Into Information Processing
07:48

Eye Tracking During A Complex Aviation Task For Insights Into Information Processing

Published on: April 4, 2025

578

GPT-4 as a Board-Certified Surgeon: A Pilot Study.

Joshua A Roshal1,2, Caitlin Silvestri3, Tejas Sathe3

  • 1University of Texas Medical Branch, 301 University Blvd, Galveston, TX 77551 USA.

Medical Science Educator
|July 8, 2025
PubMed
Summary

Large language models like GPT-4 show promise for surgical education, excelling in multiple-choice tests but struggling with complex oral board scenarios. Further development is needed for clinical decision-making applications.

Keywords:
Artificial intelligenceCompetency-based educationDigital learningEducation technologyEntrustable professional activitiesLarge language modelsSurgical education

More Related Videos

Evaluating Flight Performance and Eye Movement Patterns Using Virtual Reality Flight Simulator
03:49

Evaluating Flight Performance and Eye Movement Patterns Using Virtual Reality Flight Simulator

Published on: May 19, 2023

1.1K
Intraoperative Gastroscopy for Tumor Localization in Laparoscopic Surgery for Gastric Adenocarcinoma
10:31

Intraoperative Gastroscopy for Tumor Localization in Laparoscopic Surgery for Gastric Adenocarcinoma

Published on: August 9, 2016

12.9K

Related Experiment Videos

Last Updated: Sep 16, 2025

Eye Tracking During A Complex Aviation Task For Insights Into Information Processing
07:48

Eye Tracking During A Complex Aviation Task For Insights Into Information Processing

Published on: April 4, 2025

578
Evaluating Flight Performance and Eye Movement Patterns Using Virtual Reality Flight Simulator
03:49

Evaluating Flight Performance and Eye Movement Patterns Using Virtual Reality Flight Simulator

Published on: May 19, 2023

1.1K
Intraoperative Gastroscopy for Tumor Localization in Laparoscopic Surgery for Gastric Adenocarcinoma
10:31

Intraoperative Gastroscopy for Tumor Localization in Laparoscopic Surgery for Gastric Adenocarcinoma

Published on: August 9, 2016

12.9K

Area of Science:

  • Artificial Intelligence in Medical Education
  • Surgical Training Technologies
  • Large Language Models (LLMs)

Background:

  • Large language models (LLMs) show potential for revolutionizing surgical education.
  • Skepticism regarding LLM accuracy and reliability hinders adoption in medical training.
  • GPT-4's performance on multiple-choice questions is known, but its clinical judgment in oral examinations is less understood.

Purpose of the Study:

  • To evaluate GPT-4's general surgery knowledge using mock written and oral board-style examinations.
  • To identify areas for improvement in LLMs for surgical education and practice.
  • To assess the reliability of GPT-4 in simulating high-stakes clinical decision-making scenarios.

Main Methods:

  • GPT-4 answered 250 multiple-choice questions (MCQs) from the Surgical Council on Resident Education (SCORE) question bank.
  • GPT-4 navigated 4 oral board scenarios based on Entrustable Professional Activities (EPA) topic list.
  • Responses were independently assessed for accuracy by two former oral board examiners.

Main Results:

  • GPT-4 achieved 78.8% accuracy on MCQs, indicating a 92% probability of passing the American Board of Surgery Qualifying Examination (ABS QE).
  • GPT-4 committed critical failures in 75% of oral board scenarios (3 out of 4 cases).
  • Key failure points included incorrect timing of interventions and inappropriate surgical recommendations.

Conclusions:

  • GPT-4's high MCQ performance aligns with previous findings, but it demonstrated limitations in generating accurate long-form content for oral examinations.
  • Significant improvements are required in LLM performance for complex clinical decision-making.
  • Future research should focus on specialized datasets and advanced reinforcement learning to enhance LLM capabilities in surgical contexts.