Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 14, 2026

Sleeve Gastrectomy in Mice using Surgical Clips
05:16

Sleeve Gastrectomy in Mice using Surgical Clips

Published on: November 14, 2020

Comparative Performance of GPT-5 and Gemini Models in Decision Support for Bariatric Surgery: A Simulation-Based

Yahya Kemal Çalışkan1, Fatih Başak2

  • 1Department of General Surgery, University of Health Sciences, Kanuni Training and Research Hospital.

Surgical Laparoscopy, Endoscopy & Percutaneous Techniques
|June 12, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Bedside Triage by Large Language Models in Acute Pancreatitis: A Scenario-Based Comparative Evaluation of GPT-4, GPT-5, and Gemini.

World journal of surgery·2026
Same author

Splenectomy During Cytoreductive Surgery: Marker of Surgical Burden or Independent Predictor of Morbidity?

The American surgeon·2026
Same author

When Wrong Answers Matter: Consequence-Weighted Evaluation of Large Language Models for ERCP Triage.

The American surgeon·2026
Same author

When algorithms triage trauma: Diagnostic accuracy, undertriage risk, and prompt fragility in frontier large language models.

Surgery·2026
Same author

Trends in upper GI pathology over 18 years: a retrospective cross-sectional study.

BMC gastroenterology·2026
Same author

Management of Appendiceal Inflammatory Mass: Nonoperative Treatment, Malignancy Risk, and Surveillance.

The American surgeon·2026

GPT-5 demonstrated superior performance in bariatric surgery decision simulations compared to Gemini. This study highlights the potential of AI decision support but emphasizes the need for clinician oversight due to persistent safety concerns.

Area of Science:

  • Medical Artificial Intelligence
  • Surgical Decision Support Systems
  • Large Language Models (LLMs)

Background:

  • Reliability of AI in high-risk surgical fields like bariatric surgery is not well-established.
  • Previous AI studies focused on patient education or prediction, not perioperative decision simulation.
  • This study addresses the gap by evaluating AI in safety-focused surgical decision-making.

Purpose of the Study:

  • To compare the performance of GPT-5 and Gemini in simulated bariatric surgery decision-making.
  • To assess AI accuracy, safety, guideline adherence, and reasoning in surgical contexts.
  • To evaluate the potential of advanced LLMs as supervised decision-support tools.

Main Methods:

  • A simulation-based paired-comparison study using 60 expert-validated bariatric surgery vignettes.
Keywords:
GPT-5Geminiartificial intelligencebariatric surgeryclinical decision supportlarge language modelspatient safetysimulation

More Related Videos

A Murine Model of Vertical Sleeve Gastrectomy
06:47

A Murine Model of Vertical Sleeve Gastrectomy

Published on: December 18, 2017

Single-Anastomosis Duodeno-Ileal Bypass with Sleeve Gastrectomy Model in Mice
06:40

Single-Anastomosis Duodeno-Ileal Bypass with Sleeve Gastrectomy Model in Mice

Published on: February 10, 2023

Related Experiment Videos

Last Updated: Jun 14, 2026

Sleeve Gastrectomy in Mice using Surgical Clips
05:16

Sleeve Gastrectomy in Mice using Surgical Clips

Published on: November 14, 2020

A Murine Model of Vertical Sleeve Gastrectomy
06:47

A Murine Model of Vertical Sleeve Gastrectomy

Published on: December 18, 2017

Single-Anastomosis Duodeno-Ileal Bypass with Sleeve Gastrectomy Model in Mice
06:40

Single-Anastomosis Duodeno-Ileal Bypass with Sleeve Gastrectomy Model in Mice

Published on: February 10, 2023

  • Evaluation of GPT-5 and Gemini under standardized conditions.
  • Independent scoring of AI outputs by three bariatric surgeons for accuracy, safety, and guideline adherence.
  • Main Results:

    • GPT-5 significantly outperformed Gemini in accuracy (87.3% vs. 72.4%) and guideline adherence (78% vs. 46%).
    • Gemini produced unsafe recommendations four times more often than GPT-5 (8.4% vs. 2.1%).
    • GPT-5 showed substantial expert concordance (κ=0.71), while Gemini had moderate concordance (κ=0.54).

    Conclusions:

    • GPT-5 demonstrated superior accuracy, safety, and guideline alignment in bariatric surgery simulations.
    • Advanced LLMs show promise for supervised surgical decision support.
    • Clinical deployment requires model transparency, validation, and essential clinician oversight.