Related Experiment Video
Updated: Jun 5, 2026

Sleeve Gastrectomy in Mice using Surgical Clips
Published on: November 14, 2020
Large Language Models and Metabolic Bariatric Surgery: A Pilot Concordance Study Between ChatGPT and
Aminah Ahmed1, Shivam Bhanderi2,3, Robyn Westerman4
1Core Surgical Trainee, West Midlands Deanery, Birmingham, UK.
Introduction:
Operation selection in metabolic surgery is a complex decision making process led by a multidisciplinary team that integrates multiple anatomical, clinical, metabolic and psychosocial aspects. The ability of large language models (LLMs) has been proposed to provide capability to act as decision support tools, but their performance in replicating MDT level decision making in metabolic surgery remains uncertain.
Methods:
A retrospective pilot study was performed using anonymised data from 100 patients at a single high-volume UK NHS bariatric centre who underwent surgery up to August 2025. Preoperative demographic, anthropometric, clinical and psychosocial variables were extracted from electronic records. These were provided to ChatGPT-4 Auto using a single standardised prompt instructing the model to act collectively on behalf of the whole bariatric MDT and recommend the most appropriate metabolic operation. These recommendations were compared with formal MDT decisions and with the operation ultimately performed. Concordance was assessed using raw percentage agreement, Cohen's kappa and Stuart-Maxwell tests.
Results:
ChatGPT demonstrated 70% concordance with MDT recommendations. Concordance between MDT and operation performed was 83%, while concordance between ChatGPT and operation performed was 63%. Agreement beyond chance between ChatGPT and MDT recommendations was low (Cohen's kappa 0.036) reflecting class imbalance. Stuart-Maxwell test showed no significant difference in marginal distribution between ChatGPT and MDT recommendations. Both ChatGPT and MDT recommended bypass procedures more frequently than ultimately performed.
Conclusion:
ChatGPT overall demonstrated moderate crude agreement with bariatric MDT decision making in a real world UK NHS cohort of patients. However, limited agreement beyond chance and influence of unmeasured human factors may preclude its use in an autonomous fashion. This study establishes the requirement for larger scale evaluation of LLM clinical reasoning and supports the future exploration of them as adjunctive rather than autonomous decision support tools in metabolic surgery.

