Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
ChatGPT and other large language models in laparoscopic cholecystectomy: a multidimensional audit of reliability,
Yusuf Yunus Korkmaz1, Oğuzhan Aydın2, Feyyaz Güngör2
1Department of General Surgery, Başakşehir Çam and Sakura City Hospital, Istanbul, Turkey. yusufyunuskorkmaz@gmail.com.
Introduction:
The rapid uptake of large language models (LLMs) in surgery demands evidence of their reliability when guiding laparoscopic cholecystectomy (LC).
Methods:
An analytical cross-sectional study (April-June 2025) compared five current LLMs (ChatGPT-o3, Claude-Sonnet-4, DeepSeek-V3.5, Gemini-2.5 Flash, and Grok-3) on 24 guideline-derived questions covering the pre-, intra-, and postoperative phases of LC. Four blinded hepatobiliary surgeons rated 120 answers with the eight-item modified DISCERN (mDISCERN, 8-40) and Global Quality Score (GQS, 1-5). Readability was quantified with FRES, FKGL, SMOG, Fog, CLI, and lexical density indices, and inter-rater agreement assessed by two-way ICC.
Results:
Grok delivered the highest mean mDISCERN (36.3 ± 2.3) and GQS (4.76 ± 0.41), whereas Gemini scored lowest (29.0 ± 2.1; 3.58 ± 0.36). DeepSeek produced the most readable output (FRES ≈ 30.6; FKGL ≈ 12.1), while Claude generated the densest, least readable text (negative FRES; FKGL ≈ 18.3). Quality correlated positively with word count and lexical density (ρ ≈ 0.7) but not with syntactic complexity. Surgeon ratings showed good reliability (ICC(2,k) = 0.775; ICC(3,k) = 0.819).
Conclusions:
LLM performance for LC varies markedly; even the best-performing model stops short of full reliability, reinforcing the need for procedure-specific validation before clinical deployment. This multidimensional audit provides a reproducible benchmark for selecting and fine-tuning surgical decision-support LLMs and highlights that terminological richness, rather than sentence complexity, underpins high-quality guidance.
