Related Experiment Video
Updated: Jul 10, 2026

Preparing a Mice Model of Severe Acute Pancreatitis via a Combination of Caerulein and Lipopolysaccharide Intraperitoneal Injection
Published on: May 10, 2024
Bedside Triage by Large Language Models in Acute Pancreatitis: A Scenario-Based Comparative Evaluation of GPT-4,
Yahya Kemal Çalışkan1, Fatih Başak2, Olgun Erdem2
1Department of General Surgery, University of Health Sciences, Kanuni Training and Research Hospital, Istanbul, Turkey.
Background:
Early decision-making in acute pancreatitis (AP) involves diagnostic confirmation, early severity triage, escalation thresholds, and initiation of guideline-concordant management under time pressure and incomplete information. Large language models (LLMs) may support structured bedside reasoning, but their clinical usefulness cannot be inferred from guideline knowledge alone.
Methods:
A cross-sectional, scenario-based comparative evaluation was conducted in January 2026 using 20 AP scenarios: 15 refined hypothetical vignettes and 5 de-identified, privacy-modified real-life case patterns. GPT-4, GPT-5, and Gemini received identical single-turn prompts. Model access was through OpenAI API gpt-4-0613, OpenAI API gpt-5, and Google Vertex AI Gemini 1.0 Pro; temperature was set to 0.0, and each prompt was repeated three times per model. Outputs were scored by two independent clinician-raters using a prespecified 1-5 ordinal rubric across guideline concordance, safety, actionability, and data-synthesis quality. Two senior board-certified surgeons independently generated expert reference pathways for comparison.
Results:
GPT-5 achieved the highest guideline concordance (4.28 ± 0.38) and safety (4.20 ± 0.45) profiles. GPT-4 provided the clearest stepwise actionability (4.15 ± 0.48), whereas Gemini showed the strongest data-synthesis quality (4.22 ± 0.52). With deterministic settings, internal consistency across three repeated runs was 100%. All models demonstrated clinically relevant failure modes, particularly unwarranted certainty under missing data; this occurred in 12/20 GPT-4, 7/20 GPT-5, and 15/20 Gemini outputs.
Conclusion:
No model should be used as a stand-alone bedside decision-maker for AP. In this scenario-based early evaluation, GPT-5 was the most safety-aligned model, GPT-4 was the most operationally actionable, and Gemini was strongest for synthesis, but all require clinician oversight, prospective validation, and governance before clinical deployment.
Related Concept Videos
Acute Pancreatitis II: Clinical Manifestations and Management
Chronic Pancreatitis II: Collaborative Care
Assessment:
Acute Pancreatitis I: Introduction
Acute Pancreatitis I: Introduction
Acute pancreatitis is characterized by rapid inflammation of the pancreas, often caused by factors like gallstone blockage or excessive alcohol consumption. Chronic pancreatitis, on the other hand, is a slow, progressive inflammation that may result from long-term alcohol abuse, obstructions in the pancreatic duct, or genetic factors.
The causes of acute pancreatitis include:
