Related Experiment Video
Updated: Aug 6, 2026

08:26
Clinical Application of Single-Surgeon, Three-Port, Laparoscopic Resection for Colorectal Cancer with Natural Orifice Specimen Extraction
Published on: March 24, 2023
Diagnostic Ambiguity in Appendicitis: Comparing Failure Modes of Large Language Models and Surgeons
Yahya Kemal Çalışkan1, Fatih Başak2, Olgun Erdem2
1Department of General Surgery, University of Health Sciences, Kanuni Training and Research Hospital, Istanbul, Turkey.
The American Surgeon
|July 25, 2026
Summary
Frontier large language models (LLMs) show varied diagnostic accuracy in simulated appendicitis cases. While GPT-5 performed well, human specialists are still crucial for safe emergency triage.
Area of Science:
- Artificial Intelligence in Medicine
- Diagnostic Accuracy
- Emergency Medicine
Background:
- Triage of suspected acute appendicitis is challenging in atypical presentations, such as pediatric, geriatric, or cases with subtle inflammatory markers.
- Frontier large language models (LLMs) are being explored for emergency decision-making, but their clinical readiness requires rigorous evaluation beyond simple vignette comparisons.
- This study addresses the need to assess LLM diagnostic behavior, triage safety, and error patterns in standardized, diagnostically complex appendicitis scenarios.
Purpose of the Study:
- To evaluate and compare the diagnostic accuracy and triage safety of frontier LLMs against general surgery specialists in simulated appendicitis cases.
- To characterize the error profiles and failure modes of both AI systems and human experts in ambiguous diagnostic scenarios.
Main Methods:
- A comparative, cross-sectional diagnostic-accuracy simulation study was conducted using 150 standardized emergency department vignettes (90 appendicitis, 60 mimics), focusing on difficult cases.
- Six frontier AI systems and two board-certified general surgery specialists independently evaluated each vignette using a standardized triage prompt.
- Key outcomes measured included diagnostic accuracy, sensitivity, specificity, inter-rater agreement, triage-priority accuracy, and a qualitative error taxonomy.
Main Results:
- Significant differences in diagnostic performance were observed across AI systems and specialists (P < 0.001).
- GPT-5 demonstrated the highest diagnostic accuracy (95.3%), outperforming GPT-4o (82.0%) and slightly exceeding specialist consensus (92.7%).
- GPT-5 maintained high accuracy (93.8%) in pediatric and geriatric cases, surpassing specialists (90.0%) and GPT-4o (71.2%). Distinct AI failure modes, like hallucinations, were noted.
Conclusions:
- Performance of frontier LLMs in simulated appendicitis diagnosis varies significantly, with GPT-5 showing notable accuracy.
- The study highlights distinct human and AI failure modes in diagnostically ambiguous situations, emphasizing the need for careful interpretation of AI performance.
- Findings support supervised, diagnosis-specific AI evaluation rather than autonomous deployment in emergency triage settings.
Related Concept Videos
Appendicitis-II: Diagnostic Studies and Management
Diagnosing and managing appendicitis requires a structured and comprehensive approach that spans from initial assessment to postoperative care. Here is an overview of the process:
Diagnosing Appendicitis
It requires a multifaceted approach, starting with a detailed physical examination to pinpoint the location and nature of the pain and identify any associated symptoms. Laboratory tests play a crucial role. A complete Blood Count (CBC) typically reveals leukocytosis (an increased number of...
Diagnosing Appendicitis
It requires a multifaceted approach, starting with a detailed physical examination to pinpoint the location and nature of the pain and identify any associated symptoms. Laboratory tests play a crucial role. A complete Blood Count (CBC) typically reveals leukocytosis (an increased number of...
Appendicitis
Appendicitis is an acute inflammatory condition of the vermiform appendix, most commonly caused by obstruction of its lumen. The appendix is a narrow, blind-ended pouch that extends from the cecum, making it particularly prone to obstruction. Causes include fecaliths, lymphoid hyperplasia (often after viral infections), parasites, tumors, or foreign bodies. This obstruction initiates a cascade of pathological changes.Luminal Obstruction and Early InflammationAfter obstruction, normal mucosal...
Appendicitis-I: Introduction
The appendix, a small, narrow, blind tube extending from the inferior part of the cecum, is widely regarded as a vestigial organ, having lost much of its original function through evolution. Despite its diminished role, the appendix can become inflamed, a condition known as appendicitis.
Etiology: Appendicitis can arise from various causes, primarily rooted in the obstruction of the appendix lumen. Factors contributing to this obstruction include fecal accumulation, lymphoid hyperplasia and, in...
Etiology: Appendicitis can arise from various causes, primarily rooted in the obstruction of the appendix lumen. Factors contributing to this obstruction include fecal accumulation, lymphoid hyperplasia and, in...