Related Experiment Video
Updated: Aug 6, 2026

Clinical Application of Single-Surgeon, Three-Port, Laparoscopic Resection for Colorectal Cancer with Natural Orifice Specimen Extraction
Published on: March 24, 2023
Diagnostic Ambiguity in Appendicitis: Comparing Failure Modes of Large Language Models and Surgeons
Yahya Kemal Çalışkan1, Fatih Başak2, Olgun Erdem2
1Department of General Surgery, University of Health Sciences, Kanuni Training and Research Hospital, Istanbul, Turkey.
None:
BackgroundSuspected acute appendicitis is difficult to triage in atypical presentations, including pediatric and geriatric cases and scenarios with muted inflammatory markers. Frontier LLMs are increasingly evaluated in emergency decision-making, but vignette-based comparisons should not be interpreted as evidence of clinical readiness. This study evaluated diagnostic behavior, triage safety, and error profiles of frontier LLMs in standardized appendicitis scenarios, using general surgery specialists as a benchmark.MethodsWe performed a comparative, cross-sectional diagnostic-accuracy simulation study using 150 standardized emergency department vignettes (90 appendicitis and 60 non-surgical mimics), enriched for diagnostically difficult scenarios. Six frontier AI systems and two board-certified general surgery specialists independently evaluated each vignette using a fixed triage prompt. Outcomes were diagnostic accuracy, sensitivity, specificity, inter-rater agreement, triage-priority accuracy, and qualitative error taxonomy.ResultsOverall diagnostic performance differed across cohorts (P < 0.001). GPT-5 achieved the highest accuracy (95.3%), exceeding GPT-4o (82.0%; P < 0.001) and slightly surpassing specialist consensus (92.7%). In pediatric and geriatric vignettes, GPT-5 maintained 93.8% accuracy, compared with 90.0% for specialists and 71.2% for GPT-4o. Error analysis identified distinct failure modes, including search-induced hallucinations in web-augmented systems. These findings were interpreted alongside clinical error severity rather than as product-level superiority.ConclusionIn this vignette-based simulation study, LLM performance varied substantially across systems. The main contribution was characterization of model and human failure modes under diagnostic ambiguity, supporting supervised, diagnosis-specific evaluation rather than autonomous emergency triage deployment.
Related Concept Videos
Appendicitis-II: Diagnostic Studies and Management
Diagnosing Appendicitis
It requires a multifaceted approach, starting with a detailed physical examination to pinpoint the location and nature of the pain and identify any associated symptoms. Laboratory tests play a crucial role. A complete Blood Count (CBC) typically reveals leukocytosis (an increased number of...
Appendicitis
Appendicitis-I: Introduction
Etiology: Appendicitis can arise from various causes, primarily rooted in the obstruction of the appendix lumen. Factors contributing to this obstruction include fecal accumulation, lymphoid hyperplasia and, in...