Related Experiment Video
Updated: Sep 18, 2026

Minimally Invasive Thumb-sized Pterional Craniotomy for Surgical Clip Ligation of Unruptured Anterior Circulation Aneurysms
Published on: August 11, 2015
Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial
Anish Narayan1, Frederick Mariajoseph2,3, Malik Farooq4
1Faculty of Medicine, Monash University, Clayton, Australia. anar0024@student.monash.edu.
Abstract:
Patients increasingly consult large language models (LLMs) before specialist review, yet whether frontier models agree with one another in nuanced clinical domains such as unruptured intracranial aneurysm (UIA) management remains uncharacterised. We therefore quantified inter-model variability across ChatGPT, Gemini and Claude, anchored against neurovascular multidisciplinary team (MDT) consensus and the Unruptured Intracranial Aneurysm Treatment Score (UIATS). Sixty-seven UIA cases referred to our neurovascular service (January-December 2025) were retrospectively analysed. De-identified clinical vignettes were submitted to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking, each run five times. Within-model reproducibility was assessed using Fleiss' κ; inter-model agreement via Cohen's κ and McNemar's test; and each LLM's majority-vote anchored against MDT and UIATS using the same methods. Within-model reproducibility was almost perfect (Fleiss' κ 0.837-0.860). Pairwise inter-model agreement was asymmetric: ChatGPT-Gemini behaved near-identically (Cohen's κ = 0.850, 95% CI 0.71-0.97), whereas Claude diverged from both (κ = 0.688 and 0.667). Recommendations were non-unanimous in 13/67 cases (19.4%); Claude was the sole outlier in 8/13 (conservative in 7). Gemini was significantly more pro-treatment than Claude (McNemar p = 0.0117). Against MDT, Gemini and ChatGPT showed significant pro-treatment propensity (p = 0.0022, p = 0.0153). Claude was significantly more conservative than UIATS (p = 0.0162). Frontier LLMs are highly reproducible internally but diverge from one another asymmetrically. ChatGPT and Gemini behave near-identically while Claude diverges conservatively. Clinicians should anticipate AI-driven treatment expectations and future work should explore prompting, patient sentiment, and multicentre moderators.
