Related Experiment Video
Updated: Mar 21, 2026

Endoscopic Endonasal Trans-sphenoidal Approach: Minimally Invasive Surgery for Pituitary Adenomas
Published on: January 17, 2018
AI at the Sella Turcica: Multi-Model Large Language Model Evaluation in Pituitary Adenomas
Aynur Aliyeva1,2, Edin Nevzati1,3,4, Fabio Grassia1,3
1Department of Surgery, Denver Health and Hospital Authority, Denver, CO, USA.
Introduction:
Large language models (LLMs) are explored as clinical decision-support tools in complex medical fields. However, their reliability and clinical usefulness in multidisciplinary management of pituitary adenomas remain insufficiently evaluated using validated, clinician-based frameworks.
Research Question:
Do LLMs differ in informational quality, clinical reasoning, and expert satisfaction when applied to pituitary adenoma-related clinical scenarios?
Materials And Methods:
A prospective comparative study evaluated three LLMs: ChatGPT-5.0, Claude Opus 4.1, and Gemini 2.5 Flash. A standardized prompt set covering general knowledge, surgical decision-making, endocrine evaluation, patient education, and MRI-based scenarios was submitted to each model identically. Outputs were anonymized and independently assessed by 10 board-certified doctors using three validated instruments: the Quality Assessment of Medical Artificial Intelligence (QAMAI), the Artificial Intelligence Performance Instrument (AIPI), and the Artificial Intelligence Satisfaction and Performance Evaluation Questionnaire (AISPE-Q).
Results:
Claude Opus 4.1 achieved the highest performance across all major domains. Aggregate QAMAI scores were highest for Claude Opus 4.1 (4.39 ± 0.66), compared with ChatGPT-5.0 (4.12 ± 0.74) and Gemini 2.5 Flash (4.07 ± 0.76; p = 0.018). Clinical reasoning assessed by AIPI was superior for Claude Opus 4.1 versus Gemini 2.5 Flash and ChatGPT-5.0. Strong correlations were observed between informational quality, reasoning performance, and satisfaction.
Discussion And Conclusion:
LLMs exhibit significant variability in performance when managing pituitary adenomas. Claude Opus 4.1 demonstrated the highest levels of informational quality, reasoning depth, and expert trust. While LLMs may serve as supportive adjuncts in multidisciplinary pituitary care, structured evaluation and expert oversight remain essential before clinical integration.
Level Of Evidence:
2 - Prospective comparative diagnostic accuracy study.

