Related Experiment Video
Updated: Sep 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy and Safety of Large Language Models in Endometrial Cancer Decision Making: A Case-Based In Silico
Emanuele Perrone1, Giuseppe Parisi1, Maria Consiglia Giuliano1
1Gynecologic Oncology Unit, Department of Women and Child Health and Public Health, Fondazione Policlinico Universitario Agostino Gemelli IRCCS, Rome, Italy.
Purpose:
To compare the concordance of ChatGPT, Gemini, and Claude with a prespecified expert guideline-based reference standard in fabricated endometrial cancer clinical vignettes under standardized prompting.
Methods:
We conducted a case-based in silico benchmarking study using 35 fabricated postoperative endometrial cancer vignettes representing a broad spectrum of ESGO-ESTRO-ESP 2025 management scenarios. Each vignette was submitted to ChatGPT, Gemini, and Claude in independent chat sessions using the same standardized prompt. The primary end point was concordance with the prespecified expert reference standard, scored as 0 (discordant), 1 (partially concordant), or 2 (fully concordant). Secondary end points were major safety issues and recognition of missing critical information. A post hoc subgroup analysis evaluated the effect of guideline-informed prompting in 12 cases.
Results:
Concordance differed significantly across models (P < .001). Gemini achieved the highest performance, with a median concordance score of 2 (IQR 1-2), compared with 1 (IQR 0-1) for ChatGPT and 0 (IQR 0-1) for Claude. Fully concordant recommendations were generated in 65.7% of cases by Gemini, 2.9% by ChatGPT, and 8.6% by Claude. Major safety issues also differed across models (P = .007), occurring in 22.9% of Gemini responses, 48.6% of ChatGPT responses, and 54.3% of Claude responses. In the subset of vignettes with intentionally missing decisive information, Gemini identified the need for additional data in 83.3% of cases, compared with 50.0% for both ChatGPT and Claude. In the post hoc subgroup analysis, guideline-informed prompting significantly improved concordance for Gemini and Claude.
Conclusion:
Mainstream consumer large language models showed substantially different performance in postoperative endometrial cancer decision making. Although Gemini achieved higher concordance and fewer major safety issues than ChatGPT and Claude, no model demonstrated performance sufficient to support autonomous clinical use in multidisciplinary management.
