Related Experiment Video
Updated: May 12, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Comparative Analysis of Large Language Model Performance in Appropriate Diagnostic Imaging Modality Selection
James R Rybczyk1, Kanhai S Amin2, Varun Chamarty3
1Renaissance School of Medicine at Stony Brook University, Stony Brook, New York.
Background:
Large language models (LLMs) show promise for guiding appropriate diagnostic imaging modality selection according to ACR criteria.
Purpose:
This study compared seven LLMs - OpenEvidence (OpenEvidence Inc., Miami, Florida); OpenAI's GPT-5 Thinking and GPT-5 (OpenAI, San Francisco, California); Anthropic's Opus 4.1 and Sonnet 4.5 (Anthropic, San Francisco, California); and Google's Gemini 2.5 Pro and 2,5 Flash (Google LLC, Mountain View, California) - using 50 clinical vignettes to assess accuracy amd clinical reasoning in formulating imaging modality recommendations.
Materials And Methods:
Fifty text-based clinical vignettes were created from ACR guidelines, featuring five variants of 10 different medical complaints with subtle symptomatic or demographic alterations. A 3-point Likert scale was used to evaluate four performance metrics: imaging appropriateness, technical specificity, clinical rationale strength, and citation quality. Readability and word count were also assessed. Two blinded, independent reviewers rated the LLM outputs, with discrepancies resolved via consensus. A third reviewer was included for persistent disagreements. Analysis involved Friedman's test followed by pairwise Wilcoxon signed-rank testing with Holm correction (P < .05).
Results:
Friedman testing demonstrated significant differences across all performance domains (P≤ .031). Appropriateness scores (range 1.60-1.88 out of 2.00) revealed no significant pairwise differences. Technical specificity (range 1.82-2.00) and clinical rationale (range 1.52-1.88) showed no significant pairwise differences. Citation quality (range 0.40-2.00) was the most variable; Gemini 2.5 Pro and Gemini 2.5 Flash hallucinated citations in 80% and 76% of prompts, respectively, performing worse than all other models (P < .001). Readability scores ranged from 15.27 to 22.19, and word counts from 90.10 to 195.02.
Conclusion:
All LLMs selected appropriate imaging modalities using reasonable clinical justification. Citation validity varied widely. Ensuring congruence between clinical reasoning and cited sources is essential before successful implementation.