Related Experiment Video
Updated: Sep 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Expert Evaluation of Multimodal Large Language Models for Pediatric Rash Diagnosis: Clinical Utility,
Dilara Lahut1, Özlem Erdede1, Rabia Gönül Sezer Yamanel1
1Department of Pediatrics, Zeynep Kamil Women and Children's Diseases Training and Research Hospital, University of Health Sciences, 34668 Istanbul, Türkiye.
Abstract:
Background/Objectives: Multimodal large language models (LLMs) can interpret clinical text and images, but their performance in pediatric rash assessment remains uncertain. This study compared the clinical utility, safety, information quality, diagnostic correctness, and readability of ChatGPT, Gemini and Grok. Methods: Fifteen content-validated pediatric rash vignettes with brief histories and anonymized photographs were submitted once to each platform using a standardized zero-shot prompt. Three pediatricians blinded to platform identity independently rated the 45 responses using a five-point Clinical Utility and Safety (CUS) scale and a five-item modified DISCERN instrument. Diagnostic correctness was assessed descriptively; platform comparisons used Friedman tests with Bonferroni-adjusted Wilcoxon tests when appropriate. Results: Overall, 82.2% of CUS ratings were in categories 4-5 and 83.0% of modified DISCERN scores were ≥20/25; no rating was assigned to CUS category 1. Gemini and Grok had descriptively higher expert ratings than ChatGPT, but CUS did not differ significantly across platforms (p = 0.157), and although modified DISCERN differed globally (p = 0.038), no pairwise comparison remained significant after adjustment. In the single-query diagnostic assessment, at least one platform missed the reference diagnosis in 9/15 vignettes, and all three missed porphyria. Gemini generated the longest responses, whereas Grok produced the most linguistically complex text; neither response length nor readability was associated with expert ratings. Conclusions: The three multimodal LLMs produced predominantly clinically acceptable responses, but performance varied by vignette and platform. Because each vignette-platform combination was sampled once, diagnostic findings represent single-response observations rather than stable platform accuracy estimates. Clinical verification remains necessary.