Related Experiment Video
Updated: Sep 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Diagnostic performance of Large Language Models (LLMs) compared with physicians in sleep medicine
Anshum Patel1, Chad Ruoff2, Scott A Helgeson3
1Division of Pulmonary, Allergy and Sleep Medicine, Mayo Clinic, Jacksonville, FL, USA.
Background:
Artificial intelligence (AI), particularly large language models (LLMs) are increasingly being explored for diagnostic capabilities in medicine. Leveraging LLMs within clinical systems may augment clinicians' diagnostic reasoning. The diagnostic effectiveness of LLMs in sleep medicine remains unevaluated against expert performance in clinical case scenarios.
Objective:
To compare the diagnostic accuracy of three widely used LLMs and experienced sleep physicians on real-world clinical vignettes.
Methods:
Using sixteen diverse sleep disorder vignettes from the AASM Case Book (2019), each was independently presented to three LLMs (ChatGPT-4, Gemini 2.0, DeepSeek) and three board-certified sleep physicians. Differential diagnoses were compared to AASM reference lists (mean % matches), and final diagnoses were scored against the AASM final diagnosis (3-point Likert: 0 = No match, 1 = Partial match, 2 = Full match).
Results:
Analysis of differential diagnoses showed similar mean agreement percentages for ChatGPT-4 (76.7 %), Gemini 2.0 (77.7 %), DeepSeek (70.7 %), and physicians' average (72.9 %). Repeated measures ANOVA indicated no statistically significant difference in differential diagnostic accuracy between LLMs and physicians (p = 0.839). For final diagnoses, all three LLMs achieved an identical average concordance score (87.5 %), falling within the performance range of experienced physicians (81.3 %-96.9 %), indicating LLM diagnostic proficiency was comparable to experts on these case vignettes. Non-parametric Friedman testing showed no statistically significant difference among the individual entities (p = 0.602). Paired t-tests comparing average final diagnosis scores also showed no significant differences (p = 0.606).
Conclusions:
LLMs showed diagnostic performance comparable to experienced sleep clinicians, suggesting their potential as supplementary tools. Future research should explore broader applications and integration.
More Related Videos
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
04:33Author Spotlight: Unveiling the Connection Between Sleep Disorders and Cognitive Symptoms in Depression
Published on: April 26, 2024