Related Experiment Video
Updated: Sep 14, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Diagnostic performance of Large Language Models (LLMs) compared with physicians in sleep medicine
Anshum Patel1, Chad Ruoff2, Scott A Helgeson3
1Division of Pulmonary, Allergy and Sleep Medicine, Mayo Clinic, Jacksonville, FL, USA.
Large language models (LLMs) demonstrate comparable diagnostic accuracy to sleep physicians in clinical scenarios. These AI tools show potential as valuable supplementary resources for medical diagnosis.
Area of Science:
- Medical Artificial Intelligence
- Clinical Diagnostics
- Sleep Medicine Research
Background:
- Artificial intelligence (AI), specifically large language models (LLMs), is being explored for medical diagnostic applications.
- LLMs may enhance clinicians' diagnostic reasoning when integrated into clinical systems.
- The diagnostic performance of LLMs in sleep medicine has not been evaluated against expert physicians.
Purpose of the Study:
- To compare the diagnostic accuracy of three prominent LLMs against experienced sleep physicians.
- The study utilized real-world clinical case vignettes for evaluation.
Main Methods:
- Sixteen diverse sleep disorder vignettes from the AASM Case Book (2019) were used.
- Three LLMs (ChatGPT-4, Gemini 2.0, DeepSeek) and three board-certified sleep physicians independently assessed the vignettes.
- Differential diagnoses were compared to AASM reference lists, and final diagnoses were scored using a 3-point Likert scale.
Main Results:
- LLMs showed comparable mean agreement percentages for differential diagnoses (70.7%-77.7%) to physicians (72.9%).
- No statistically significant difference was found in differential diagnostic accuracy between LLMs and physicians (p=0.839).
- LLMs achieved an average final diagnosis concordance score of 87.5%, within the range of expert physicians (81.3%-96.9%).
Conclusions:
- LLMs exhibit diagnostic performance comparable to experienced sleep clinicians.
- LLMs show potential as supplementary tools to aid in medical diagnosis.
- Further research is recommended to explore broader applications and integration of LLMs in clinical practice.
More Related Videos
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
04:33Author Spotlight: Unveiling the Connection Between Sleep Disorders and Cognitive Symptoms in Depression
Published on: April 26, 2024