Related Experiment Video
Updated: May 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessment of frontier Large Language Models in sleep medicine
Anshum Patel1, Het Contractor1, Hayden Heninger1
1Division of Pulmonary, Allergy and Sleep Medicine, Mayo Clinic, Jacksonville, FL, United States.
Frontiers in Digital Health
|May 15, 2026
Summary
Frontier large language models (LLMs) like ChatGPT-5 and Grok-4 excel at sleep medicine knowledge recall but struggle with complex differential diagnosis generation. Performance was similar between models, indicating current LLMs are better for focused support than broad hypothesis generation.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Diagnostics
- Sleep Medicine
Background:
- Large language models (LLMs) are increasingly evaluated for medical applications.
- Frontier LLMs show promise but require rigorous assessment in specialized domains.
- Sleep medicine presents unique diagnostic reasoning challenges.
Purpose of the Study:
- To compare the diagnostic reasoning and knowledge recall capabilities of ChatGPT-5 and Grok-4 in sleep medicine.
- To assess LLM performance on clinical vignettes and multiple-choice questions relevant to sleep disorders.
Main Methods:
- Two LLMs, ChatGPT-5 and Grok-4, were tested on 79 sleep medicine case vignettes and 897 multiple-choice questions.
- Performance metrics included exact match for diagnosis, F1-score for differential diagnosis, and proportion correct for MCQs.
- Statistical comparison used the Mann-Whitney U test.
Main Results:
- Both models achieved high accuracy for final diagnosis (92.4%) and MCQs (ChatGPT-5: 93.0%, Grok-4: 92.8%).
- Differential diagnosis generation yielded modest F1-scores (ChatGPT-5: 0.55, Grok-4: 0.59).
- No statistically significant performance differences were observed between the two LLMs.
Conclusions:
- Frontier LLMs demonstrate strong knowledge recall but limited complex clinical reasoning in sleep medicine.
- Current general-purpose LLMs are more suitable for focused knowledge support than broad diagnostic hypothesis generation.
- Future research should explore domain-adapted models and clinician-in-the-loop systems to enhance LLM utility in clinical practice.
