Related Experiment Video
Updated: Sep 15, 2025

07:54
Drug-Induced Sleep Endoscopy DISE with Target Controlled Infusion TCI and Bispectral Analysis in Obstructive Sleep Apnea
Published on: December 6, 2016
20.0K
Evaluating Locally Run Large Language Models for Obstructive Sleep Apnea Diagnosis and Treatment: A Real-World
Christopher Seifen1, Tilman Huppertz1, Katharina Bahr-Hamm1
1Sleep Medicine Center & Department of Otolaryngology, Head and Neck Surgery, University Medical Center Mainz, Mainz, Germany.
Nature and Science of Sleep
|July 15, 2025
Summary
Locally run large language models (LLMs) show potential for sleep medicine but require significant improvement for interpreting polysomnography (PSG) results. While accurate for treatment recommendations, LLM performance in diagnosing sleep apnea severity needs refinement before clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Sleep Science
- Medical Informatics
Background:
- Sleep medicine is resource-intensive, creating a need for AI solutions.
- Locally run large language models (LLMs) address data protection concerns for clinical AI implementation.
- This study evaluates the first use of local LLMs for real-world polysomnography (PSG) interpretation.
Purpose of the Study:
- To assess the performance of locally run LLMs in interpreting PSG data.
- To compare the diagnostic and therapeutic recommendations of LLMs against a sleep physician.
- To determine the feasibility of using local LLMs in clinical sleep medicine.
Main Methods:
- Thirty patients with suspected obstructive sleep apnea (OSA) were selected.
- Polysomnography (PSG) results were interpreted by a sleep physician and three local LLMs (Gemma2, Llama3, Mistral Nemo).
- Concordance was measured for diagnosis, OSA severity, and treatment recommendations (aPAP).
Main Results:
- LLM concordance for OSA severity ranged from 33% (Gemma2) to 50% (Llama3).
- Mistral Nemo achieved 90% concordance for automatic positive airway pressure (aPAP) recommendations.
- Gemma2 and Llama3 showed 83% concordance for aPAP recommendations.
Conclusions:
- Locally run LLMs offer data security benefits for sleep medicine applications.
- Current LLM performance in interpreting PSG data, particularly for diagnosis, requires substantial improvement.
- Further refinement and fine-tuning are necessary before routine clinical implementation of local LLMs for PSG interpretation.

