Related Experiment Video
Updated: Aug 8, 2026

07:54
Drug-Induced Sleep Endoscopy (DISE) with Target Controlled Infusion (TCI) and Bispectral Analysis in Obstructive Sleep Apnea
Published on: December 6, 2016
Performance of nine large language models on real-world polysomnography interpretation for obstructive sleep apnea: a
Abdurrahman Koç1, Abdullah Enes Ataş2, Şebnem Yosunkaya3
1Department of Pulmonary Medicine, Meram State Hospital, Meram, Konya, 42090, Türkiye. abdurrahman.koc1@saglik.gov.tr.
Summary
Large language models show promise in interpreting sleep apnea reports, but performance varies by model and case complexity. Selected models may aid specialists in diagnosis and treatment recommendations.
Area of Science:
- Artificial Intelligence in Medicine
- Sleep Medicine Research
- Clinical Decision Support Systems
Background:
- Large language models (LLMs) show potential for medical applications, but their real-world performance in interpreting complex diagnostic reports like polysomnography (PSG) for obstructive sleep apnea (OSA) is not well-established.
- Previous studies were limited by small sample sizes, single-model comparisons, and narrow outcome metrics.
Purpose of the Study:
- To compare the diagnostic and therapeutic capabilities of nine contemporary LLMs in interpreting real-world PSG reports for OSA.
- To benchmark LLM performance against expert consensus in diagnosing OSA and generating treatment recommendations.
Main Methods:
- Retrospective analysis of 212 de-identified Turkish PSG reports from a university sleep laboratory, classified as simple or complex.
- Nine LLMs were prompted with English queries, and their outputs were evaluated by two sleep specialists using a four-dimensional rubric (diagnostic accuracy, recommendation quality, safety, parameter coverage).
- Statistical analysis using generalized estimating equations accounted for within-case clustering across 1,908 model-case assessments.
Main Results:
- Fully correct diagnostic rates ranged from 75.0% to 86.3%, with Claude 4.1 Opus, ChatGPT-5, and Gemini 2.5 Pro performing comparably at the top tier.
- All models showed decreased performance on complex OSA cases (11.8-25.7 percentage point decline).
- Safety rates exceeded 88% across all models; moderate OSA was the most challenging diagnosis, and hypoxemia particularly affected lower-ranked models' accuracy.
Conclusions:
- Contemporary LLMs demonstrate significant potential for interpreting PSG reports for OSA, but their performance is not yet perfect and varies by model and case complexity.
- Selected LLMs show promise as adjunctive decision-support tools for sleep specialists, requiring careful oversight in clinical practice.

