Related Experiment Video
Updated: Aug 8, 2026

Drug-Induced Sleep Endoscopy (DISE) with Target Controlled Infusion (TCI) and Bispectral Analysis in Obstructive Sleep Apnea
Published on: December 6, 2016
Performance of nine large language models on real-world polysomnography interpretation for obstructive sleep apnea: a
Abdurrahman Koç1, Abdullah Enes Ataş2, Şebnem Yosunkaya3
1Department of Pulmonary Medicine, Meram State Hospital, Meram, Konya, 42090, Türkiye. abdurrahman.koc1@saglik.gov.tr.
Purpose:
The diagnostic and therapeutic capabilities of large language models in interpreting real-world polysomnography reports for obstructive sleep apnea remain insufficiently characterized, with prior studies limited by small samples, single-model designs, and unidimensional outcome measures. This study compared the clinical performance of nine contemporary large language models in diagnosing obstructive sleep apnea and generating treatment recommendations from polysomnography reports, benchmarked against expert consensus.
Methods:
Two hundred twelve polysomnography records from a university sleep laboratory were retrospectively classified as simple (n = 153) or complex (n = 59) based on AASM criteria. De-identified reports, originally in Turkish, were submitted to nine large language models using an English prompt within a one-week frozen evaluation window. Model outputs were independently scored by two sleep medicine specialists using a four-dimensional rubric encompassing diagnostic accuracy, recommendation quality, safety, and parameter coverage. Generalized estimating equations accounted for within-case clustering across 1,908 model-case assessments.
Results:
Fully correct diagnostic rates ranged from 75.0% to 86.3%, with Claude 4.1 Opus, ChatGPT-5, and Gemini 2.5 Pro forming a statistically indistinguishable top tier. All models demonstrated significant performance degradation on complex cases (11.8-25.7 percentage point decline). Safety rates exceeded 88% across all models. Moderate obstructive sleep apnea was the most challenging diagnostic category. Hypoxemia disproportionately impaired diagnostic accuracy in lower-ranked models.
Conclusion:
Contemporary large language models demonstrate promising yet imperfect capacity for real-world polysomnography report interpretation in obstructive sleep apnea, with performance varying by model and case complexity. These findings support selected models as potential adjunctive decision-support tools under specialist oversight.

