Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
AI meets sleep surgery: assessing drug-induced sleep endoscopy interpretation with a large language model
Sholem Hack1, Shibli Alsleibi2,3, Shay Shemesh4
1City St George's University London School of Medicine, Program Delivered by University of Nicosia at the Chaim Sheba Medical Center, Ramat Gan, Israel.
A large language model shows promise in interpreting drug-induced sleep endoscopy (DISE) videos for obstructive sleep apnea, matching expert performance in most areas and providing safe, reproducible treatment recommendations.
Area of Science:
- Artificial intelligence in medicine
- Otolaryngology
- Sleep surgery
Background:
- Obstructive sleep apnea (OSA) diagnosis often relies on subjective interpretation of drug-induced sleep endoscopy (DISE) videos.
- Objective interpretation of DISE videos is crucial for effective OSA treatment planning.
Purpose of the Study:
- To assess the accuracy and safety of a large language model (LLM) in interpreting DISE videos.
- To compare LLM performance against expert human raters and clinical reports for OSA treatment recommendations.
Main Methods:
- A prospective, blinded study involving 16 adult patients undergoing DISE.
- Independent review of anonymized DISE videos and clinical data by sleep surgery experts, a resident, and an LLM.
- Evaluation of rater concordance with a reference standard using Cohen's kappa and intraclass correlation coefficients.
Main Results:
- The LLM demonstrated expert-level agreement in interpreting velum, oropharynx, tongue base, and jaw thrust response.
- Moderate agreement was observed for epiglottic collapse, similar to human raters.
- LLM-generated treatment recommendations were safe, clinically consistent, and highly reproducible.
Conclusions:
- A large language model can accurately interpret DISE videos and provide safe, clinically relevant treatment recommendations for OSA.
- The LLM's performance approximates that of expert clinicians in most interpretative domains.
- Further multicenter studies are warranted to validate these findings and assess generalizability.
More Related Videos
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
07:54Drug-Induced Sleep Endoscopy DISE with Target Controlled Infusion TCI and Bispectral Analysis in Obstructive Sleep Apnea
Published on: December 6, 2016