Related Experiment Video
Updated: Jan 12, 2026

Drug-Induced Sleep Endoscopy DISE with Target Controlled Infusion TCI and Bispectral Analysis in Obstructive Sleep Apnea
Published on: December 6, 2016
Experts V/S AI´s 2.0: Comparative evaluation of AI models and expert consensus in obstructive sleep apnea assessment
Giovanni Cammaroto1,2, Felipe Ahumada Mira2,3, Valentin Favier2,4
1Head and Neck Department, ENT & Oral Surgery Unity, G.B. Morgagni, L. Pierantoni Hospital, Forlì, Italy.
Purpose:
This study aims to compare the evaluation of obstructive sleep apnea (OSA) by ten super-experts using responses from a 10-question survey answered by 3 different artificial intelligence chatbots, Chat GPT-3.5, Chat GPT-4.0, Gemini, and a panel of 100 otolaryngologists specialized in sleep medicine.
Methods:
A 10-question survey regarding OSA management was answered by Chat GPT-3.5, Chat GPT-4.0, Gemini, and a panel of 100 otolaryngologists. The responses were assessed by ten super-experts in sleep medicine for their agreement with expert consensus, using a Likert scale. Statistical analyses were performed to evaluate the level of agreement and significance.
Result:
Expert consensus had the highest mean score (4.5 ± 0.9), significantly outperforming all AI models. ChatGPT-3.5 was the best among AI systems, with a score of 4.1 ± 1.2 (p=0.003), followed by ChatGPT-4 with 3.9 ± 1.4 (p<0.001) and Gemini with 3.6 ± 1.5 (p<0.001). Perfect agreement with expert consensus was achieved in specific scenarios, particularly regarding indications for bariatric surgery and lateral pharyngoplasty. However, there were significant differences in complex clinical scenarios that required integration of multiple factors, particularly in therapeutic management questions where the performance of AI models was significantly below that of expert consensus (p<0.01).
Conclusions:
Although AI models are promising in the management of OSA, especially for well-defined clinical scenarios, they at present serve best as complementary tools rather than replacements for expert clinical judgment. Most surprisingly, ChatGPT-3.5 outperformed its newer versions in many aspects, indicating that model updates with general capabilities may not always lead to better performance in specialized medical domains. These findings emphasize the potential of AI as a supportive resource while emphasizing the continuing need for human expertise in complex clinical decision-making.
More Related Videos
Related Concept Videos
Sleep Apnea
The condition is more prevalent among...
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
Cardiopulmonary Resuscitation II: ACLS Airway Management
Physical Assessment of the Respiratory Tract II: Inspection
Chest Configuration
The chest configuration...
Assessment of Ventilation I: Respiratory Rate
A Ventilation assessment is critical for monitoring a patient's health status. Respiration, one of the most accessible vital signs, provides insights into the function of numerous body systems and can indicate serious health issues, such as brainstem injuries from head trauma.
Critical Guidelines for Assessing Ventilation:

