Related Experiment Video
Updated: May 2, 2026

Neuro-rehabilitation Approach for Sudden Sensorineural Hearing Loss
Published on: January 25, 2016
Can Chatbots Please Both Patients and Experts? Benchmarking AI and Clinical Guidelines for Hearing Loss
Sholem Hack1, Ben Gvili2, Idit Tessler2
1City St. Georges University London School of Medicine, Program Delivered by University of Nicosia at the Chaim Sheba Medical Center.
Objective:
To evaluate the reliability, accuracy, and clarity of responses generated by 7 contemporary artificial intelligence chatbots in answering patient-focused questions about age-related hearing loss and sudden sensorineural hearing loss, and to compare these outputs to expert-authored guideline responses as well as layperson ratings.
Study Design:
Cross-sectional.
Setting:
Academic medical center.
Patients:
Not applicable. Ten independent layperson raters, all over the age of 18, recruited from personal networks, assessed a subset of chatbot and expert responses.
Interventions:
Patient-centered questions, derived from official clinical practice guidelines for hearing loss, were submitted to 7 artificial intelligence chatbots from 3 major development groups. Responses were rated by a blinded panel of 5 otolaryngologists for accuracy, extensiveness, misleading content, quality of cited references, and overall reliability. A panel of 10 independent layperson raters, all over the age of 18, recruited from personal networks, assessed a subset of chatbot and expert responses.
Main Outcome Measures:
Proportion of chatbot answers rated fully accurate by expert panel; mean layperson clarity and trustworthiness scores; frequency of misleading information and high-quality references.
Results:
The most advanced chatbots achieved full guideline-concordant accuracy for up to 50% of questions, while earlier models ranged from 25% to 37.5%. All models performed highly for extensiveness and reference quality. Layperson ratings were highest for gold-standard expert answers, the latest chatbots approached these levels for both clarity and trustworthiness (mean scores: 4.7 to 4.8 out of 5; 95% CI: 4.67-4.85), and differences between models were of moderate-to-large effect size (η2=0.29 to 0.30). Misleading content was rare and typically not clinically significant.
Conclusions:
Modern artificial intelligence chatbots can provide clear and generally reliable patient education for hearing loss, but full guideline concordance remains inconsistent. Expert oversight is advised to ensure clinical accuracy.
Related Concept Videos
Current Trends in Nursing II
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...

