Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Validity, reliability, and readability of Artificial Intelligence chatbots as public sources of information on
Mohamad Amin Pourhoseingholi1,2, Catherine Killan1,2, Sara Rafiee3
1Hearing Sciences, Mental Health and Clinical Neurosciences, School of Medicine, University of Nottingham, Nottingham, United Kingdom.
Objective:
To assess the validity, reliability and readability of four AI chatbots for hearing-health information.
Design And Study Sample:
Three audiologists created 100 questions covering adult hearing loss, paediatric hearing, hearing aids, tinnitus and cochlear implants (20 each). Questions were submitted twice to ChatGPT-3.5, Bing AI, Gemini and Perplexity. Answers were scored for factual accuracy and completeness on a five-point Global Quality Score. Validity was defined using low (score = 5) and high (score ≥ 4) thresholds. Internal consistency was estimated with Cronbach's α; readability with the Flesch Reading Ease Score (FRES) and Flesch-Kincaid Grade Level (FKGL). All scoring was completed independently by two blinded reviewers; discrepancies were resolved by consensus.
Results:
Under the low threshold ChatGPT-3.5 and Perplexity were most valid (84% and 79%); high-threshold validity fell to 37% and 34%. Perplexity had the highest overall reliability (α = 0.83) yet α dropped below 0.70 for cochlear-implant, tinnitus and hearing-aid questions. 84% percent of outputs were "Difficult"/"Very Difficult" and 68% read at college level.
Conclusions:
AI chatbots deliver generally accurate hearing-health content, but high-threshold accuracy, domain-specific reliability and readability remain suboptimal. They should supplement, not replace the professional counselling. Continued optimisation and external validation are needed before routine clinical recommendation.
