Related Experiment Video
Updated: Aug 6, 2026

An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
Published on: March 13, 2026
Quality of Tinnitus Information From Generative AI Systems and Web Search: An Expert‑Rated Comparative Study
Sholem Hack1, Idit Tessler2,3, David Saltz4
1City St. George's University London School of Medicine, Program Delivered by University of Nicosia at the Chaim Sheba Medical Center.
Objective:
To compare the quality of tinnitus-related information generated by multiple generative artificial intelligence (GenAI) systems and web search using expert evaluation.
Study Design:
Cross-sectional comparative study.
Setting:
Digital platforms evaluated in their native public interfaces.
Patients:
Not applicable. Thirty commonly searched tinnitus-related questions derived from United States Google Trends data (2020-2025).
Interventions:
Questions were submitted to 6 GenAI systems (OpenEvidence, Claude, DeepSeek, GPT-5, Gemini, GPT-4) and Google Search (first organic result). Responses were independently rated by 6 experts using the QAMAI framework.
Main Outcome Measures:
Mean expert-rated quality scores across 5 domains (accuracy, clarity, relevance, completeness, and usefulness).
Results:
Overall quality differed significantly across systems (Friedman P<0.001; Kendall W=0.34). OpenEvidence achieved the highest mean score (4.45±0.72; 95% CI: 4.40-4.49), followed by Claude (4.00±1.02), DeepSeek (3.92±1.13), GPT-4 (3.89±0.84), Gemini (3.62±0.98), and GPT-5 (3.30±1.11). Google Search scored lowest (2.27±1.12; 95% CI: 2.20-2.35). Completeness was the lowest-performing domain across systems (range: 1.70-4.41). Pairwise comparisons showed significant differences between OpenEvidence and all other systems (effect size r=0.49-0.86). Inter-rater reliability was high (ICC=0.82). Readability demonstrated an inverse pattern relative to expert-rated quality. OpenEvidence demonstrated the lowest readability (Flesch-Kincaid Grade Level 17.5; Flesch Reading Ease 2.2), corresponding to a postgraduate reading level, whereas general-purpose LLMs produced more accessible responses at a sixth-seventh grade reading level.
Conclusions:
The quality of tinnitus information varies substantially across digital platforms. While GenAI systems generally outperform web search, deficiencies in completeness persist. Readability analysis revealed an inverse relationship between expert-rated quality and response accessibility, suggesting that clinician and patient assessments of informational value may not always align. These findings highlight the need for continued evaluation and clinician oversight to ensure safe, comprehensive, and accessible patient-facing information.
