Related Experiment Video
Updated: Oct 10, 2026

Screening of Axonal Degeneration in Carpal Tunnel Syndrome Using Ultrasonography and Nerve Conduction Studies
Published on: January 11, 2019
Reliability, quality, and readability of large language models in carpal tunnel syndrome: A comparative
Bilgehan Kolutek Ay1, Cem Zafer Yıldır1
1Department of Physical Medicine and Rehabilitation, Sütçü İmam University Faculty of Medicine, Kahramanmaraş, Türkiye.
Abstract:
BackgroundAI-based chatbots are increasingly used for health information, while cross-linguistic studies suggest that response performance may vary by language. Given the high prevalence of carpal tunnel syndrome (CTS), this study evaluated the reliability, quality, patient education suitability, and readability of AI-generated responses to CTS questions in English and Turkish.Materials and methodsSeventeen patient-centered CTS questions were directed to ChatGPT® (GPT-5 Mini), Gemini® 3 Flash, and Claude® 4.5 Sonnet, using zero-shot prompting across 102 independent sessions in English and Turkish. Responses were evaluated for reliability, quality, and patient education using mDISCERN (MD), Global Quality Score (GQS), and the Patient Education Materials Assessment Tool (PEMAT). Readability was assessed via Flesch Reading Ease (English) and Ateşman Index (Turkish). Friedman and Wilcoxon tests with Bonferroni correction were applied.ResultsSignificant inter-model differences were observed in English across all metrics (p < 0.05), with Gemini achieving higher MD and PEMAT scores than ChatGPT and Claude (p < 0.01). Cross-language differences were observed in mDISCERN scores, while Gemini showing a significant difference in favor of English (r = 0.93; p < 0.001). None of the models met recommended readability levels in either language.ConclusionAI-generated CTS information varied by model and language. Differences in MD scores should be interpreted cautiously because this instrument partly reflects the provision of references and does not directly assess factual clinical accuracy. None of the evaluated models achieved recommended readability levels for patient education. These findings support the use of AIs as supplementary rather than standalone sources of CTS-related patient information.