Related Experiment Video
Updated: May 28, 2026

An Automated Squint Method for Time-syncing Behavior and Brain Dynamics in Mouse Pain Studies
Published on: November 1, 2024
Evaluation of Arabic-Language AI Chatbot Responses to Migraine-Related Questions: A Comparative Cross-Sectional Study
Danah Aljaafari1, Hussain Khalifa Aljumah2, Mujtaba Abbas Alzuwayr2
1Department of Neurology, College of Medicine, Imam Abdulrahman bin Faisal University, Dammam 34212, Saudi Arabia.
Abstract:
Background/Objectives: Migraine is a common and disabling neurological disorder, and many individuals increasingly seek information online. With the growing use of large language models (LLMs), such as ChatGPT, for patient education, concerns have emerged regarding the quality and reliability of the responses they generate, particularly in Arabic, where evidence remains limited. This study aimed to evaluate the reliability, quality, and accuracy of Arabic-language responses to frequently asked questions (FAQs) about migraine. Methods: A total of 25 FAQs were selected using a multisource approach and entered into four LLMs (ChatGPT-4.1, Gemini 3 Flash, DeepSeek-V3.2, and Grok 4.1), generating 100 responses. Responses were evaluated by a panel of expert neurologists using the modified DISCERN (mDISCERN), Global Quality Scale (GQS), and an accuracy scale. Inter-rater reliability was assessed using the intraclass correlation coefficient (ICC). Results: Significant differences were observed between chatbots for mDISCERN and GQS (both p < 0.001), whereas accuracy did not differ significantly across models (p = 0.072). DeepSeek and Grok demonstrated the highest mDISCERN scores (34.07 ± 1.31 and 34.29 ± 2.59, respectively), while DeepSeek achieved the highest GQS (4.95 ± 0.13). The clearest between-model differences were observed in source transparency and communication of uncertainty. Inter-rater reliability was good across all instruments (ICC range, 0.799-0.831). Conclusions: Medical content generated by the chatbots was broadly comparable, whereas important differences were observed in how that content was communicated. These tools may support patient education; however, their use should remain guided by clinical oversight and professional judgment.
