Related Experiment Video
Updated: May 28, 2026

05:49
An Automated Squint Method for Time-syncing Behavior and Brain Dynamics in Mouse Pain Studies
Published on: November 1, 2024
Evaluation of Arabic-Language AI Chatbot Responses to Migraine-Related Questions: A Comparative Cross-Sectional Study
Danah Aljaafari1, Hussain Khalifa Aljumah2, Mujtaba Abbas Alzuwayr2
1Department of Neurology, College of Medicine, Imam Abdulrahman bin Faisal University, Dammam 34212, Saudi Arabia.
Journal of Clinical Medicine
|May 27, 2026
Summary
Large language models (LLMs) provide comparable migraine information in Arabic, but vary in clarity and source transparency. Clinical oversight is essential when using these AI tools for patient education.
Area of Science:
- Neurology
- Artificial Intelligence
- Medical Informatics
Background:
- Migraine is a prevalent and debilitating neurological condition.
- Patients increasingly use online resources, including AI chatbots, for health information.
- The reliability of Arabic-language AI responses for migraine FAQs is largely unexamined.
Purpose of the Study:
- To assess the reliability, quality, and accuracy of Arabic responses from four leading large language models (LLMs) to common migraine questions.
- To compare the performance of ChatGPT-4.1, Gemini 3 Flash, DeepSeek-V3.2, and Grok 4.1 in generating migraine information.
Main Methods:
- 25 migraine-related FAQs were curated from multiple sources.
- Responses were generated using four distinct LLMs, yielding 100 total answers.
- Expert neurologists evaluated responses using modified DISCERN (mDISCERN), Global Quality Scale (GQS), and an accuracy scale, with inter-rater reliability assessed via ICC.
Main Results:
- Significant variations in mDISCERN and GQS scores were found across LLMs (p < 0.001), with DeepSeek and Grok performing highest.
- No significant differences in accuracy were observed between the models (p = 0.072).
- Source transparency and communication of uncertainty showed the most notable differences between chatbots.
Conclusions:
- Arabic medical content generated by LLMs for migraine FAQs is generally comparable in accuracy.
- Significant differences exist in the clarity, source transparency, and uncertainty communication of AI-generated responses.
- While LLMs can aid patient education, clinical judgment and oversight are crucial for responsible use.
