Related Experiment Video
Updated: Aug 5, 2026

Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
LLMs in Medical Education for Autism Caregivers: A Comparative Evaluation of Accuracy, Readability, Actionability,
Shahid Akhtar Akhund1, Asma Alsaleh2, Naheed Haroon Kazi3
1Department of Medical Education and Anatomy and Genetics, College of Medicine, Alfaisal University, Riyadh 11533, Saudi Arabia.
Abstract:
Background: Family caregivers of children with autism spectrum disorder (ASD) increasingly utilize large language models (LLMs) for health information. This study presents a systematic comparative evaluation of three widely used LLMs as localized ASD health information tools in Saudi Arabia. Methods: Twenty-four clinically validated, caregiver-oriented questions were posed to Google Gemini 1.5 Pro, OpenAI ChatGPT (GPT-4o), and DeepSeek-V3 using a standardized prompt. Three expert raters independently evaluated responses across four dimensions: scientific accuracy, PEMAT-P understandability, PEMAT-P actionability, and neurodiversity (ND)-affirming language. Readability was assessed via Flesch-Kincaid Grade Level (FKGL) and SMOG indices. Non-parametric Kruskal-Wallis tests with post hoc Mann-Whitney U comparisons and one-sample t-tests were applied. Results: Gemini achieved the highest mean accuracy (2.96/3.00), significantly outperforming DeepSeek (p = 0.003, r = -0.37). Accuracy failures across all LLMs clustered on regional epidemiological, genetic risk, and financial inquiries. ChatGPT achieved significantly higher understandability than Gemini (p < 0.001, r = 0.55), while DeepSeek achieved significantly higher actionability than Gemini (p < 0.001, r = 0.59). However, all three LLM scores fell short of the Agency for Healthcare Research and Quality (AHRQ) 80% actionability benchmark (all p < 0.001). All LLMs exceeded patient education readability benchmarks (FKGL ≤ 6, SMOG ≤ 8; all p < 0.001); ChatGPT was the most readable (FKGL = 7.86; SMOG = 9.97) and Gemini the most complex. No model differed significantly on ND-affirming language, defaulting to a mixed medical-affirming register. Conclusions: Evaluated LLMs demonstrated distinct, specialized strengths: Gemini was the most accurate, ChatGPT the most readable, and DeepSeek the most actionable. Importantly, all models failed to meet established consumer education standards for readability and actionability. LLMs require extensive plain-language adaptation and cultural customization. Clinicians must guide families on navigating LLM outputs, particularly concerning country-specific epidemiological, economic, and healthcare service queries.
Related Concept Videos
Learning Disabilities
Dyslexia
Dyslexia is a...
Autism Spectrum Disorder
These core symptoms manifest differently among individuals, ranging from mild to severe. The disorder's complexity extends beyond its clinical presentation, encompassing a diverse range of biological, cognitive, and sociocultural influences.
Health Literacy
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
