Related Experiment Video
Updated: Aug 5, 2026

A Workflow to Quantitatively Determine Age-Related Macular Degeneration Lesion-Specific Variations in Fundus Autofluorescence
Published on: May 26, 2023
Mapping the Reliability-Readability Gap in the Education of Patients With Age-Related Macular Degeneration Across 6
Zhili Lu1, Haixing Cao1, Cong Ma1
1Department of Ophthalmology, First Affiliated Hospital of Dalian Medical University, 222 Zhongshan Road, Dalian City, Liaoning Province, Dalian, Liaoning, 116011, China, 86 18098876399.
Background:
Artificial intelligence-generated health information is increasingly used by patients, but its reliability, visible transparency indicators, and readability remain uncertain in specialized ophthalmic conditions such as age-related macular degeneration (AMD).
Objective:
This study aimed to evaluate and compare the informational reliability, visible transparency indicators, overall quality, and readability of responses generated by 6 publicly accessible large language models (LLMs) to AMD-related patient-facing prompts under a zero-shot, single-turn prompting scenario.
Methods:
Thirty English-language AMD-related prompts were curated from Google Trends, the 2023 Chinese AMD guideline, and the 2025 American Academy of Ophthalmology Preferred Practice Pattern. Chinese guideline-derived prompts were translated and reviewed before model querying. Each finalized prompt was entered verbatim into ChatGPT-5.1-auto, DeepSeek-v3.2, Gemini-2.5-Flash-Thinking, Grok 4, Claude-Sonnet 4.5, and Qwen3-Max between October 10 and November 25, 2025. Two senior ophthalmologists (ZL and XM) blinded to model identity independently scored all responses using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale, and Journal of the American Medical Association benchmark criteria, with adjudication for disagreements. Readability was assessed using 6 standard formulas against a sixth-grade benchmark. Between-model differences were analyzed using Friedman tests with Holm-adjusted pairwise comparisons.
Results:
A total of 180 responses were analyzed. Interrater agreement was substantial to near-perfect across reliability instruments (κ=0.72-0.97). No model met the recommended sixth-grade readability target. Grok 4 achieved the highest scores on reliability-related instruments, including DISCERN (mean 46.40, SD 7.43) and EQIP (mean 74.33, SD 9.07), whereas DeepSeek-v3.2 generated the most readable responses, with the highest Flesch Reading Ease Score (mean 48.23, SD 9.16) and lowest Flesch-Kincaid Grade Level (mean 9.95, SD 1.87). Significant between-model differences were observed across all reliability and readability metrics (all P<.001).
Conclusions:
Under zero-shot, single-turn prompting conditions, the evaluated public LLMs showed substantial model-dependent differences in AMD-related patient education quality and readability. No model met the sixth-grade readability benchmark, including those with comparatively stronger reliability performance. These findings support clinician oversight, readability optimization, and further evaluation before LLM-generated AMD information is used directly in patient-facing settings.

