Related Experiment Video
Updated: Aug 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models for diabetic retinopathy multiple-choice question generation in clinical ophthalmic
Xue Qin1, Ping Song1, Zhipeng Yan1
1The Affiliated Eye Hospital of Nanjing Medical University, Nanjing, Jiangsu, China.
Background:
Large language models (LLMs) are increasingly used in medical education, but their ability to generate ophthalmic multiple-choice questions (MCQs) remains unclear. Diabetic retinopathy (DR), a core ophthalmic training topic, provides a framework for evaluating LLM-based item generation.
Methods:
Five publicly accessible LLMs completed 60 predefined DR MCQ tasks under a standardized Chinese single-turn prompt and blueprint, yielding 300 items. Evaluation included structural completeness, format compliance, keyed-answer accuracy, textual features, response time, and blinded expert ratings across six educational domains. Because the same 60 tasks were completed by all five models, between-model comparisons were performed using paired task-level analyses. Continuous and ordinal outcomes were compared using Friedman tests, followed by Bonferroni-corrected paired Wilcoxon signed-rank tests when appropriate. Inter-rater reliability was assessed using intraclass correlation coefficients, and Spearman analyses examined associations between output features and expert-rated quality.
Results:
Between-model differences were observed in all textual variables and response time (all Friedman test p < 0.001). Gemini 3 was fastest (6.47 ± 1.27 s), whereas Qwen3-Max-Thinking was slowest (21.99 ± 4.38 s). Gemini 3 and ChatGPT-5.4 produced the longest responses (239.67 ± 29.12 and 245.98 ± 39.61 characters, respectively). ICCs ranged from 0.794 to 0.859. Between-model differences were found for content rigor, clarity, distractor quality, cognitive-level alignment, overall usability, and mean score, whereas educational usefulness did not reach statistical significance. For overall usability, the overall Friedman test showed only a weak difference, and no pairwise comparison remained significant after Bonferroni correction. Gemini 3 showed the highest proportion of directly usable items (90.00%), followed by ChatGPT-5.4 (86.67%). Longer explanations were associated with higher expert-rated quality, whereas longer response time showed no quality advantage.
Conclusion:
All five LLMs generated structurally complete and format-compliant DR MCQ drafts, but differences remained in factual accuracy, expert-rated item quality, output style, and usability. Gemini 3 and ChatGPT-5.4 showed the most favorable balance between correctness and expert-rated usability, supporting LLMs as assisted item-generation tools rather than replacements for expert review.