Related Experiment Video
Updated: Sep 4, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models using the Type 2 Diabetes Health Education guideline: a comparative analysis of
Zhaoxia Huang1, Yuxin Bai2, Junyue Luo2
1Department of Ophthalmology, The Affiliated Hospital of Southwest Medical University, Luzhou, China.
Objective:
To systematically evaluate the quality and readability of health information generated by four large language models (LLMs) in response to inquiries regarding type 2 diabetes mellitus (T2DM), using an authoritative Chinese clinical guideline as the reference standard.
Methods:
A total of 124 standardized questions were extracted from the Chinese Type 2 Diabetes Popular Science Guidelines. Six endocrinologists and diabetes specialists conducted independent, blind evaluations using the CLEAR tool (Completeness, Lack of false Information, Evidence, Appropriateness, Relevance) and PEMAT-P (Patient Education Materials Assessment Tool for Printable materials). Response characteristics were also recorded. Between-model differences were tested using the Kruskal-Wallis H test with Bonferroni pairwise comparisons.
Results:
All four models achieved total CLEAR scores within the "very good" range (19-25), with no significant differences seen between models (χ2 = 1.985, p = 0.576). No significant differences were observed in the dimensions of Lack of false information (χ2 = 7.644, p = 0.054), Evidence (χ2 = 2.309, p = 0.511), and Relevance (χ2 = 7.516, p = 0.057). However, significant differences emerged in Completeness (χ2 = 47.661, p < 0.001) and Appropriateness (χ2 = 88.360, p < 0.001). Claude-4.0 received the lowest score in Completeness (median 4.00, IQR 3.00-5.00) but achieved the highest ranking in Appropriateness (median 4.00, IQR 4.00-5.00). On the PEMAT-P, understandability differed significantly across models (χ2 = 159.120, p < 0.001), yet all models surpassed the 70% threshold, with ChatGPT-4.1 highest (median 91.91%, IQR 91.91-100.00%). However, despite significant differences among the various models (χ2 = 354.023, p < 0.001), only ERNIE Bot 4.5 Turbo (median 75.00%, IQR75.00-75.00%) surpassed the 70% threshold, with no single model demonstrating consistent superiority across all dimensions.
Conclusion:
Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education. Future development should prioritize stronger step-by-step behavioral guidance and differentiated, scenario-specific model deployment to enhance their value in patient-facing diabetes self-management support.