Related Experiment Video
Updated: Aug 31, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing Guideline Knowledge Alignment of Large Language Models in Ophthalmology: A Preclinical Benchmarking Study
Hongxia Lu1,2, Yan Huo2,3, Ruisi Xie2,3
1Clinical College of Ophthalmology, Tianjin Medical University, Tianjin, China, tijmu.edu.cn.
Purpose:
To evaluate the guideline knowledge alignment of two large language models (LLMs), GPT-5.5 Instant and DeepSeek-V4, and to determine their preclinical reliability as reference tools in refractive surgery.
Methods:
Using the 38 evidence-based recommendations of the international keratorefractive lenticule extraction (KLEx) guidelines as the gold standard, both LLMs were evaluated in their default configurations. Two ophthalmologists independently assessed the clinical safety and medical accuracy of the model responses using a 5-point Likert scale (1-5 points). Agreement between each model's recommendation strength and the guideline was quantified by intraclass correlation coefficient (ICC), structural reliability by the DISCERN scale, and readability by the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL) indices.
Results:
Likert ratings did not differ between GPT-5.5 Instant (4.96 ± 0.206) and DeepSeek-V4 (4.89 ± 0.385; p = 0.134). Across 114 independent generations, the ICC for agreement with the guideline was 0.884 (95% CI, 0.836-0.918) for GPT-5.5 Instant and 0.739 (0.640-0.813) for DeepSeek-V4 (both p < 0.001). DISCERN scores were 70.18 ± 5.16 and 68.05 ± 5.41 (p = 0.085); FRE, 9.93 ± 8.33 and 3.74 ± 5.24 (p < 0.001); and FKGL, 16.07 ± 2.21 and 18.95 ± 1.96 (p < 0.001).
Conclusion:
Both models aligned closely with the guidelines, with significant concordance in GRADE-based recommendation strength. They may serve as preclinical reference tools in refractive surgery but not as a substitute for specialist judgment.
