Related Experiment Video
Updated: Jun 20, 2026

Inducement and Evaluation of a Murine Model of Experimental Myopia
Published on: January 22, 2019
Evaluation of large language models for providing educational information in orthokeratology care
Yangyi Huang1, Runhan Shi2, Can Chen1
1Eye Institute and Department of Ophthalmology, Eye & ENT Hospital, Fudan University, Shanghai 200031, China; NHC Key Laboratory of Myopia (Fudan University), Key Laboratory of Myopia, Chinese Academy of Medical Sciences, Shanghai 200031, China; Shanghai Research Center of Ophthalmology and Optometry, China; Shanghai Engineering Research Center of Laser and Autostereoscopic 3D for Vision Care, China.
Background:
Large language models (LLMs) are gaining popularity in solving ophthalmic problems. However, their efficacy in patient education regarding orthokeratology, one of the main myopia control strategies, has yet to be determined.
Methods:
This cross-sectional study established a question bank consisting of 24 orthokeratology-related questions used as queries for GTP-4, Qwen-72B, and Yi-34B to prompt responses in Chinese. Objective evaluations were conducted using an online platform. Subjective evaluations including correctness, relevance, readability, applicability, safety, clarity, helpfulness, and satisfaction were performed by experienced ophthalmologists and parents of myopic children using a 5-point Likert scale. The overall standardized scores were also calculated.
Results:
The word count of the responses from Qwen-72B (199.42 ± 76.82) was the lowest (P < 0.001), with no significant differences in recommended age among the LLMs. GPT-4 (3.79 ± 1.03) scored lower in readability than Yi-34B (4.65 ± 0.51) and Qwen-72B (4.65 ± 0.61) (P < 0.001). No significant differences in safety, relevance, correctness, and applicability were observed across the three LLMs. Parental evaluations rated all LLMs an average score exceeding 4.7 points, with GPT-4 outperforming the others in helpfulness (P = 0.004) and satisfaction (P = 0.016). Qwen-72B's overall standardized scores surpassed those of the other two LLMs (P = 0.048).
Conclusions:
GPT-4 and the Chinese LLM Qwen-72B produced accurate and beneficial responses to inquiries on orthokeratology. Further enhancement to bolster precision is essential, particularly within diverse linguistic contexts.
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
Introduction to Language of Pathophysiology ll

