Related Experiment Video
Updated: Jul 2, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for Ophthalmology Training in China: A Prospective Evaluation
Zuhui Zhang1, Changke Huang1, Xinxin Yu1
1National Clinical Research Center for Ocular Diseases, Eye Hospital, Wenzhou Medical University, Wenzhou, China.
Purpose:
This study explored large language models (LLMs) as a scalable solution to the global shortage and uneven distribution of ophthalmologists, particularly their actual effectiveness and potential risks in ophthalmic training.
Design:
This is a prospective study.
Subjects:
Eleven leading LLMs (ERNIE Bot 4.5 Turbo, Dou Bao, Tongyi Qianwen 3, DeepSeek-R1, ChatGPT-4o, Gemini 2.5 Flash, Tencent Yuanbao T1, HuatuoGPT II, Baichuan4-Turbo, Kimi k1.5, GLM-4) and 10 resident physicians (RPs) from a tertiary ophthalmology hospital in China.
Methods:
Phase 1: all LLMs were tested on the Chinese and English versions of the Chinese National Health Professional Technical Qualification Examination (Intermediate Level) in Ophthalmology (CNHPTQE-O). Phase 2: the best-performing LLM was used to assist the 10 RPs in 2 tasks: (1) answering the same CNHPTQE-O text questions and (2) classifying 4 types of keratitis images. Resident physicians completed unassisted and assisted phases with a 1-month washout period.
Main Outcome Measures:
The primary outcomes were accuracy (%) on the CNHPTQE-O for LLMs, and change in accuracy for RPs with versus without LLM assistance (text and image tasks). The secondary outcomes included postassessment survey ratings and confusion matrix analysis.
Results:
Several Chinese LLMs, especially ERNIE Bot 4.5 Turbo, demonstrated superior performance on the CNHPTQE-O, achieving accuracies of 98.00% (Chinese) and 86.50% (English). ERNIE Bot 4.5 Turbo significantly outperformed all RPs on the Chinese examination (P = 0.001). With LLM assistance, all 10 RPs passed the text examination; the mean accuracy improved from 60.75% to 79.00% (mean difference 18.25%, 95% confidence interval: 10.30%-26.20%, P = 0.001). Questionnaire feedback was positive. However, on the keratitis image task, LLM assistance did not improve RP accuracy (41.25% vs. 40.56%, P = 0.662); questionnaire feedback was markedly less favorable.
Conclusions:
Large language models possess a solid foundation in ophthalmic knowledge and can effectively enhance trainee performance in text-based assessments, demonstrating clear potential as a training aid. However, their limitations in image-assisted diagnostic tasks and the associated risk of "artificial ignorance" should not be overlooked.
Financial Disclosures:
The author has no/the authors have no proprietary or commercial interest in any materials discussed in this article.
