Related Experiment Video
Updated: May 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Accuracy, Safety, and Equity of Multimodal Large Language Models for Real-World Ocular Surface Disease
Zitao Liu1, Shenchuan Qian2, Xiaorui Bao2
1Cixi Biomedical Research Institute, Wenzhou Medical University, Zhejiang, China; National Clinical Research Center for Ocular Diseases, Eye Hospital, Wenzhou Medical University, Wenzhou, China; State Key Laboratory of Ophthalmology, Optometry and Visual Science, Eye Hospital, Wenzhou Medical University, Wenzhou, China.
Abstract:
Ocular surface diseases impose a substantial burden in resource-limited regions, where diagnostic access remains scarce. Although multimodal large language models (M-LLMs) show potential for clinical support, their real-world accuracy, safety, and equity have not been systematically evaluated. A comprehensive assessment was performed on five leading M-LLMs (GPT-5, Gemini-2.5 Pro, Gemini-2.5 Flash, GLM-4.5V, and Claude-Sonnet-4.5) for ocular surface disease diagnosis using a retrospective multimodal data set of 259 representative cases from Aksu, China, including anterior segment photographs and structured clinical data. Model predictions were compared with consensus diagnoses from three board-certified ophthalmologists. Diagnostic accuracy, disease-specific sensitivity, interrun stability, safety risks, demographic bias, and cost-effectiveness were analyzed. Overall accuracy ranged from 77.50% (GLM-4.5V) to 83.73% (Gemini-2.5 Pro). High sensitivity was observed for data-driven conditions, such as dry eye disease (91.5% to 100.0%), whereas morphologic reasoning failures were prominent, with pterygium sensitivity declining to 0.0% to 57.7%. Potentially unsafe recommendations appeared in 24% to 45% of outputs. Significant demographic biases were detected, including age-related shortcut learning in pseudophakia misclassification and sex disparities in Meibomian gland dysfunction diagnosis. Patient-level interrun stability was low (2% to 39%), and per-case costs varied 15-fold ($0.004-$0.060) without correlation to accuracy (P > 0.05). Current M-LLMs remain unsuitable for autonomous clinical deployment, highlighting the need for human oversight and supporting their near-term role in supervised triage workflows in underserved settings.