Related Experiment Video
Updated: Jul 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmark evaluation of multi-modal large language models for ophthalmic diagnosis in real world
Shoujun Huang1, Junhong Chen1, Jiaoman Wang2
1College of Mathematical Medicine, Zhejiang Normal University, Jinhua, China.
Abstract:
Multimodal large language models (MLLMs) are increasingly demonstrating substantial potential in the medical domain, particularly in image-intensive specialties such as ophthalmology. Although cutting-edge models like ChatGPT-4o and Qwen-VL 2.5 have shown strong performance on general-domain tasks, real-world clinical benchmarks for rigorously assessing their diagnostic capabilities in specialized medical contexts remain limited. To address this gap, we constructed a carefully curated benchmark dataset comprising 295 pathologically confirmed ophthalmic cases with representative clinical presentations. Using this dataset, we systematically evaluated nine leading MLLMs, including both open-source and proprietary models. The results showed that models such as HAIBU-ReMUD and ChatGPT-4o achieved comparatively strong diagnostic accuracy and consistency, with performance in some settings approaching that of human experts. These findings suggest that current MLLMs are showing encouraging feasibility for real-world clinical applications and provide a basis for further investigation of their integration into ophthalmology practice.