Related Experiment Video
Updated: Jun 12, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating reasoning in multimodal large language models for ophthalmology: a bilingual benchmark study using
Houfa Yin1,2, Kaikai Zhao1,2,3, Danli Shi4,5
1Eye Center of Second Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou, Zhejiang, People's Republic of China.
The British Journal of Ophthalmology
|June 10, 2026
Summary
Vision-language large language models (LLMs) show promise in ophthalmology, with reasoning-enabled prompts improving performance and interpretability. Rigorous evaluation is crucial for safe application in education and clinics.
Area of Science:
- Artificial Intelligence in Medicine
- Ophthalmology Research
- Multimodal Learning
Background:
- Large language models (LLMs) excel in text but struggle with multimodal data integration, crucial for ophthalmology.
- The study addresses the underexplored area of vision-language LLMs in ophthalmic question-answering.
Purpose of the Study:
- To evaluate the accuracy and reasoning capabilities of multimodal LLMs on complex ophthalmology questions.
- To assess the impact of reasoning-enabled prompts on LLM performance and interpretability in ophthalmology.
Main Methods:
- Three multimodal LLMs (CLM-V, ChatGPT-5, MiniCPM-V 4.5) were tested on 316 bilingual ophthalmology questions with clinical vignettes and images.
- Models were evaluated using reasoning-enabled and disabled prompts, with accuracy measured against reference standards.
- Reasoning quality was assessed via automated scoring and expert review, including qualitative case analysis.
Main Results:
- Reasoning-enabled prompts numerically improved AI-assisted scores across all models and datasets.
- ChatGPT-5 consistently ranked highest, with substantial inter-rater agreement (κ=0.87) in human evaluations.
- Qualitative analysis indicated reasoning-enabled outputs were often more interpretable, though benefits varied by model and dataset.
Conclusions:
- Multimodal LLMs show potential for ophthalmic question-answering, with reasoning prompts enhancing interpretability and performance.
- Limitations in subspecialty robustness and image interpretation highlight the need for rigorous reasoning evaluation.
- Careful assessment of reasoning is essential for safe educational and clinical deployment of these AI tools.