Related Experiment Video
Updated: Sep 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative performance of contemporary multimodal large language models in retinal imaging question answering
Xiaochen Gu1, Yizhou Yang2, Xuanqiao Lin2
1National Clinical Research Center for Ocular Diseases, Eye Hospital, Wenzhou Medical University, Wenzhou, China.
Background:
Contemporary multimodal large language models (LLMs) can process both clinical text and medical images, but their reliability in retinal imaging-based question answering remains uncertain. This study evaluated the performance of six contemporary multimodal LLMs on retinal multiple-choice questions derived from OCTCases.
Methods:
This cross-sectional benchmark study included 226 multiple-choice questions from 78 OCTCases retinal cases, comprising 151 image-based and 75 nonimage-based questions. Six multimodal LLMs were evaluated: ChatGPT-5.5 Instant, ChatGPT-5.5 Thinking, Grok 4, Gemini 3, DeepSeek V4, and Kimi 2.5. Model-selected answers were compared with the OCTCases answer key, which was reviewed by three retina specialists. Accuracy was assessed overall and by question subset. Pairwise model comparisons, inter-model agreement, voting-based group-answer performance, item difficulty distribution, and interface latency were analyzed.
Results:
Overall accuracy ranged from 59.3 to 77.4%, with the highest accuracy observed for ChatGPT-5.5 Thinking, followed by Gemini 3 and ChatGPT-5.5 Instant. DeepSeek V4 showed the lowest overall accuracy. All models performed better on nonimage-based questions than on image-based questions. In the image-based subset, Gemini 3 achieved the highest accuracy, whereas DeepSeek V4 showed the lowest accuracy. Under the majority-vote rule, group answers were generated for 182 of 226 questions and achieved an accuracy of 89.6% among questions with a determinate group answer. The plurality-vote rule generated group answers for more questions but with lower accuracy. Inter-model agreement was generally higher for nonimage-based than image-based questions, and all questions answered incorrectly by all six models were image-based. Interface latency varied across models and showed no consistent association with response correctness.
Conclusion:
Contemporary multimodal LLMs showed promising but uneven performance on OCTCases retinal questions. Their performance was consistently better on nonimage-based than image-based questions, indicating that retinal image interpretation remains a major limitation. Voting-based group answers may provide a useful reliability signal when model agreement is present, whereas model disagreement may help identify difficult or visually ambiguous questions requiring specialist review. These findings support continued benchmarking and cautious, specialist-supervised use of multimodal LLMs in retinal education and structured imaging question answering.