Related Experiment Video
Updated: May 11, 2026

12:23
Dynamic Visual Tests to Identify and Quantify Visual Damage and Repair Following Demyelination in Optic Neuritis Patients
Published on: April 14, 2014
14.5K
Performance of Foundation Models vs Physicians in Textual and Multimodal Ophthalmological Questions.
Henry Rocha1, Yu Jeat Chong2, Arun James Thirunavukarasu3,4
1University of Cambridge School of Clinical Medicine, University of Cambridge, Cambridge, United Kingdom.
JAMA Ophthalmology
|November 13, 2025
Summary
Foundation models (FMs) show strong performance in ophthalmology text-based questions, matching expert ophthalmologists. However, their ability to interpret multimodal data like images remains limited, requiring further development for clinical use.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Education
Background:
- Growing interest in large language models (LLMs) for clinical knowledge and reasoning in ophthalmology.
- Limited research on the multimodal capabilities of LLMs in ophthalmology, specifically image and table interpretation.
- Need to assess advanced foundation models (FMs) in ophthalmology examination contexts.
Purpose of the Study:
- To evaluate the multimodal performance of seven leading foundation models (FMs) in answering ophthalmology examination questions.
- To compare FM performance against ophthalmology trainees and physicians.
- To assess both textual and multimodal question-answering capabilities.
Main Methods:
- Cross-sectional study using Fellowship of the Royal College of Ophthalmologists part 2 written examination preparation materials.
- Seven foundation models (GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet, Llama-3.2-11B, DeepSeek V3, Qwen2.5-Max, Qwen2.5-VL-72B) were evaluated.
- Performance measured by accuracy compared to textbook answers, with head-to-head comparisons against human physicians.
Main Results:
- For textual questions, Claude 3.5 Sonnet achieved the highest accuracy (77.7%), outperforming trainees and junior physicians, and comparable to expert ophthalmologists.
- GPT-4o demonstrated strong performance in textual questions (69.9%), surpassing older LLMs like GPT-4 and GPT-3.5.
- For multimodal questions, GPT-4o (57.5%) outperformed other FMs but was less accurate than expert ophthalmologists and trainees.
Conclusions:
- Current FMs show significant improvement in ophthalmological knowledge reasoning for textual queries, comparable to expert ophthalmologists.
- FMs show potential as medical assistants for textual ophthalmology questions.
- The multimodal capabilities of current FMs in ophthalmology remain limited and require further research and fine-tuning with diverse ophthalmic data.

