Related Experiment Video
Updated: Apr 14, 2026

07:11
Assessing Early Stage Open-Angle Glaucoma in Patients by Isolated-Check Visual Evoked Potential
Published on: May 25, 2020
6.9K
Comparative Performance of Gemini 3 Pro and GPT-5 Family Models on Ophthalmology Board-Style Questions.
Ryan S Shean1, Jayanth Kumar Mallapu2, Tathya Shah1
1Keck School of Medicine, University of Southern California, Los Angeles, California.
Ophthalmology Science
|April 13, 2026
Summary
Gemini 3 Pro demonstrated superior performance on ophthalmology board-style questions compared to GPT models. Further research is needed to address limitations in image-based and complex reasoning tasks.
Area of Science:
- Artificial Intelligence in Medicine
- Ophthalmology Education
- Large Language Models
Background:
- Large language models (LLMs) are increasingly used in medical education.
- Evaluating LLM performance on specialized board-style questions is crucial for assessing their utility.
Purpose of the Study:
- To compare the performance of Gemini and GPT large language models on ophthalmology board-style questions.
- To analyze performance variations based on subspecialty, cognitive complexity, and question type.
Main Methods:
- A cross-sectional evaluation of 12 large language model configurations using 500 multiple-choice ophthalmology questions.
- Questions were categorized by subspecialty, image vs. text, and cognitive complexity (first, second, third order).
- Performance metrics included accuracy, paired discordance, and analysis of variance, with human benchmarking.
Main Results:
- Gemini 3 Pro achieved the highest accuracy (94.0%), outperforming all GPT-5 family variants.
- Models performed significantly better on American Academy of Ophthalmology questions (94.4%) than StatPearls (81.9%).
- Accuracy decreased with increasing cognitive complexity and was notably lower for image-based questions (10-22 point decrement).
Conclusions:
- Gemini 3 Pro exhibits strong general-purpose performance on ophthalmology board-style questions.
- Persistent limitations exist for image-based and third-order cognitive complexity questions.
- Ongoing benchmarking with clinically relevant datasets is essential for advancing LLM capabilities in ophthalmology.

