Related Experiment Video
Updated: Sep 14, 2025

Quantitative Fundus Autofluorescence for the Evaluation of Retinal Diseases
Published on: March 11, 2016
Assessing the proficiency of large language models on funduscopic disease knowledge
Jun-Yi Wu1, Yan-Mei Zeng2, Xian-Zhe Qian2
1Department of Ophthalmology, Wuhan Fourth Hospital, Wuhan 430033, Hubei Province, China.
Aim:
To assess the performance of five distinct large language models (LLMs; ChatGPT-3.5, ChatGPT-4, PaLM2, Claude 2, and SenseNova) in comparison to two human cohorts (a group of funduscopic disease experts and a group of ophthalmologists) on the specialized subject of funduscopic disease.
Methods:
Five distinct LLMs and two distinct human groups independently completed a 100-item funduscopic disease test. The performance of these entities was assessed by comparing their average scores, response stability, and answer confidence, thereby establishing a basis for evaluation.
Results:
Among all the LLMs, ChatGPT-4 and PaLM2 exhibited the most substantial average correlation. Additionally, ChatGPT-4 achieved the highest average score and demonstrated the utmost confidence during the exam. In comparison to human cohorts, ChatGPT-4 exhibited comparable performance to ophthalmologists, albeit falling short of the expertise demonstrated by funduscopic disease specialists.
Conclusion:
The study provides evidence of the exceptional performance of ChatGPT-4 in the domain of funduscopic disease. With continued enhancements, validated LLMs have the potential to yield unforeseen advantages in enhancing healthcare for both patients and physicians.
More Related Videos
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
08:54Author Spotlight: Understanding Age-Related Macular Degeneration Pathophysiology with QAF Workflow
Published on: May 26, 2023