Related Experiment Video
Updated: Jun 1, 2025

07:44
In vivo Structural Assessments of Ocular Disease in Rodent Models using Optical Coherence Tomography
Published on: July 24, 2020
2.8K
Assessing the possibility of using large language models in ocular surface diseases
Qian Ling1, Zi-Song Xu1, Yan-Mei Zeng1
1Department of Ophthalmology, the First Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang 330006, Jiangxi Province, China.
International Journal of Ophthalmology
|January 20, 2025
Summary
Large language models (LLMs) show promise in ophthalmology. GPT-4 excels in answering ocular surface disease questions, outperforming other LLMs and nearing expert human performance.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Education
Background:
- Ocular surface diseases (OSDs) encompass a range of conditions affecting the eye's front surface.
- Accurate diagnosis and management of OSDs require specialized knowledge.
- Large language models (LLMs) are emerging as potential tools in medical information retrieval and education.
Purpose of the Study:
- To evaluate the accuracy and performance of five different LLMs in answering specialized questions on ocular surface diseases.
- To compare LLM performance against human ophthalmology trainees and physicians.
Main Methods:
- A 100-question multiple-choice exam on OSDs was developed by ophthalmology professors.
- Five LLMs (ChatGPT-4, ChatGPT-3.5, Claude 2, PaLM2, SenseNova) were tested.
- LLM responses were compared to those of human groups (chief physician, attending physician, trainee, graduate student).
Main Results:
- ChatGPT-4 demonstrated the highest performance among all tested LLMs.
- ChatGPT-4's overall score surpassed all human groups except potentially chief physicians.
- ChatGPT-4 and PaLM2 showed high accuracy and credibility, with ChatGPT-4 achieving a 59% success rate.
Conclusions:
- ChatGPT-4 exhibits superior performance in relevance and confidence for OSD-related questions.
- PaLM2 also demonstrated strong accuracy and confidence, ranking second.
- LLMs, particularly GPT-4, show significant potential as valuable educational and clinical resources in specialized fields like OSDs.

