Related Experiment Video
Updated: Sep 10, 2025

07:12
Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
559
Specialized curricula for training vision language models in retinal image analysis.
Robbie Holland1, Thomas R P Taylor2, Christopher Holmes3
1Biomedical Image Analysis, Department of Computing, Imperial College London, London, United Kingdom. robbie.holland@stanford.edu.
NPJ Digital Medicine
|August 19, 2025
Summary
A new training curriculum significantly improved a specialized vision-language model (VLM) for age-related macular degeneration (AMD) tasks. The enhanced VLM now rivals junior ophthalmologists, outperforming general models like ChatGPT-4o.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Imaging Analysis
Background:
- Clinicians face significant time burdens reviewing medical images and transcribing findings.
- Foundation models integrating visual and textual data show potential for reducing workloads but clinical value is uncertain.
- Current general and medical vision-language models (VLMs) underperform ophthalmologists in age-related macular degeneration (AMD) tasks.
Purpose of the Study:
- To develop and evaluate a specialized training curriculum to optimize VLMs for clinical decision-making in ophthalmology.
- To assess the performance of a curriculum-enhanced VLM against general foundation models and human experts in AMD-related tasks.
Main Methods:
- A dedicated training curriculum was designed by domain specialists to optimize VLMs for clinical decision-making.
- The performance of the specialized model (RetinaVLM-Specialist) was compared against foundation medical VLMs and ChatGPT-4o in AMD disease staging and referral tasks.
- A reader study involving senior ophthalmologists evaluated the accuracy of reports generated by RetinaVLM-Specialist versus ChatGPT-4o.
Main Results:
- RetinaVLM-Specialist significantly outperformed foundation medical VLMs and ChatGPT-4o in AMD disease staging (F1: 0.63 vs. 0.33) and referral (0.67 vs. 0.50).
- The specialized VLM achieved performance comparable to junior ophthalmologists.
- Senior ophthalmologists found RetinaVLM-Specialist reports substantially more accurate than ChatGPT-4o reports (64.3% vs. 14.3%).
Conclusions:
- A curriculum-based training approach can effectively adapt foundation models for specialized medical applications.
- Optimized VLMs demonstrate potential to assist clinicians in real-world medical tasks, improving accuracy and efficiency.
- This study provides a blueprint for enhancing AI models in ophthalmology and other medical fields.
Related Concept Videos
Vision
55.3K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
55.3K
Visual System
685
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
685

