Related Experiment Video
Updated: Jun 28, 2026

07:12
A Gaze-Contingent Display Framework for Perceptual Learning Research with Simulated Central Vision Loss
Published on: April 11, 2025
Benchmarking and fine-tuning vision-language models on a visual question answering dataset for myopic maculopathy
Tsun Hei Yip1, Pusheng Xu1, Zirong Liu1
1Department of Ophthalmology, LKS Faculty of Medicine, The University of Hong Kong, Hong Kong, China.
Summary
A new visual question answering (VQA) dataset for myopic maculopathy (MM) was created to train and assess vision-language models (VLMs). The fine-tuned InternVL3-8B model demonstrated superior performance, highlighting the potential for specialized VLMs in ophthalmology.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Imaging
Background:
- Myopic maculopathy (MM) poses a significant challenge in ophthalmology.
- Developing advanced tools for diagnosing and monitoring MM is crucial.
- Vision-language models (VLMs) show promise in interpreting medical images.
Purpose of the Study:
- To establish a comprehensive visual question answering (VQA) dataset specifically for myopic maculopathy (MM).
- To facilitate the fine-tuning and evaluation of vision-language models (VLMs) for MM-related tasks.
- To advance AI-driven diagnostic capabilities in ophthalmology.
Main Methods:
- A cross-sectional study was conducted using colour fundus photographs (CFPs).
- A novel MM-VQA dataset was constructed, including clinical captions and question-answer pairs (TFQ and OEQ), generated using GPT-5 and manually verified.
- The InternVL3-8B model was fine-tuned on this dataset and benchmarked against other leading VLMs.
Main Results:
- The MM-VQA dataset contains 2,591 CFPs and 19,648 question-answer pairs.
- The fine-tuned InternVL3-8B achieved a high overall accuracy of 0.746, outperforming several other models.
- The model demonstrated strong performance in both true/false questions (0.919 accuracy) and open-ended questions (0.572 accuracy).
Conclusions:
- The developed MM-VQA dataset is a valuable resource for advancing VLM research in ophthalmology.
- This dataset supports the creation of specialized VLMs for improved myopic maculopathy assessment.
- The findings underscore the potential of fine-tuned VLMs in clinical ophthalmology applications.

