Related Experiment Video
Updated: Jul 12, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language
Zhenyue Qin1, Yang Liu2, Y U Yin3
1School of Medicine, Yale University, New Haven, CT, USA.
Summary
A new large-scale multimodal ophthalmology dataset, LMOD+, aids artificial intelligence (AI) development for diagnosing eye diseases. This benchmark dataset and evaluation pipeline aim to advance multimodal large language models (MLLMs) in ophthalmology.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Imaging
Background:
- Vision-threatening eye diseases present a significant global health burden.
- Diagnosis is hampered by workforce shortages, delays, and limited access to specialists.
- Artificial intelligence (AI), particularly multimodal large language models (MLLMs), shows promise for improving ophthalmic diagnostics.
Purpose of the Study:
- To address the lack of comprehensive benchmark datasets for developing and evaluating MLLMs in ophthalmology.
- To introduce LMOD+, a large-scale multimodal ophthalmology benchmark dataset.
- To facilitate the advancement of AI in ophthalmic applications.
Main Methods:
- Developed LMOD+, a dataset with 32,633 instances across 12 ophthalmic conditions and 5 imaging modalities.
- Integrated imaging, anatomical structures, demographics, and free-text annotations.
- Created a unified data curation pipeline for MLLM development and evaluated 24 state-of-the-art MLLMs.
Main Results:
- LMOD+ supports tasks including anatomical recognition, disease screening, staging, and demographic prediction.
- Evaluated MLLMs demonstrated potential in disease screening and anatomical recognition, with some achieving over 57% accuracy in zero-shot settings.
- Performance on challenging tasks like disease staging remained suboptimal, highlighting the gap between general MLLMs and specialized ophthalmic needs.
Conclusions:
- LMOD+ is a valuable resource for advancing MLLMs in ophthalmology.
- Current MLLMs show promise but require further development for complex ophthalmic tasks.
- Public release of the dataset, pipeline, and leaderboard aims to foster community-driven AI innovation in eye care.
