Related Experiment Video
Updated: Feb 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
BEnchmarking Large Language Models for Ophthalmology (BELO): An Expert-Curated Data Set and Evaluation Framework for
Sahana Srinivasan1,2, Xuguang Ai3, Thaddaeus Wai Soon Lo4
1Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Purpose:
Current benchmarks evaluating large language models (LLMs) in ophthalmology are narrow and disproportionately prioritize accuracy. We introduce BEnchmarking LLMs for Ophthalmology (BELO), a standardized evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BEnchmarking LLMs for Ophthalmology assesses ophthalmology-related knowledge and reasoning quality.
Subjects:
This study did not involve human participation.
Design:
Cross-sectional study.
Methods:
Using keyword matching and a fine-tuned PubMed Bidirectional Encoder Representations from Transformers model, we curated ophthalmology-specific multiple-choice questions (MCQs) from diverse medical data sets (Basic and Clinical Science Course [BCSC], Multi-Subject Multi-Choice Dataset for Medical domain [MedMCQA], Medical Question Answering [MedQA], Biomedical Semantic Indexing and Question Answering [BioASQ], and PubMed Question Answering [PubMedQA]). The data set underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by 3 senior ophthalmologists. To illustrate BELO's utility, we evaluated 8 LLMs (OpenAI o1, o3-mini, GPT-5, GPT-4o, DeepSeek-R1, MedGemma-4B, Llama-3-8B, and Gemini 1.5 Pro).
Main Outcome Measures:
The 8 LLMs were evaluated in terms using accuracy, macro-F1, and 5 text-generation metrics (Recall-Oriented Understudy for Gisting Evaluation, BERTScore, BARTScore, Metric for Evaluation of Translation with Explicit Ordering, and AlignScore). In a further evaluation involving human experts, 2 ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness.
Results:
BEnchmarking LLMs for Ophthalmology consists of 900 high-quality, expert-reviewed questions aggregated from 5 sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). To demonstrate BELO's utility, we conducted a series of benchmarking exercises. In the quantitative evaluation, GPT-5 achieved the highest accuracy (0.90, 95% confidence interval [CI]: 0.89-0.92) and macro-F1 score (0.91, 95% CI: 0.89-0.93). On the other hand, the models' performance on text-generation metrics varied and were generally suboptimal, with scores ranging from 20.4 to 72.0 (out of 100, excluding the BARTScore metric), indicating room for improvement in clinical reasoning. In expert evaluations, GPT-4o was rated highest for accuracy and readability, while Gemini 1.5 Pro scored highest for completeness. A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO data set will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models.
Conclusions:
BEnchmarking LLMs for Ophthalmology provides a robust clinically relevant benchmark for evaluating both the accuracy and reasoning capabilities of current and emerging LLMs in ophthalmology. Future BELO benchmarking efforts will expand to include vision-based question answering and clinical scenario management tasks.
Financial Disclosures:
The authors have no proprietary or commercial interest in any materials discussed in this article.
