Related Experiment Video
Updated: Feb 17, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.2K
BEnchmarking Large Language Models for Ophthalmology (BELO): An Expert-Curated Data Set and Evaluation Framework for
Sahana Srinivasan1,2, Xuguang Ai3, Thaddaeus Wai Soon Lo4
1Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Ophthalmology Science
|February 16, 2026
Summary
A new benchmark, BEnchmarking LLMs for Ophthalmology (BELO), evaluates ophthalmology large language models (LLMs) on knowledge and reasoning. GPT-5 showed top accuracy, while GPT-4o and Gemini 1.5 Pro excelled in qualitative expert reviews.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Informatics
Background:
- Current large language model (LLM) benchmarks in ophthalmology are limited, primarily focusing on accuracy.
- There is a need for a comprehensive evaluation that assesses both knowledge recall and reasoning abilities.
Purpose of the Study:
- To introduce BEnchmarking LLMs for Ophthalmology (BELO), a standardized, expert-validated benchmark for evaluating LLMs in ophthalmology.
- To assess the performance of various LLMs on ophthalmology-related knowledge and reasoning tasks.
Main Methods:
- Developed BELO through multiple rounds of expert review by 13 ophthalmologists.
- Curated 900 ophthalmology-specific multiple-choice questions from diverse medical datasets.
- Evaluated 8 LLMs using quantitative metrics (accuracy, macro-F1) and qualitative expert assessments.
Main Results:
- GPT-5 achieved the highest quantitative scores for accuracy (0.90) and macro-F1 (0.91).
- Qualitative expert evaluations showed GPT-4o rated highest for accuracy and readability, and Gemini 1.5 Pro for completeness.
- LLM performance on text-generation metrics indicated room for improvement in clinical reasoning.
Conclusions:
- BELO offers a robust, clinically relevant benchmark for evaluating LLMs in ophthalmology.
- The benchmark assesses both accuracy and reasoning capabilities of current and emerging LLMs.
- Future iterations will incorporate vision-based QA and clinical scenario management.
