BEnchmarking Large Language Models for Ophthalmology (BELO): An Expert-Curated Data Set and Evaluation Framework for

Sahana Srinivasan1,2, Xuguang Ai3, Thaddaeus Wai Soon Lo4

  • 1Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.

Ophthalmology Science
|February 16, 2026
PubMed
Summary

A new benchmark, BEnchmarking LLMs for Ophthalmology (BELO), evaluates ophthalmology large language models (LLMs) on knowledge and reasoning. GPT-5 showed top accuracy, while GPT-4o and Gemini 1.5 Pro excelled in qualitative expert reviews.

Related Concept Videos