Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Feb 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.2K

BEnchmarking Large Language Models for Ophthalmology (BELO): An Expert-Curated Data Set and Evaluation Framework for

Sahana Srinivasan1,2, Xuguang Ai3, Thaddaeus Wai Soon Lo4

  • 1Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.

Ophthalmology Science
|February 16, 2026
PubMed

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Large language models perpetuate bias in palliative care: Development and analysis of the Palliative Care Adversarial Dataset (PCAD).

PLOS digital health·2026
Same author

Memorization in large language models in medicine prevalence characteristics and implications.

Nature communications·2026
Same author

Evaluating the Potential Impact of AI on Urinary Tract Infection Diagnosis in the Emergency Department Across Demographic Groups: Retrospective Cohort Study.

JMIR AI·2026
Same author

Intraoperative Optical Coherence Tomography Features of Epiretinal Human Amniotic Membrane Graft Under Different Tamponade Agents.

Journal of vitreoretinal diseases·2026
Same author

Effectiveness of a Hybrid Telemedicine Model for Glaucoma Management.

Journal of glaucoma·2026
Same author

Time and person sensitive foundation model for disease prediction and risk stratification.

NPJ digital medicine·2026
Summary

A new benchmark, BEnchmarking LLMs for Ophthalmology (BELO), evaluates ophthalmology large language models (LLMs) on knowledge and reasoning. GPT-5 showed top accuracy, while GPT-4o and Gemini 1.5 Pro excelled in qualitative expert reviews.

Area of Science:

  • Ophthalmology
  • Artificial Intelligence
  • Medical Informatics

Background:

  • Current large language model (LLM) benchmarks in ophthalmology are limited, primarily focusing on accuracy.
  • There is a need for a comprehensive evaluation that assesses both knowledge recall and reasoning abilities.

Purpose of the Study:

  • To introduce BEnchmarking LLMs for Ophthalmology (BELO), a standardized, expert-validated benchmark for evaluating LLMs in ophthalmology.
  • To assess the performance of various LLMs on ophthalmology-related knowledge and reasoning tasks.

Main Methods:

  • Developed BELO through multiple rounds of expert review by 13 ophthalmologists.
  • Curated 900 ophthalmology-specific multiple-choice questions from diverse medical datasets.
  • Evaluated 8 LLMs using quantitative metrics (accuracy, macro-F1) and qualitative expert assessments.
Keywords:
Benchmark data setExpert curatedLarge language modelsOphthalmological reasoningQuestion–answer

More Related Videos

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.6K

Related Experiment Videos

Last Updated: Feb 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.2K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.6K

Main Results:

  • GPT-5 achieved the highest quantitative scores for accuracy (0.90) and macro-F1 (0.91).
  • Qualitative expert evaluations showed GPT-4o rated highest for accuracy and readability, and Gemini 1.5 Pro for completeness.
  • LLM performance on text-generation metrics indicated room for improvement in clinical reasoning.

Conclusions:

  • BELO offers a robust, clinically relevant benchmark for evaluating LLMs in ophthalmology.
  • The benchmark assesses both accuracy and reasoning capabilities of current and emerging LLMs.
  • Future iterations will incorporate vision-based QA and clinical scenario management.