Related Experiment Video
Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of ophthalmic large language models: quantitative vs. qualitative methods
Ting Fang Tan1, Arun J Thirunavukarasu2, Chrystie Quek1
1Singapore National Eye Centre, Singapore Eye Research Institute, Singapore, Singapore.
Purpose Of Review:
Alongside the development of large language models (LLMs) and generative artificial intelligence (AI) applications across a diverse range of clinical applications in Ophthalmology, this review highlights the importance of evaluation of LLM applications by discussing evaluation metrics commonly adopted.
Recent Findings:
Generative AI applications have demonstrated encouraging performance in clinical applications of Ophthalmology. Beyond accuracy, evaluation in the form of quantitative and qualitative metrics facilitate a more nuanced assessment of LLM output responses. Several challenges limit evaluation including the lack of consensus on standardized benchmarks, and limited availability of robust and curated clinical datasets.
Summary:
This review outlines the spectrum of quantitative and qualitative evaluation metrics adopted in existing studies, highlights key challenges in LLM evaluation, to catalyze further work towards standardized and domain-specific evaluation. Robust evaluation to effectively validate clinical LLM applications is crucial in closing the gap towards clinical integration.

