Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Qualitative Analysis01:10

Qualitative Analysis

674
Qualitative analysis is the process of identifying elements, ions, or compounds in an unknown sample. It is the first and most fundamental type of analysis based on the hierarchy of analytical goals. This hierarchy is significant as it provides a structured approach to scientific research, with qualitative analysis serving as the initial step, providing essential information before moving on to quantitative or other forms of analysis.
There are two main approaches to qualitative analysis:...
674

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A S.C.O.R.E. framework for evaluating open-ended responses from large language models in healthcare.

Cell reports. Medicine·2026
Same author

Regional and temporal trends in antimicrobial susceptibility among isolates from bacterial keratitis: a systematic review and meta-analysis.

The Lancet. Microbe·2026
Same author

Corrigendum to "Oculomics and AI: The eye as a biomarker for health span" [Asia-Pac J Ophthalmol 15 (1) (2026) 100282].

Asia-Pacific journal of ophthalmology (Philadelphia, Pa.)·2026
Same author

AI-induced never-skilling in medical education.

Nature medicine·2026
Same author

How to meaningfully evaluate AI in clinical medicine.

Nature medicine·2026
Same author

Application of artificial intelligence and robotics in ophthalmic practice.

Asia-Pacific journal of ophthalmology (Philadelphia, Pa.)·2026

Related Experiment Video

Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

680

Evaluation of ophthalmic large language models: quantitative vs. qualitative methods.

Ting Fang Tan1, Arun J Thirunavukarasu2, Chrystie Quek1

  • 1Singapore National Eye Centre, Singapore Eye Research Institute, Singapore, Singapore.

Current Opinion in Ophthalmology
|September 5, 2025
PubMed
Summary

Evaluating large language models (LLMs) and generative artificial intelligence (AI) in ophthalmology requires diverse metrics beyond accuracy. Standardized benchmarks and curated datasets are crucial for robust clinical validation and integration.

Keywords:
GPTevaluationgenerative artificial intelligencelarge language modelsophthalmology

More Related Videos

A Method to Quantify Visual Information Processing in Children Using Eye Tracking
09:47

A Method to Quantify Visual Information Processing in Children Using Eye Tracking

Published on: July 9, 2016

17.6K
Author Spotlight: Quantifying Rough Eye Phenotypes in Drosophila Models of Amyotrophic Lateral Sclerosis with Frontotemporal Dementia
05:25

Author Spotlight: Quantifying Rough Eye Phenotypes in Drosophila Models of Amyotrophic Lateral Sclerosis with Frontotemporal Dementia

Published on: October 4, 2024

1.0K

Related Experiment Videos

Last Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

680
A Method to Quantify Visual Information Processing in Children Using Eye Tracking
09:47

A Method to Quantify Visual Information Processing in Children Using Eye Tracking

Published on: July 9, 2016

17.6K
Author Spotlight: Quantifying Rough Eye Phenotypes in Drosophila Models of Amyotrophic Lateral Sclerosis with Frontotemporal Dementia
05:25

Author Spotlight: Quantifying Rough Eye Phenotypes in Drosophila Models of Amyotrophic Lateral Sclerosis with Frontotemporal Dementia

Published on: October 4, 2024

1.0K

Area of Science:

  • Ophthalmology
  • Artificial Intelligence
  • Medical Informatics

Background:

  • Large language models (LLMs) and generative artificial intelligence (AI) are increasingly applied in clinical ophthalmology.
  • Evaluating the performance of these AI tools is critical for their safe and effective integration into healthcare.
  • Current evaluation methods often focus solely on accuracy, potentially overlooking other crucial aspects of AI performance.

Purpose of the Study:

  • To review and highlight the importance of evaluation metrics for LLM and generative AI applications in ophthalmology.
  • To discuss commonly adopted quantitative and qualitative evaluation metrics.
  • To identify challenges hindering robust AI evaluation in this domain.

Main Methods:

  • Systematic review of existing literature on LLM and AI evaluation in ophthalmology.
  • Analysis of quantitative and qualitative metrics used in published studies.
  • Identification of common challenges and limitations in current evaluation practices.

Main Results:

  • Generative AI shows promising performance in various ophthalmology clinical applications.
  • Beyond accuracy, quantitative and qualitative metrics offer a more comprehensive assessment of LLM outputs.
  • Key challenges include a lack of standardized benchmarks and limited availability of high-quality clinical datasets.

Conclusions:

  • A spectrum of evaluation metrics exists, but standardization is lacking.
  • Addressing challenges in dataset curation and benchmark development is essential.
  • Robust, domain-specific evaluation is crucial for validating clinical AI applications and facilitating their integration into ophthalmology practice.