Related Experiment Video
Updated: Jul 30, 2026

09:22
Establishing a Porcine Ex Vivo Cornea Model for Studying Drug Treatments against Bacterial Keratitis
Published on: May 12, 2020
6.1K
To Evaluate the Efficacy of Zero-Shot Prompting Using Large Language Models in the Extraction of Microbial Keratitis
Lokeshwari Aruljyothi1, Tushar Mungle2, Maria A Woodward3
1Department of Cornea and Refractive Services, Aravind Eye Hospital, Salem, Tamil Nadu, India.
Cornea
|December 11, 2025
Summary
Large Language Models (LLMs) like GPT-4o effectively extract microbial keratitis (MK) descriptors from clinical notes, showing good agreement with human experts. LLM performance depends on electronic health record data quality.
Area of Science:
- Ophthalmology
- Medical Informatics
- Artificial Intelligence
Background:
- Microbial keratitis (MK) is a significant cause of vision loss.
- Accurate extraction of MK descriptors from clinical notes is crucial for patient care and research.
- Electronic health records (EHRs) contain valuable clinical information but often require sophisticated methods for data extraction.
Purpose of the Study:
- To evaluate the efficacy of Large Language Models (LLMs) in extracting microbial keratitis (MK) descriptors from clinician notes.
- To compare LLM-generated descriptors with those identified by expert human annotators using a zero-shot prompting approach.
- To assess the performance of GPT-4o and GPT-4o mini in identifying MK characteristics such as centrality, infiltrate depth, and thinning.
Main Methods:
- A dataset of 215 patients with culture-proven MK was analyzed.
- Free-text clinical notes from the first corneal examination were used for descriptor extraction.
- GPT-4o and GPT-4o mini were prompted to identify centrality, infiltrate depth, and thinning, with results compared to expert consensus using Cohen Kappa scores, sensitivity, and specificity.
Main Results:
- GPT-4o achieved high mean sensitivity (92-97%) and specificity (93-99%) for MK descriptors.
- Cohen Kappa scores for GPT-4o indicated good agreement with human annotations (0.73-0.88).
- GPT-4o mini showed lower performance than GPT-4o, though not statistically significant, highlighting potential trade-offs in LLM architectures.
Conclusions:
- LLMs, including GPT-4o and GPT-4o mini, demonstrate substantial agreement with human annotations for extracting MK descriptors.
- The accuracy of LLM-based extraction is influenced by the quality and consistency of EHR documentation.
- Consideration of the efficiency-performance trade-off is essential when implementing LLMs for large-scale analysis of MK data.

